Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
TilQazyna 's Collections
First Releases (2024)
Kazakh Terminology
Kazakh Morphology and POS
Speech and OCR
Til Web, Crawls and Archive
Til Instruct — Task Datasets
Tokenizers
Til 256k Research Ladder
Kazakh GEC — All Models
Til — Multilingual Models
Til Core — Kazakh-only Models
Til Flagship Datasets

Til Flagship Datasets

updated 25 days ago

Core datasets for pretraining, instruction tuning, speech and language tasks. Основные датасеты для предобучения, инструкций, речи и языковых задач.

Upvote
-

  • TilQazyna/Til-Corpus

    Updated 25 days ago • 126 • 1

    Note Start here for pretraining text.


  • TilQazyna/Til-Instruct

    Viewer • Updated 26 days ago • 6.26M • 18

  • TilQazyna/Til-Parallel

    Updated 26 days ago • 16

  • TilQazyna/Til-Audio

    Viewer • Updated 26 days ago • 380k • 23

  • TilQazyna/Til-GEC

    Viewer • Updated 25 days ago • 4.62M • 38

    Note Training data behind every GEC model in this collection.


  • TilQazyna/Til-Morphology

    Updated 26 days ago • 17 • 1

  • TilQazyna/Til-Terminology

    Viewer • Updated 26 days ago • 317k • 10

  • TilQazyna/Til-Classification

    Viewer • Updated 26 days ago • 91.8k • 13
Upvote
-
  • Collection guide
  • Browse collections
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs