Back to catalog

Neurotatarlar

A volunteer open-source community building NLP infrastructure for the Tatar language - text corpora, speech recognition, and language models.

github.com

About this project

  • 14 repositories on GitHub, including the curated "awesome-tatar" list
  • Developing the "yazam.tatar" app for grammar checking
  • Aims to build the largest open corpus of Tatar texts for training language models

TurkicASR

Language project#Language#Technology#Speech recognition#Open sourceTatarsBashkirs
  • Recognizes speech in ten Turkic languages
  • Supports the Tatar and Bashkir languages
  • The model and code are published under the CC BY 4.0 license
Learn more

Bashkir Language Audio2Text

Language project#Language#Technology#Speech recognition#Open sourceBashkirs
  • Processes audiobooks, poetry, and ELAN field recordings
  • Converts materials into a format similar to Mozilla Common Voice
  • Prepares data for Whisper and Wav2Vec2
  • The collection includes 583 audio files of books and stories
Learn more

Fastmorph

Language project#Language#Technology#Language model#Open sourceTatars
  • Created for the Written Tatar Language Corpus
  • Searches by word forms, lemmas, grammatical tags, and patterns
  • Supports distance constraints between words
  • Source code published under the GPL-3.0 license
Learn more
Муэллиф