Back to catalog

Bashkir Language Audio2Text

Tools for collecting and preparing a Bashkir speech corpus for speech-recognition models.

github.com

About this project

  • Processes audiobooks, poetry, and ELAN field recordings
  • Converts materials into a format similar to Mozilla Common Voice
  • Prepares data for Whisper and Wav2Vec2
  • The collection includes 583 audio files of books and stories

TurkicASR

Language project#Language#Technology#Speech recognition#Open sourceTatarsBashkirs
  • Recognizes speech in ten Turkic languages
  • Supports the Tatar and Bashkir languages
  • The model and code are published under the CC BY 4.0 license
Learn more

Neurotatarlar

Language project#Language#Society#Technology#Speech recognition#Open sourceTatars
  • 14 repositories on GitHub, including the curated "awesome-tatar" list
  • Developing the "yazam.tatar" app for grammar checking
  • Aims to build the largest open corpus of Tatar texts for training language models
Learn more

Bashkir Drama Corpus

Language project#Language#Technology#Open sourceBashkirs
  • Plays are provided in TEI-P5 format
  • Compatible with DraCor tools and API
  • Data is developed in an open repository
Learn more
Муэллиф