Back to catalog

Qirimtatar Embedding Crawlers

A set of crawlers for building a Crimean Tatar text corpus and training language models.

github.com

About this project

  • Collects texts from open web sources and digital archives
  • Uses Tesseract for text recognition in PDFs and images
  • Publishes the extracted documents separately from the source code
  • Processes the Qirim Junior archives and the texts of Shamil Aladin

Qırımtatar TTS Datasets

Language project#Language#Technology#Text to speech#Dataset#Open sourceCrimean Tatars
  • Apache 2.0 license, maintained by developer egorsmkv
  • About 8.8 hours of audio in OPUS 48 kHz format
  • Mirrored on Hugging Face with an interactive demo
Learn more

Qırımtatar TTS

Language project#Language#Technology#Text to speech#Open source#InteractiveCrimean Tatars
  • Converts Crimean Tatar text into speech
  • The interactive version runs on Hugging Face Spaces
  • Source code is published under the MIT license
Learn more

Bashkir Drama Corpus

Language project#Language#Technology#Open sourceBashkirs
  • Plays are provided in TEI-P5 format
  • Compatible with DraCor tools and API
  • Data is developed in an open repository
Learn more
Муэллиф