Back to catalog

Tugantel corpus of the Tatar language

A publicly available corpus of contemporary Tatar for research, teaching, and self-study.

tugantel.tatar

About this project

  • Search works by lexeme, word form, and grammatical features
  • Developed with staff of the Academy of Sciences of Tatarstan and Kazan Federal University
  • The corpus includes material from Tatar newspapers, magazines, and publishing houses

Aigiz Kunafin

AuthorLanguage project#Developer#Language#Technology#Text to speech#Language modelBashkirsTatars
  • Published 19 models and 13 datasets on Hugging Face for Bashkir and Tatar speech
  • Author of the tools bashspell, bashkort_translate_bot and bashkir-stress
  • Winner of the Enterprise RAG Challenge
Learn more

Fastmorph

Language project#Language#Technology#Language model#Open sourceTatars
  • Created for the Written Tatar Language Corpus
  • Searches by word forms, lemmas, grammatical tags, and patterns
  • Supports distance constraints between words
  • Source code published under the GPL-3.0 license
Learn more

Ilshat Säetov

AuthorLanguage project#Linguist#Language#Literature#Technology#Keyboard#Text recognition#Open sourceTatars
  • Associated member of CETOBaC at EHESS in Paris
  • Author of TATlit, a language model for literary Tatar
  • Author of an OCR model for Tatar Cyrillic and a macOS keyboard layout
Learn more
Муэллиф