Back to catalog

Tweety Tatar 7B

An open 7-billion-parameter language model with a native tokenizer for Tatar, part of the Tweeties research series for low-resource languages.

huggingface.co

About this project

  • Created by an international multilingual-NLP research group
  • Adapted via trans-tokenization and fine-tuned on a Tatar corpus
  • The series also includes specialized model variants

Aigiz Kunafin

AuthorLanguage project#Developer#Language#Technology#Text to speech#Language modelBashkirsTatars
  • Published 19 models and 13 datasets on Hugging Face for Bashkir and Tatar speech
  • Author of the tools bashspell, bashkort_translate_bot and bashkir-stress
  • Winner of the Enterprise RAG Challenge
Learn more

Fastmorph

Language project#Language#Technology#Language model#Open sourceTatars
  • Created for the Written Tatar Language Corpus
  • Searches by word forms, lemmas, grammatical tags, and patterns
  • Supports distance constraints between words
  • Source code published under the GPL-3.0 license
Learn more

Written Corpus of the Tatar Language

Language project#Language#Technology#Language modelTatars
  • The corpus contains hundreds of millions of Tatar words
  • Supports search by word form and grammatical features
  • Includes a thesaurus and a spell-checking service
Learn more
Муэллиф