Back to catalog

Bashkir Drama Corpus

A digital corpus of Bashkir drama within the DraCor ecosystem.

github.com

About this project

  • Plays are provided in TEI-P5 format
  • Compatible with DraCor tools and API
  • Data is developed in an open repository

Bashkir Language Audio2Text

Language project#Language#Technology#Speech recognition#Open sourceBashkirs
  • Processes audiobooks, poetry, and ELAN field recordings
  • Converts materials into a format similar to Mozilla Common Voice
  • Prepares data for Whisper and Wav2Vec2
  • The collection includes 583 audio files of books and stories
Learn more

Bashkir Corpus

Language project#Language#Technology#Open sourceBashkirs
  • About 20.9 million tokens
  • GPL-3.0 license, accepts pull requests
  • Includes both public-domain works and copyrighted text excerpts
Learn more

Spoken Corpus of the Bashkir Language

Language project#Language#Technology#Open sourceBashkirs
  • Collects materials of live Bashkir speech
  • The data is intended for linguistic analysis
  • Published under the CC BY-SA 4.0 license
Learn more
Муэллиф