Bashkir Language Audio2Text
Tools for collecting and preparing a Bashkir speech corpus for speech-recognition models.
About this project
- Processes audiobooks, poetry, and ELAN field recordings
- Converts materials into a format similar to Mozilla Common Voice
- Prepares data for Whisper and Wav2Vec2
- The collection includes 583 audio files of books and stories
