Qirimtatar Embedding Crawlers
A set of crawlers for building a Crimean Tatar text corpus and training language models.
About this project
- Collects texts from open web sources and digital archives
- Uses Tesseract for text recognition in PDFs and images
- Publishes the extracted documents separately from the source code
- Processes the Qirim Junior archives and the texts of Shamil Aladin