Skip to content
Advertisement
MultimodalTabularText

OpenSTT annotate by Balalaika

OpenSTT annotate by Balalaika

OpenSTT Annotated by Balalaika !IMPORTANT Official dataset for our INTERSPEECH 2026 paper “A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models” (arXiv:2507.13563). Part of the Balalaika Russian speech data-processing pipeline — code: If you use this resource, please cite it. A curated Russian speech dataset for advanced speech generative tasks. Overview OpenSTT… See the full description on the dataset page: balalaika.

Source: Hugging Face Hub (lab260/openstt_balalaika). Metadata imported from the dataset’s Hub tags.

Advertisement