Skip to content
Advertisement
MultimodalTabularText

ESpeech datasets annotate by Balalaika

ESpeech datasets annotate by Balalaika

ESpeech datasets (w/o podcasts) Annotated by Balalaika !IMPORTANT Official dataset for our INTERSPEECH 2026 paper “A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models” (arXiv:2507.13563). Part of the Balalaika Russian speech data-processing pipeline — code: If you use this resource, please cite it. A curated Russian speech dataset for advanced speech generative tasks.… See the full description on the dataset page: balalaika.

Source: Hugging Face Hub (lab260/espeech_balalaika). Metadata imported from the dataset’s Hub tags.

Advertisement