Multilingual
575 results
EuroSpeech
EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech…
nemotron-cc-translated
Helsinki-NLP/nemotron-cc-translated nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the…
fineweb-edu-translated
Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are…
yodas2_sidon
YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis…
common_corpus
Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising…
shofo-tiktok-general-small
Shofo TikTok General (Small) Overview Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive…
AutoMathText-V2
🚀 AutoMathText-V2: A 2.46 Trillion Token AI-Curated STEM Pretraining Dataset 🎉 AutoMathText-v2 has surpassed 1.5 million downloads!…
The Cross-lingual TRansfer Evaluation of Multilingual Encoders for Speech (XTREME-S) benchmark is a
The Cross-lingual TRansfer Evaluation of Multilingual Encoders for Speech (XTREME-S) benchmark is a benchmark designed to evaluate speech…
VoxSafeBench
VoxSafeBench Demopage: demopage/ Code: This dataset is uploaded as raw files (JSONL + audio), not parquet. Subset: Safety-tier1…
wikipedia-monthly
🚀 Wikipedia Monthly Last updated: March 14, 2026, 21:06 UTC This repository provides monthly, multilingual dumps of Wikipedia,…