Skip to content
Advertisement

10K–100K

748 results

Image

WorldMemArena

WorldMemArena WorldMemArena is a large-scale multimodal memory benchmark designed to evaluate how well AI systems retain, update, and…

10K–100K·CC-BY-NC·Images (folder)
AudioMultimodalText

Irodori-Ja-Spk2-10k

SynDataLab/Irodori-Ja-Spk2-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker…

10K–100K·CC-BY-NC-SA·Parquet
AudioMultimodalText

irodori-refs-10k-v2

Irodori TTS Reference Voices v2 (10K) 10,000 reference voices generated with Aratako/Irodori-TTS-500M-v2-VoiceDesign (no ref=True) using a richer…

10K–100K·Apache-2.0·Parquet
AudioMultimodalText

IndicTTS_Bengali

Bengali Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Bengali…

10K–100K·CC-BY·Parquet
AudioMultimodalText

Habibi

A systematic and standardized benchmark for the multi-dialect Arabic zero-shot TTS task Paper: "Habibi: Laying the Open-Source Foundation…

10K–100K·Apache-2.0·Parquet
Audio

Urdu-Munch-Lina

Urdu-Munch-Lina Processed version of zuhri025/Urdu-Munch with LinaCodec encoding. Dataset Structure This dataset contains 2 batches of audio data…

10K–100K·MIT
Audio

AISHELL-3

AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It…

10K–100K·Apache-2.0·Audio (folder)
AudioMultimodalText

bashkort_tts_dataset

Bashkort TTS Dataset The largest open dataset for speech synthesis in the Bashkir language — featuring multi-speaker recordings…

10K–100K·CC-BY·Parquet
AudioMultimodalText

wuwa-voice-EN

wuwa-voice-EN wuwa voice EN is a dataset of voice line from Wuthering Waves Attribute Value Language English Total…

10K–100K·Parquet
Advertisement