Open
2,922 results
Cosmos3-DROID
DROID: Distributed Robot Interaction Dataset Dataset Summary DROID (Distributed Robot Interaction Dataset) is a large-scale "in-the-wild" robot…
bc_z_lerobot
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v2.0", "robot type": "google robot", "total…
open-swara
Open Swara The largest open-source humanized voice library -- 4,065 voice samples across 44 languages, free forever under…
Audio2Tool — Spoken Tool-Calling Benchmark
Audio2Tool — Spoken Tool-Calling Benchmark
Suno AI Music Dataset — Multi-Genre Curated
Suno AI Music Dataset — Multi-Genre Curated
libritts
Dataset Card for LibriTTS LibriTTS is a multi-speaker English corpus of approximately 585 hours of read English speech…
SciTS: Scientific Time Series Understanding and Generation with LLMs
SciTS: Scientific Time Series Understanding and Generation with LLMs
FalAR
FalAR FalAR is a large-scale, speaker-annotated European Portuguese speech corpus built from recordings of parliamentary sessions of the…
esc50
The dataset is available under the terms of the Creative Commons Attribution Non-Commercial license. K. J. Piczak. ESC:…
russia_voices
Russian voices for train AI. ♀ Male and ♂ Female Male voices - 497 pcs. Female voices -…
EuroSpeech-24kHz
EuroSpeech 24 kHz Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary…
Multilingual MFA-Aligned Speech Dataset
Multilingual MFA-Aligned Speech Dataset