Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
VoxBox This dataset is a curated collection of bilingual speech corpora annotated clean transcriptions and rich metadata incluing age, gender, and emotion. Dataset Structure . ├── audios/ │ └── aishell-3/ Audio files (organised by sub-corpus) │ └── … └── metadata/ ├── aishell-3.jsonl ├── casia.jsonl ├── commonvoice cn.jsonl ├── … └── wenetspeech4tts.jsonl JSONL metadata files Each JSONL file corresponds to a… See the full description on the dataset page:
Source: Hugging Face Hub (SparkAudio/voxbox). Metadata imported from the dataset’s Hub tags.