Skip to content
Advertisement

VoxBox This dataset is a curated collection of bilingual speech corpora annotated clean transcriptions and rich metadata incluing age, gender, and emotion. Dataset Structure . ├── audios/ │ └── aishell-3/ Audio files (organised by sub-corpus) │ └── … └── metadata/ ├── aishell-3.jsonl ├── casia.jsonl ├── commonvoice cn.jsonl ├── … └── wenetspeech4tts.jsonl JSONL metadata files Each JSONL file corresponds to a… See the full description on the dataset page:

Source: Hugging Face Hub (SparkAudio/voxbox). Metadata imported from the dataset’s Hub tags.

Advertisement