Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Synthetic ASR data — hi Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream ASR…
Synthetic ASR data — hi Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream ASR finetuning. Total audio: 63.5 hr across short (5s) and long (30s) length buckets, each in clean and augmented variants. Loading from datasets import load dataset ds = load dataset(” /synthetic-asr-hi”, “short clean”) print(ds “train” 0 “audio” ) {“array”: np.ndarray, “sampling rate”: 16000, “path”: “…”}… See the full description on the dataset page:
Source: Hugging Face Hub (silvermango9927/synthetic-asr-hi). Metadata imported from the dataset’s Hub tags.