Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
NaturalVoices Restored (48 kHz, Sidon + UTMOS-filtered)
NaturalVoices — Sidon-Restored, UTMOS-Filtered (48 kHz) High-quality English speech derived from NaturalVoices VC 870h (JHU SmileLab), restored with Sidon v0.1 and kept only where restoration measurably improved perceptual quality (UTMOS gate). Each clip ships with rich per-utterance metadata (transcript, speaker age/gender, speaking rate, emotion, and quality scores) so it is ready for TTS / voice-cloning / ASR / paralinguistic research. 494,903 clips · 736.7 hours (clips ≥… See the full description on the dataset page: 737h 48k.
Source: Hugging Face Hub (PleasedPenguin/naturalvoice_737h_48k). Metadata imported from the dataset’s Hub tags.