Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Munch Hashed Index – Lightweight Audio Reference Dataset 📖 Overview Munch Hashed Index is a lightweight reference dataset that provides SHA-256 hashes for all audio files in the Munch Urdu TTS Dataset. Instead of storing 1.27 TB of raw audio, this index stores only metadata and cryptographic hashes, enabling: ✅ Fast duplicate detection across 4.17 million audio samples ✅ Efficient dataset exploration without downloading terabytes ✅ Quick metadata queries (voice… See the full description on the dataset page: data.
Source: Hugging Face Hub (humair025/hashed_data). Metadata imported from the dataset’s Hub tags.