Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
TTS-Italian A high-quality Italian speech dataset for text-to-speech and automatic speech recognition. Data Sources Derived from LibriVox Italian — volunteer-read Italian public domain audiobooks hosted on archive.org. Books: 23 Italian-language audiobooks (Dante, Pirandello, Verga, De Amicis, Collodi, Pascoli, etc.) License: Public Domain Processing: Standardized to 24kHz mono, WhisperX transcription (large-v3) with word-level alignment, segmented at word boundaries… See the full description on the dataset page:
Source: Hugging Face Hub (datadriven-company/TTS-Italian). Metadata imported from the dataset’s Hub tags.