Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
TAGARELA: A Portuguese Speech Dataset From Podcasts TAGARELA is a large-scale Portuguese speech dataset built from podcast audio and curated for speech technology research, especially Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The dataset contains more than 8,972 hours of Portuguese speech derived from the Cem Mil Podcasts collection. It includes Brazilian Portuguese and European Portuguese speech, processed through a pipeline involving audio standardization… See the full description on the dataset page:
Source: Hugging Face Hub (freds0/TAGARELA). Metadata imported from the dataset’s Hub tags.