Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Khmer ASR Cultural Dataset 727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided. Language: Khmer (khm). Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata. Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page:
Source: Hugging Face Hub (DDD-Cambodia/khmer-speech-dataset). Metadata imported from the dataset’s Hub tags.