Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
AVQA JSONL (Audio Multiple-Choice QA)
Summary 摘要 This dataset is collected from the AVQA training subset (train qa.json). We converted the data to the R1-AQA format, where each line in the text file represents a JSON object with specific keys. The AVQA training set originally consists of approximately 40k samples. However, we use only about 38k samples because some data sources have become invalid (e.g. link failure, or less than 10 seconds). Given that there is no quick link to the audio mentioned in the above two… See the full description on the dataset page:
Source: Hugging Face Hub (Joysw909/AVQA). Metadata imported from the dataset’s Hub tags.