Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Dataset Card for LongVALE Uses This dataset is designed for training and evaluating models on omni-modal (vision-audio-language-event) fine-grained…
Dataset Card for LongVALE Uses This dataset is designed for training and evaluating models on omni-modal (vision-audio-language-event) fine-grained video understanding tasks. It is intended for academic research and educational purposes only. For data generated using third-party models (e.g., Gemini-1.5-Pro, GPT-4o, Qwen-Audio), users must comply with the respective model providers’ usage policies. Data Sources LongVALE comprises 8,411 long videos (549 hours) with 105… See the full description on the dataset page:
Source: Hugging Face Hub (ttgeng233/LongVALE). Metadata imported from the dataset’s Hub tags.