Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
🚁 UAVReason: VQA / Caption / Generation Annotations A UAV-native benchmark for aerial visual reasoning, captioning, temporal understanding, and cross-modal generation 📌 News Paper: Can Vision-Language Models Think from the Sky? Unifying UAV Reasoning and Generation arXiv: arXiv:2604.05377 Depth data: jarvissun/UAVReason depth This dataset is released as part of UAVReason, introduced in the paper above.Please cite the paper if you use this dataset.… See the full description on the dataset page: vqa.
Source: Hugging Face Hub (jarvissun/UAVReason_vqa). Metadata imported from the dataset’s Hub tags.