nuScenes
1,000 driving scenes with 360° camera, LiDAR and radar.
Dataset Card for NaviTrace NaviTrace is a Visual Question Answering benchmark for evaluating how well vision-language models (VLMs) understand embodiment-specific navigation in real-world scenes. Each sample presents a first-person image of an outdoor environment paired with a natural language navigation instruction. The task is to predict a 2D navigation trace — a sequence of waypoints in image space — that a given embodiment type would follow to complete the instruction. The… See the full description on the dataset page:
Source: Hugging Face Hub (Voxel51/NaviTrace). Metadata imported from the dataset’s Hub tags.