Skip to content
Advertisement

Highlights 40,000 high-quality video-based spatial reasoning training samples Built primarily from five large-scale real-world indoor 3D scene datasets (ScanNet, ScanNet++, S3DIS, ARKitScenes, Aria Digital Twin), plus a small ProcTHOR simulated subset Covers diverse spatial skills: geometric perception, spatial relations, counting, and temporal / appearance-order reasoning over video Filtered with rejection sampling using Qwen3-VL-8B-Instruct to reduce ambiguous or low-quality… See the full description on the dataset page:

Source: Hugging Face Hub (hyin-ustc/CIRCLE-40K). Metadata imported from the dataset’s Hub tags.

Advertisement