Skip to content
Advertisement

CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models Dataset Description CrossVid is a large-scale, multi-task dataset designed to advance cross-video understanding capabilities in vision-language models. The dataset encompasses 10 diverse task types that require models to reason across multiple videos, understand temporal dynamics, spatial relationships, and complex narrative structures. Unlike existing benchmarks… See the full description on the dataset page:

Source: Hugging Face Hub (Chuntianli/CrossVid). Metadata imported from the dataset’s Hub tags.

Advertisement