LLaVA-Video-178K
Dataset Card for LLaVA-Video-178K Uses This dataset is used for the training of the LLaVA-Video model. We only…
339 results
Dataset Card for LLaVA-Video-178K Uses This dataset is used for the training of the LLaVA-Video model. We only…
GSA volc - GSA Embodied Perception Training Dataset Large-scale Grounding-Spatial-Affordance (GSA) training data for embodied perception Teacher…
MMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?) 🌐 Homepage 🤗 Dataset 🤗 Paper…
A multi-hazard, multi-sensor, and multi-task vision-language dataset for global-scale disaster assessment and response.
RealText-V2: A Large-Scale Multilingual Document Forgery Analysis Benchmark 💾 Dataset Description RealText-V2 is a large-scale multilingual document…
ShareGPT4Video Captions Dataset Card
ChartNet: A Million-Scale Multimodal Dataset for Chart Understanding 🌐 Homepage 📖 arXiv 📝 Changelog June 3, 2026 —…
LongVT-Source This repository contains the source video and image files for the LongVT project. Overview LongVT is an…
Document Haystack Dataset This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document…
GeoDrive-Bench A multi-country driving scene benchmark for evaluating vision-language models on culture- and region-specific traffic knowledge.…
MVBench Important Update 18/10/2024 Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded…
This dataset was converted from using the following script. import json import os from datasets import Dataset, DatasetDict,…
VisionArena-Battle: 30K Real-World Image Conversations with Pairwise Preference Votes 200k single and multi-turn chats between users and VLM's…
personalized visual instruction tuning