AudioMultimodalText
Daily-Omni (QA repackaged for lmms-eval)
Daily-Omni (QA repackaged for lmms-eval)
1K–10K·CC-BY-NC-SA·Parquet
54 results
Daily-Omni (QA repackaged for lmms-eval)
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries 🤗 About This Benchmark This repository hosts the full…
Dataset Card for LLaVA-Video-178K Uses This dataset is used for the training of the LLaVA-Video model. We only…
LongVT-Source This repository contains the source video and image files for the LongVT project. Overview LongVT is an…
LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family, covering large-scale video and spatial reasoning corpora used…