Skip to content
Advertisement

Question Answering

190 results

MultimodalTabularText

ArabicMMLU

Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha…

10K–100K·CC-BY-NC·CSV
MultimodalTabularText

car-bench-dataset

CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It…

1M–10M·MIT·JSON
ImageMultimodalTabular

MentalBlackboard

MentalBlackboard Benchmark This repository contains the MentalBlackboard benchmark with multiple tasks: - Prediction - Planning Dataset Sources

1K–10K·CC-BY·Parquet
MultimodalTabularText

full-modality-data

Full Modality Dataset Statistics Video Statistics Total Videos: 28,472 Total Duration: 1422.33 hours Average Duration: 179.84 seconds Median…

1M–10M·MIT·Parquet
Video

MVLU

MLVU: Multi-task Long Video Understanding Benchmark This repo contains the annotation data and evaluation code for the paper…

CC-BY-NC-SA
Video

VEU-Bench

Video Editing Understanding(VEU) Benchmark 🖥 Project Page Widely shared videos on the internet are often edited. Recently, although…

Apache-2.0
AudioMultimodalText

AudioMarathon

AudioMarathon AudioMarathon is a long-context audio benchmark for evaluating multimodal LLMs on speech, music, environmental audio, and meetings.…

<1K·Custom / Research-only·CSV
Advertisement