Skip to content
Advertisement

Question Answering

190 results

ImageMultimodalText

MMMU_Pro

MMMU-Pro (A More Robust Multi-discipline Multimodal Understanding Benchmark) 🌐 Homepage 🏆 Leaderboard 🤗 Dataset 🤗 Paper 📖 arXiv…

1K–10K·Apache-2.0·Parquet
AudioMultimodalText

C3T

C3T: Cross-modal Capabilities Conservation Test Dataset Description C3T (Cross-modal Capabilities Conservation Test) is a benchmark for assessing the…

100K–1M·CC-BY
MultimodalTabularText

NautData

NautData Paper Project Page Code NautData is a large-scale underwater instruction-following dataset containing 1.45 million image-text pairs. It…

1M–10M·Apache-2.0·JSON
ImageMultimodalText

S-Chain

S-Chain: Structured Visual Chain-of-Thought for Medicine ⭐ If you find this project helpful, please consider giving it a…

100K–1M·Apache-2.0·Parquet
3D / Point CloudImageMultimodal

BenchCAD

BenchCAD Three-config dataset for CAD evaluation: edit-bench — held-out CAD edit benchmark. code gen — 17,900 synthetic CadQuery…

10K–100K·CC-BY·Parquet
MultimodalTabularText

medicine-tasks

Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the evaluation datasets for our paper…

1K–10K·JSON
MultimodalTabularText

reward-bench-2

Code Leaderboard Results Paper RewardBench 2 Evaluation Dataset Card The RewardBench 2 evaluation dataset is the new version…

1K–10K·ODC-BY·Parquet
Advertisement