Question Answering
190 results
Daily-Omni (QA repackaged for lmms-eval)
Daily-Omni (QA repackaged for lmms-eval)
AVQA JSONL (Audio Multiple-Choice QA)
AVQA JSONL (Audio Multiple-Choice QA)
AIR-Bench-Dataset
AIR-Bench Arxiv: is the AIR-Bench dataset download page.AIR-Bench encompasses two dimensions: foundation and chat benchmarks. The former consists…
prompts.chat
a.k.a. Awesome ChatGPT Prompts This is a Dataset Repository mirror of prompts.chat — a social platform for AI…
browsecomp-plus
BrowseComp-Plus BrowseComp-Plus is a new benchmark for Deep-Research system, isolating the effect of the retriever and the LLM…
databricks-dolly-15k
Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several…
LongBench-v2
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: 💻 Github Repo: 📚…
AutoMathText-V2
🚀 AutoMathText-V2: A 2.46 Trillion Token AI-Curated STEM Pretraining Dataset 🎉 AutoMathText-v2 has surpassed 1.5 million downloads!…