Skip to content
Advertisement

<1K

439 results

Audio

MIT_environmental_impulse_responses

MIT Environmental Impulse Response Dataset The audio recordings in this dataset are originally created by the Computational Audition…

<1K·Custom / Research-only·Audio (folder)
Audio

french-tts-conversational-dataset

French Conversational TTS Dataset Dataset Description This dataset contains high-fidelity French text-to-speech audio clips generated using Mistral's…

<1K·Apache-2.0·Audio (folder)
Audio

GTSinger

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks Yu Zhang , Changhao…

<1K·Audio (folder)
Audio

AIR-Bench-Dataset

AIR-Bench Arxiv: is the AIR-Bench dataset download page.AIR-Bench encompasses two dimensions: foundation and chat benchmarks. The former consists…

<1K·CC-BY-NC·Audio (folder)
Text

WideSearch

WideSearch: Benchmarking Agentic Broad Info-Seeking Dataset Summary WideSearch is a benchmark designed to evaluate the capabilities of Large…

<1K·Custom / Research-only·JSON
Text

browsecomp-plus

BrowseComp-Plus BrowseComp-Plus is a new benchmark for Deep-Research system, isolating the effect of the retriever and the LLM…

<1K·MIT·Parquet
MultimodalTabularText

gaming-500-hours

Gaming Dataset (gaming-1) — 494.7 Hours Native PC/console gameplay screen-recordings, organized by game. Each workflow is one play…

<1K·JSON
Text

aime25

AIME 25 American Invitational Mathematics Examination (AIME) 2025 Citation If you use the AIME25 dataset in your research,…

<1K·Apache-2.0·JSON
Text

aime_2024

Dataset card for AIME 2024 This dataset consists of 30 problems from the 2024 AIME I and AIME…

<1K·Parquet
Text

recipes

🦛 Chonkie Recipes 🍳 Chonkie loves to cook up a storm in the kitchen This repository contains all…

<1K·Apache-2.0
Text

LongBench-v2

LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: 💻 Github Repo: 📚…

<1K·Apache-2.0·JSON
Text

SWE-bench_Pro

Dataset Summary SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks.…

<1K·Parquet
Text

SWE-bench_Verified

Dataset Summary SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been…

<1K·Parquet
AudioMultimodalText

coraal

Corpus of Regional African American Language (CORAAL) Dataset link: Hugging Face Hub preparation scripts: coraal Citation Kendall, Tyler…

<1K·CC-BY-NC-SA·Parquet
Text

requests

Open LLM Leaderboard Requests This repository contains the request files of models that have been submitted to the…

<1K·Apache-2.0·JSON
Text

gdpval

Dataset for GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks. Paper Blog Site 220 real-world knowledge…

<1K·Parquet
Text

SWE-bench_Lite

Dataset Summary SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues…

<1K·Parquet
Advertisement