open-asr-leaderboard
ESB Test Sets: Parquet & Sorted This dataset takes the open-asr-leaderboard/datasets-test-only data and sorts each split by audio…
1,940 results
ESB Test Sets: Parquet & Sorted This dataset takes the open-asr-leaderboard/datasets-test-only data and sorts each split by audio…
FreeSound.org LAION-640k Dataset
AVQA JSONL (Audio Multiple-Choice QA)
NatureLM-audio-training
Synthetic ASR data — zh Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream…
Malaysian Youtube Malaysian and Singaporean youtube channels, total up to 60k audio files with total 18.7k hours. URLs…
MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations Humans rely on multisensory integration to perceive…
Keiran Paster , Marco Dos Santos , Zhangir Azerbayev, Jimmy Ba GitHub ArXiv PDF OpenWebMath is a dataset…
Chinese Fineweb Edu Dataset V2.1 中文 English OpenCSG Community 👾github wechat Twitter 📖Technical Report The Chinese Fineweb Edu…
PIAST Dataset This repo is for downloading transcribed MIDI & and text data of the PIAST Dataset. The…
WideSearch: Benchmarking Agentic Broad Info-Seeking Dataset Summary WideSearch is a benchmark designed to evaluate the capabilities of Large…
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries 🤗 About This Benchmark This repository hosts the full…
DCLM-baseline Note: this is an identical copy of where all the files have been mapped to a parquet…
LOTSA Data The Large-scale Open Time Series Archive (LOTSA) is a collection of open time series datasets for…
STSBenchmark An MTEB dataset Massive Text Embedding Benchmark Semantic Textual Similarity Benchmark (STSbenchmark) dataset. Task category t2t Domains…