Open
2,922 results
Marco_Longspeech
Marco-LongSpeech Dataset Marco-LongSpeech is a multi-task long speech understanding dataset containing 8 different speech understanding tasks…
open-web-math
Keiran Paster , Marco Dos Santos , Zhangir Azerbayev, Jimmy Ba GitHub ArXiv PDF OpenWebMath is a dataset…
Fineweb-Edu-Chinese-V2.1
Chinese Fineweb Edu Dataset V2.1 中文 English OpenCSG Community 👾github wechat Twitter 📖Technical Report The Chinese Fineweb Edu…
PIAST
PIAST Dataset This repo is for downloading transcribed MIDI & and text data of the PIAST Dataset. The…
WideSearch
WideSearch: Benchmarking Agentic Broad Info-Seeking Dataset Summary WideSearch is a benchmark designed to evaluate the capabilities of Large…
FLARE
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries 🤗 About This Benchmark This repository hosts the full…
dclm-baseline-1.0-parquet
DCLM-baseline Note: this is an identical copy of where all the files have been mapped to a parquet…
lotsa_data
LOTSA Data The Large-scale Open Time Series Archive (LOTSA) is a collection of open time series datasets for…
stsbenchmark-sts
STSBenchmark An MTEB dataset Massive Text Embedding Benchmark Semantic Textual Similarity Benchmark (STSbenchmark) dataset. Task category t2t Domains…
Mathematics Aptitude Test of Heuristics (MATH)
Mathematics Aptitude Test of Heuristics (MATH)
music-arena-dataset
Music Arena Dataset This is the official dataset from Music Arena, an open platform for evaluating text-to-music (TTM)…
Argimi-Ardian-Finance-10k-text
The ArGiMI Ardian datasets : Text-only version The ArGiMi project is committed to open-source principles and data sharing.…