Skip to content
Advertisement

Datasets

1,940 results

AudioMultimodalText

synthetic-asr-zh

Synthetic ASR data — zh Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream…

10K–100K·JSON
AudioMultimodalText

malaysian-youtube

Malaysian Youtube Malaysian and Singaporean youtube channels, total up to 60k audio files with total 18.7k hours. URLs…

10K–100K·Parquet
AudioMultimodalText

MRSAudio

MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations Humans rely on multisensory integration to perceive…

100K–1M·CC-BY·CSV
Text

open-web-math

Keiran Paster , Marco Dos Santos , Zhangir Azerbayev, Jimmy Ba GitHub ArXiv PDF OpenWebMath is a dataset…

1M–10M·Parquet
Text

wmdp

Dataset Card for WMDP The Weapons of Mass Destruction Proxy (WMDP) benchmark is a dataset of multiple-choice questions…

1K–10K·MIT·Parquet
Text

Fineweb-Edu-Chinese-V2.1

Chinese Fineweb Edu Dataset V2.1 中文 English OpenCSG Community 👾github wechat Twitter 📖Technical Report The Chinese Fineweb Edu…

100M–1B·Apache-2.0·Parquet
AudioMultimodalText

PIAST

PIAST Dataset This repo is for downloading transcribed MIDI & and text data of the PIAST Dataset. The…

MIT
Text

WideSearch

WideSearch: Benchmarking Agentic Broad Info-Seeking Dataset Summary WideSearch is a benchmark designed to evaluate the capabilities of Large…

<1K·Custom / Research-only·JSON
MultimodalTextVideo

FLARE

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries 🤗 About This Benchmark This repository hosts the full…

100K–1M·CC-BY·JSON
Text

stsbenchmark-sts

STSBenchmark An MTEB dataset Massive Text Embedding Benchmark Semantic Textual Similarity Benchmark (STSbenchmark) dataset. Task category t2t Domains…

1K–10K·Custom / Research-only·JSON
Advertisement