Skip to content
Advertisement

Datasets

3,216 results

Text

Argimi-Ardian-Finance-10k-text

The ArGiMI Ardian datasets : Text-only version The ArGiMi project is committed to open-source principles and data sharing.…

1M–10M·CC-BY·WebDataset
Text

BLiMP

BLiMP

10K–100K·CC-BY·Parquet
Text

hh-rlhf

Dataset Card for HH-RLHF Dataset Summary This repository provides access to two different kinds of data: Human preference…

100K–1M·MIT·JSON
Text

prompts.chat

a.k.a. Awesome ChatGPT Prompts This is a Dataset Repository mirror of prompts.chat — a social platform for AI…

1K–10K·CC0·CSV
Text

browsecomp-plus

BrowseComp-Plus BrowseComp-Plus is a new benchmark for Deep-Research system, isolating the effect of the retriever and the LLM…

<1K·MIT·Parquet
Text

SEC-EDGAR

Datamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data,…

1M–10M·Apache-2.0
AudioMultimodalText

X-Voice-Dataset-Train

X-Voice Training Dataset Overview The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech…

10M–100M·Custom / Research-only·WebDataset
MultimodalTabularText

gaming-500-hours

Gaming Dataset (gaming-1) — 494.7 Hours Native PC/console gameplay screen-recordings, organized by game. Each workflow is one play…

<1K·JSON
Text

bbh

BIG-bench Hard dataset homepage: @article{suzgun2022challenging, title={Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them},…

1K–10K·Parquet
Text

nfcorpus

NFCorpus An MTEB dataset Massive Text Embedding Benchmark NFCorpus: A Full-Text Learning to Rank Dataset for Medical Information…

100K–1M·JSON
Text

TweetEval

TweetEval

100K–1M·Custom / Research-only·Parquet
Advertisement