Skip to content
Advertisement

Open

2,922 results

Text

prompts.chat

a.k.a. Awesome ChatGPT Prompts This is a Dataset Repository mirror of prompts.chat — a social platform for AI…

1K–10K·CC0·CSV
Text

browsecomp-plus

BrowseComp-Plus BrowseComp-Plus is a new benchmark for Deep-Research system, isolating the effect of the retriever and the LLM…

<1K·MIT·Parquet
Text

SEC-EDGAR

Datamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data,…

1M–10M·Apache-2.0
AudioMultimodalText

X-Voice-Dataset-Train

X-Voice Training Dataset Overview The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech…

10M–100M·Custom / Research-only·WebDataset
MultimodalTabularText

gaming-500-hours

Gaming Dataset (gaming-1) — 494.7 Hours Native PC/console gameplay screen-recordings, organized by game. Each workflow is one play…

<1K·JSON
Text

bbh

BIG-bench Hard dataset homepage: @article{suzgun2022challenging, title={Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them},…

1K–10K·Parquet
Text

nfcorpus

NFCorpus An MTEB dataset Massive Text Embedding Benchmark NFCorpus: A Full-Text Learning to Rank Dataset for Medical Information…

100K–1M·JSON
Text

TweetEval

TweetEval

100K–1M·Custom / Research-only·Parquet
Text

Danbooruwildcards

This is a set of wildcards for danbooru tags. Artist:Prompts for random artist styles, covering approximately 0.6M different…

10M–100M·Custom / Research-only·Text (raw)
Text

sts12-sts

STS12 An MTEB dataset Massive Text Embedding Benchmark SemEval-2012 Task 6. Task category t2t Domains Encyclopaedic, News, Written…

1K–10K·Custom / Research-only·JSON
Advertisement