Skip to content
Advertisement

Text Generation

217 results

MultimodalTabularText

PKU-SafeRLHF

Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are…

100K–1M·CC-BY-NC·JSON
MultimodalTabularText

the-stack-smol

Dataset Description A small subset (~0.1%) of the-stack dataset, each programming language has 10,000 random samples from the…

100K–1M·JSON
MultimodalTabularText

car-bench-dataset

CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It…

1M–10M·MIT·JSON
AudioMultimodalText

legco-speech

香港立法會會議語音數據集 本數據集係由香港立法會會議製成嘅大規模語音數據集。原始錄音總時長 22,196 個鐘,切分語音後總時長 20,471 個鐘。數據集分兩個子集,raw同segmented,分別為原始錄音同VAD識別切分後嘅語音。 數據集製作流程 先去香港特別行政區立法會…

1M–10M·CC0·Parquet
Audio

Marco_Longspeech

Marco-LongSpeech Dataset Marco-LongSpeech is a multi-task long speech understanding dataset containing 8 different speech understanding tasks…

10K–100K·Apache-2.0
Text

wmdp

Dataset Card for WMDP The Weapons of Mass Destruction Proxy (WMDP) benchmark is a dataset of multiple-choice questions…

1K–10K·MIT·Parquet
Text

Fineweb-Edu-Chinese-V2.1

Chinese Fineweb Edu Dataset V2.1 中文 English OpenCSG Community 👾github wechat Twitter 📖Technical Report The Chinese Fineweb Edu…

100M–1B·Apache-2.0·Parquet
Text

Argimi-Ardian-Finance-10k-text

The ArGiMI Ardian datasets : Text-only version The ArGiMi project is committed to open-source principles and data sharing.…

1M–10M·CC-BY·WebDataset
Text

prompts.chat

a.k.a. Awesome ChatGPT Prompts This is a Dataset Repository mirror of prompts.chat — a social platform for AI…

1K–10K·CC0·CSV
Text

SEC-EDGAR

Datamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data,…

1M–10M·Apache-2.0
Advertisement