Skip to content
Advertisement

CC0

85 results

MultimodalTabularText

Mana-TTS

ManaTTS-Persian-Speech-Dataset ManaTTS is the largest publicly available single-speaker Persian corpus, comprising over 114 hours of high-quality…

10K–100K·CC0·Parquet
MultimodalTabularText

iris

Iris Species Dataset The Iris dataset was used in R.A. Fisher's classic 1936 paper, The Use of Multiple…

<1K·CC0·CSV
MultimodalTabularText

GlotCC-V1

Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is…

>1B·CC0
Video

FakeParts_Legacy

FakeParts: A New Family of AI-Generated DeepFakes Abstract We introduce FakeParts, a new class of deepfakes characterized by…

<1K·CC0
AudioMultimodalText

legco-speech

香港立法會會議語音數據集 本數據集係由香港立法會會議製成嘅大規模語音數據集。原始錄音總時長 22,196 個鐘,切分語音後總時長 20,471 個鐘。數據集分兩個子集,raw同segmented,分別為原始錄音同VAD識別切分後嘅語音。 數據集製作流程 先去香港特別行政區立法會…

1M–10M·CC0·Parquet
AudioMultimodalText

commonvoice22_sidon

CV22-Sidon Overview This dataset hosts a release of Mozilla Common Voice 22 restored with the Sidon speech restoration…

10M–100M·CC0·WebDataset
Text

prompts.chat

a.k.a. Awesome ChatGPT Prompts This is a Dataset Repository mirror of prompts.chat — a social platform for AI…

1K–10K·CC0·CSV
Advertisement