CC0
85 results
EmbodiedEval
This repository contains the dataset of the paper EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents. Github repository: Project…
Mana-TTS
ManaTTS-Persian-Speech-Dataset ManaTTS is the largest publicly available single-speaker Persian corpus, comprising over 114 hours of high-quality…
KLING AI Generative Media Dataset
KLING AI Generative Media Dataset
crystallography-open-database
Crystallography Open Database (COD) — Full Snapshot A complete mirror of the Crystallography Open Database (COD) as a…
us-names-by-state
US Baby names The SSA dataset with baby names: Coniferest We use this dataset in the active anomaly…
Chronicling America – Historic American Newspapers 1770–1810
Chronicling America - Historic American Newspapers 1770–1810
adult-census-income
Adult Census Income Dataset The following was retrieved from UCI machine learning repository. This data was extracted from…
iris
Iris Species Dataset The Iris dataset was used in R.A. Fisher's classic 1936 paper, The Use of Multiple…
GlotCC-V1
Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is…
Kenyan Animal Behavior Recognition (KABR) Mini-Scene Raw Videos
Kenyan Animal Behavior Recognition (KABR) Mini-Scene Raw Videos
FakeParts_Legacy
FakeParts: A New Family of AI-Generated DeepFakes Abstract We introduce FakeParts, a new class of deepfakes characterized by…
legco-speech
香港立法會會議語音數據集 本數據集係由香港立法會會議製成嘅大規模語音數據集。原始錄音總時長 22,196 個鐘,切分語音後總時長 20,471 個鐘。數據集分兩個子集,raw同segmented,分別為原始錄音同VAD識別切分後嘅語音。 數據集製作流程 先去香港特別行政區立法會…
commonvoice22_sidon
CV22-Sidon Overview This dataset hosts a release of Mozilla Common Voice 22 restored with the Sidon speech restoration…
prompts.chat
a.k.a. Awesome ChatGPT Prompts This is a Dataset Repository mirror of prompts.chat — a social platform for AI…