Text
2,175 results
iris
Iris Species Dataset The Iris dataset was used in R.A. Fisher's classic 1936 paper, The Use of Multiple…
SynData
SynData 中文说明 Demo If the video cannot be displayed in your environment, open it directly: assets/syndata-demo.mp4 1. Overview…
kinetics400
Kinetics-400 Video Dataset This dataset is derived from the Kinetics-400 dataset, which is released under the Creative Commons…
car-bench-dataset
CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It…
GlotCC-V1
Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is…
OpenAssistant Conversations Release 2
OpenAssistant Conversations Release 2
OpenAssistant Conversations
OpenAssistant Conversations
jigsaw-toxic-comment-classification-challenge
Dataset Description You are provided with a large number of Wikipedia comments which have been labeled by human…
plurel
PluRel A large collection of 2000 synthetic relational databases (seeds plurel-0 … plurel-1999) generated for pretraining relational/tabular…
Chronos datasets
Chronos datasets
FineWeb-Edu 100BT (Shuffled)
FineWeb-Edu 100BT (Shuffled)
Magicoder-OSS-Instruct-75K
This is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy…
EuroWeb-2512
EuroWeb-2512 EuroWeb is a dataset of collecting multilingual web data from various sources. It was processed with standard…
finemath
📐 FineMath What is it? 📐 FineMath consists of 34B tokens (FineMath-3+) and 54B tokens (FineMath-3+ with InfiMM-WebMath-3+)…