Skip to content
Advertisement

Tabular

1,117 results

MultimodalTabularText

PKU-SafeRLHF

Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are…

100K–1M·CC-BY-NC·JSON
MultimodalTabularText

Scientific-Summaries

Scientific Summaries 22 million LLM-generated structured summaries of scientific papers, enriched with OpenAlex scholarly metadata. Each paper has…

10M–100M·CC-BY·Parquet
MultimodalTabularText

the-stack-smol

Dataset Description A small subset (~0.1%) of the-stack dataset, each programming language has 10,000 random samples from the…

100K–1M·JSON
MultimodalTabularText

bio-lens

🌿 iNaturalist Bronze Dataset (Research-Grade, Deduplicated) Description This dataset contains research-grade observations from iNaturalist, processed…

>1B·Other
MultimodalTabularText

aime_2025

AIME 2025 This dataset contains 30 problems from the 2025 AIME tests, including: AIME I: 15 problems AIME…

<1K·Parquet
MultimodalTabularText

iris

Iris Species Dataset The Iris dataset was used in R.A. Fisher's classic 1936 paper, The Use of Multiple…

<1K·CC0·CSV
3D / Point CloudMultimodalTabular

SynData

SynData 中文说明 Demo If the video cannot be displayed in your environment, open it directly: assets/syndata-demo.mp4 1. Overview…

100K–1M·CC-BY·Parquet
MultimodalTabularText

kinetics400

Kinetics-400 Video Dataset This dataset is derived from the Kinetics-400 dataset, which is released under the Creative Commons…

10K–100K·Parquet
MultimodalTabularText

car-bench-dataset

CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It…

1M–10M·MIT·JSON
MultimodalTabularText

GlotCC-V1

Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is…

>1B·CC0
Advertisement