Skip to content
Advertisement

10M–100M

169 results

MultimodalTabularText

SylReg

Dataset for SylReg This repository contains the datasets and alignments associated with the paper Speaker-Disentangled Chunk-Wise Regression for…

10M–100M·CC-BY-NC-SA·Parquet
MultimodalTabularText

TencentGR-1M

TencentGR-1M Dataset Paper Project Page Code TAAC2025 Preliminary Round Dataset (2025年腾讯广告算法大赛初赛数据集) TencentGR-1M Dataset is a large-scale,…

10M–100M·CC-BY·Parquet
MultimodalTabularText

ScaleEdit-12M

ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework       📌 Overview The largest open-source…

10M–100M·CC-BY-NC-SA·Parquet
ImageMultimodalTabular

Open-Qwen2VL-Data

Introduction This repository contains the data for Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources.…

10M–100M·MIT·Parquet
MultimodalTabularText

MegaMath-Web-Pro-Max

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling The Curation of MegaMath-Web-Pro-Max Step 1: Uniformly and randomly sample…

10M–100M·ODC-BY·Parquet
MultimodalTabularText

Classifiers-Data

Text Quality Classifier Training Dataset This dataset is specifically designed for training text quality assessment classifiers, containing annotated…

10M–100M·Parquet
MultimodalTabularText

Scientific-Summaries

Scientific Summaries 22 million LLM-generated structured summaries of scientific papers, enriched with OpenAlex scholarly metadata. Each paper has…

10M–100M·CC-BY·Parquet
MultimodalTabularText

finemath

📐 FineMath What is it? 📐 FineMath consists of 34B tokens (FineMath-3+) and 54B tokens (FineMath-3+ with InfiMM-WebMath-3+)…

10M–100M·ODC-BY·Parquet
Advertisement