Skip to content
Advertisement

Tabular

1,117 results

MultimodalTabularText

TencentGR-1M

TencentGR-1M Dataset Paper Project Page Code TAAC2025 Preliminary Round Dataset (2025年腾讯广告算法大赛初赛数据集) TencentGR-1M Dataset is a large-scale,…

10M–100M·CC-BY·Parquet
MultimodalTabularText

KMMLU

KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging…

100K–1M·Other·CSV
MultimodalTabularText

smollm-chunked

FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora This repository contains part of the FAISS indices…

100M–1B·ODC-BY·Arrow
Tabular

Nemotron-ClimbMix

ClimbMix Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠…

100M–1B·CC-BY-NC·JSON
MultimodalTabularText

GPT-NL_Public_Corpus

Dataset Card GPT-NL Public Corpus The GPT-NL Public Corpus is the largest permissively licensed Dutch-language resource available for…

100M–1B·CC-BY·Parquet
MultimodalTabularText

ragbench

RAGBench Dataset Overview RAGBEnch is a large-scale RAG benchmark dataset of 100k RAG examples. It covers five unique…

10K–100K·CC-BY·Parquet
ImageMultimodalTabular

Core-S2L2A

Core-S2L2A Contains a global coverage of Sentinel-2 (Level 2A) patches, each of size 1,068 x 1,068 pixels. Source…

1M–10M·CC-BY-SA·Parquet
MultimodalTabularText

fake_news

TODO: Add YAML tags here. Copy-paste the tags obtained with the online tagging app: annotations creators: - no-annotation…

10K–100K·Parquet
ImageMultimodalTabular

filtered-wit

Filtered WIT, an Image-Text Dataset. A reliable Dataset to run Image-Text models. You can find WIT, Wikipedia Image…

1M–10M·Parquet
Advertisement