Skip to content
Advertisement

Open

2,922 results

Tabular

Nemotron-ClimbMix

ClimbMix Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠…

100M–1B·CC-BY-NC·JSON
MultimodalTabularText

GPT-NL_Public_Corpus

Dataset Card GPT-NL Public Corpus The GPT-NL Public Corpus is the largest permissively licensed Dutch-language resource available for…

100M–1B·CC-BY·Parquet
MultimodalTabularText

ragbench

RAGBench Dataset Overview RAGBEnch is a large-scale RAG benchmark dataset of 100k RAG examples. It covers five unique…

10K–100K·CC-BY·Parquet
ImageMultimodalTabular

Core-S2L2A

Core-S2L2A Contains a global coverage of Sentinel-2 (Level 2A) patches, each of size 1,068 x 1,068 pixels. Source…

1M–10M·CC-BY-SA·Parquet
MultimodalTabularText

fake_news

TODO: Add YAML tags here. Copy-paste the tags obtained with the online tagging app: annotations creators: - no-annotation…

10K–100K·Parquet
ImageMultimodalTabular

filtered-wit

Filtered WIT, an Image-Text Dataset. A reliable Dataset to run Image-Text models. You can find WIT, Wikipedia Image…

1M–10M·Parquet
MultimodalTabularText

SurvHTE-Bench

SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis Paper: ICLR 2026 — SurvHTE-Bench: A Benchmark…

1M–10M·CC-BY·Parquet
MultimodalTabularText

ScaleEdit-12M

ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework       📌 Overview The largest open-source…

10M–100M·CC-BY-NC-SA·Parquet
MultimodalTabularText

AlgoTune

Website Paper Code How good are language models at coming up with new algorithms? To try to answer…

<1K·MIT·JSON
MultimodalTabularText

toxic-chat

Update 01/31/2024 We update the OpenAI Moderation API results for ToxicChat (0124) based on their updated moderation model…

10K–100K·CC-BY-NC·CSV
Advertisement