Skip to content
Advertisement

10K–100K

748 results

MultimodalTabularText

uci-har-federated

UCI Human Activity Recognition (HAR) Dataset Dataset Description The UCI Human Activity Recognition dataset is a widely-used benchmark…

10K–100K·MIT·Parquet
MultimodalTabularText

pg-en

Overview Property Value Source Project Gutenberg (English catalog) Snapshot 2026-07-02-18-47-04 Total files 50871 Total Tokens (BPE) ~7.14 billion…

10K–100K·MIT·Parquet
Tabular

pack_toothbrush_Nov19-advantages

Advantage Values for villekuosmanen/pack toothbrush Nov19 Pre-computed advantage values for offline RL training. Source Dataset: villekuosmanen/pack…

10K–100K·Apache-2.0·Parquet
MultimodalTabularText

GSM-Symbolic

GSM-Symbolic This project accompanies the research paper, GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language…

10K–100K·CC-BY-NC-ND·JSON
ImageMultimodalTabular

zendo-synthetic-data

Zendo Synthetic Visual Reasoning Dataset Synthetic Zendo-style scenes with associated rules and per-scene tensor representations. Each scene either…

10K–100K·CC-BY·Parquet
MultimodalTabularText

OceanTACO

Dataset Card: OceanTACO Dataset Summary This dataset is a multi-source collection of global ocean sea surface measurements, integrating…

10K–100K·CC-BY
MultimodalTabularText

liars-bench

Liars' Bench Liars' Bench is a benchmark for evaluating lie-detectors for language models. For details, see paper here.…

10K–100K·Custom / Research-only·Parquet
MultimodalTabularText

Code-Contests-Plus

CodeContests+: A Competitive Programming Dataset with High-Quality Test Cases Introduction CodeContests+ is a competitive programming problem dataset…

10K–100K·CC-BY·Parquet
MultimodalTabularText

ragbench

RAGBench Dataset Overview RAGBEnch is a large-scale RAG benchmark dataset of 100k RAG examples. It covers five unique…

10K–100K·CC-BY·Parquet
Advertisement