English
1,233 results
osm-polygon-selection
osm-polygon-selection dataset A curated set of OpenStreetMap polygons from 310 geographic units — sovereign countries plus sub-country regions…
pile-deduped-pythia-preshuffled
This dataset contains the fully prepared data, which has been tokenized and pre-shuffled, used to train the Pythia…
medicine-tasks
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the evaluation datasets for our paper…
Hate Speech and Offensive Language
Hate Speech and Offensive Language
reward-bench-2
Code Leaderboard Results Paper RewardBench 2 Evaluation Dataset Card The RewardBench 2 evaluation dataset is the new version…
MVTamperBench
MVTamperBench Dataset Overview MVTamperBench is a robust benchmark designed to evaluate Vision-Language Models (VLMs) against adversarial video…
MVBench
MVBench Important Update 18/10/2024 Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded…
embeddings-pre-training-curated
Embeddings pre-training curated data This dataset is the English subset of lightonai/embeddings-pre-training, assembled to reproduce the English data…
lex-au – Commonwealth Acts as AKN 3.0 XML
lex-au - Commonwealth Acts as AKN 3.0 XML
distilabel-math-preference-dpo
Dataset Card for "distilabel-math-preference-dpo" More Information needed
agent-llm-traces
Multi-Benchmark LLM Agent Traces A comprehensive dataset of OpenTelemetry traces capturing LLM inference behavior across multiple agent frameworks,…
KodCode-V1
🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset…
India Index & Options 1-minute OHLC
India Index & Options 1-minute OHLC
PKU-SafeRLHF-10K
Paper You can find more information in our paper. Dataset Paper:
binance-btcusdt
BTCUSDT Perpetual Futures — 5-Minute Feature Dataset Complete historical dataset for Binance BTCUSDT USDT-Margined Perpetual Futures, covering…