in-the-wild-jailbreak-prompts
In-The-Wild Jailbreak Prompts on LLMs This is the official repository for the ACM CCS 2024 paper "Do Anything…
1,117 results
In-The-Wild Jailbreak Prompts on LLMs This is the official repository for the ACM CCS 2024 paper "Do Anything…
Game of 24 Mathematical Puzzle Dataset
Vietnamese Án lệ + Bản án Corpus
Changelog NEW Changes March 11th 2026 Added new split: arxiv papers, sourced from the Hugging Face /api/papers endpoint…
MMLU-ProX MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to…
Dataset Card: OceanTACO Dataset Summary This dataset is a multi-source collection of global ocean sea surface measurements, integrating…
Dataset Card for Demo1 Dataset Summary This is a demo dataset. It consists in two files data/train.csv and…
Foursquare OS Places is now a gated dataset on Hugging Face. Read more about why we are making…
Betty Dota 2 — Decision Context Dataset Overview 9,385 professional Dota 2 matches parsed from replay files (.dem)…
NoeFlandre/osm-polygon-wikidata-only OSM polygons tagged with a wikidata= reference, enriched with Wikipedia and Wikivoyage text across all available…
Dataset Description A small subset of the-stack dataset, with 87 programming languages, each has 10,000 random samples from…
CTU relational datasets (redelex) Relational databases from the CTU Prague Relational Learning Repository (a.k.a. the CTU relational repository),…
Dataset Card for "pickapic v2" please pay attention - the URLs will be temporariliy unavailabe - but you…
VLM TSR Test Data (Part 1/4)
WhestBench 2026: ARC White-Box Estimation Challenge
Dataset Card for "QM9" QM9 dataset from Ruddigkeit et al., 2012; Ramakrishnan et al., 2014. Original data downloaded…
cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance…
Crystallography Open Database (COD) — Full Snapshot A complete mirror of the Crystallography Open Database (COD) as a…
Liars' Bench Liars' Bench is a benchmark for evaluating lie-detectors for language models. For details, see paper here.…
Dataset Card for CodeForces-CoTs Dataset description CodeForces-CoTs is a large-scale dataset for training reasoning models on competitive…