reward-bench-2
Code Leaderboard Results Paper RewardBench 2 Evaluation Dataset Card The RewardBench 2 evaluation dataset is the new version…
2,175 results
Code Leaderboard Results Paper RewardBench 2 Evaluation Dataset Card The RewardBench 2 evaluation dataset is the new version…
MVTamperBench Dataset Overview MVTamperBench is a robust benchmark designed to evaluate Vision-Language Models (VLMs) against adversarial video…
Languages: 简体中文 · English dojo forex kline — FX Daily Bars Overview Daily OHLC and amplitude for major…
Dataset Summary This dataset is a modified version of the Emilia corpus, converted into parquet format to facilitate…
AIDev: Studying AI Coding Agents on GitHub (The Rise of AI Teammates in Software Engineering 3.0) Papers: The…
MVBench Important Update 18/10/2024 Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded…
Industry models play a vital role in promoting the intelligent transformation and innovative development of enterprises. High-quality industry…
Core-S1RTC Contains a global coverage of Sentinel-1 (RTC) patches, each of size 1,068 x 1,068 pixels. Source Sensing…
starcoderdata-python-edu StarCoder Training Dataset Cleaned and Scored Dataset Details Dataset Description This dataset is a filtered version of…
Embeddings pre-training curated data This dataset is the English subset of lightonai/embeddings-pre-training, assembled to reproduce the English data…
SoC Builder RTL Dataset v1 (Experiment)
lex-au - Commonwealth Acts as AKN 3.0 XML
Dataset Card for "github-code-haskell-file" Rows: 339k Download Size: 806M This dataset is extracted from github-code-clean. Each row also…
Dataset Card for "distilabel-math-preference-dpo" More Information needed
💻 Stack-Edu Stack-Edu is a 125B token dataset of educational code filtered from The Stack v2, precisely the…
wikipedia-multilingual-synthetic-ir-query This dataset contains multilingual Wikipedia-derived synthetic query-document pairs for information…
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha…
Multi-Benchmark LLM Agent Traces A comprehensive dataset of OpenTelemetry traces capturing LLM inference behavior across multiple agent frameworks,…
Indoor Anomaly Detection & Path Obstruction Monitoring Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from…