Multimodal
2,368 results
Code-Contests-Plus
CodeContests+: A Competitive Programming Dataset with High-Quality Test Cases Introduction CodeContests+ is a competitive programming problem dataset…
challenge2026_dataset
Challenge 2026 Dataset This dataset provides training/evaluation data for the 2026 Global Humanoid Robot Challenge Simulation Competition. Data…
synthetic-bathroom-dataset-for-robotic-perception
Synthetic Bathroom Dataset for Robotic Perception Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl.…
dojo_fin_indicators
Languages: 简体中文 · English dojo fin indicators — Financial Metrics Overview Multi-period financial statement derivatives per symbol: income…
Consumer Price Indices — Africa (FAOSTAT)
Consumer Price Indices — Africa (FAOSTAT)
Wan2.2-Syn-121x704x1280_32k
FastVideo Synthetic Wan2.2 720P dataset FastVideo Team Paper Github Project Page Abstract Scaling video diffusion transformers (DiTs) is…
TencentGR-1M
TencentGR-1M Dataset Paper Project Page Code TAAC2025 Preliminary Round Dataset (2025年腾讯广告算法大赛初赛数据集) TencentGR-1M Dataset is a large-scale,…
CLUE: Chinese Language Understanding Evaluation benchmark
CLUE: Chinese Language Understanding Evaluation benchmark
bias_in_bios
Bias in Bios Bias in Bios was created by (De-Artega et al., 2019) and published under the MIT…
KMMLU
KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging…
smollm-chunked
FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora This repository contains part of the FAISS indices…
GPT-NL_Public_Corpus
Dataset Card GPT-NL Public Corpus The GPT-NL Public Corpus is the largest permissively licensed Dutch-language resource available for…
ragbench
RAGBench Dataset Overview RAGBEnch is a large-scale RAG benchmark dataset of 100k RAG examples. It covers five unique…
Core-S2L2A
Core-S2L2A Contains a global coverage of Sentinel-2 (Level 2A) patches, each of size 1,068 x 1,068 pixels. Source…