Wan2.2-Syn-121x704x1280_32k
FastVideo Synthetic Wan2.2 720P dataset FastVideo Team Paper Github Project Page Abstract Scaling video diffusion transformers (DiTs) is…
1,117 results
FastVideo Synthetic Wan2.2 720P dataset FastVideo Team Paper Github Project Page Abstract Scaling video diffusion transformers (DiTs) is…
TencentGR-1M Dataset Paper Project Page Code TAAC2025 Preliminary Round Dataset (2025年腾讯广告算法大赛初赛数据集) TencentGR-1M Dataset is a large-scale,…
CLUE: Chinese Language Understanding Evaluation benchmark
Bias in Bios Bias in Bios was created by (De-Artega et al., 2019) and published under the MIT…
KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging…
FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora This repository contains part of the FAISS indices…
ClimbMix Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠…
Dataset Card GPT-NL Public Corpus The GPT-NL Public Corpus is the largest permissively licensed Dutch-language resource available for…
RAGBench Dataset Overview RAGBEnch is a large-scale RAG benchmark dataset of 100k RAG examples. It covers five unique…
Core-S2L2A Contains a global coverage of Sentinel-2 (Level 2A) patches, each of size 1,068 x 1,068 pixels. Source…
TODO: Add YAML tags here. Copy-paste the tags obtained with the online tagging app: annotations creators: - no-annotation…
Filtered WIT, an Image-Text Dataset. A reliable Dataset to run Image-Text models. You can find WIT, Wikipedia Image…
This dataset contains the subset of ArXiv papers with the "cs.LG" tag to indicate the paper is about…