Skip to content
Advertisement
MultimodalTextVideo

REVISOR-25k

REVISOR-25k A multi-task video understanding dataset for training video LLMs with reinforcement learning (GRPO). The dataset contains ~25k samples…

REVISOR-25k A multi-task video understanding dataset for training video LLMs with reinforcement learning (GRPO). The dataset contains ~25k samples spanning Video QA and Temporal Grounding tasks. Dataset Structure The dataset is organized into 4 subsets: Subset Task Samples Description video r1 Video QA 20,855 Multiple-choice video question answering time r1 Temporal Grounding 2,500 Locate time intervals in videos cg bench Temporal Grounding 1,167… See the full description on the dataset page:

Source: Hugging Face Hub (williamljz/REVISOR-25k). Metadata imported from the dataset’s Hub tags.

Advertisement