Skip to content
Advertisement

Multimodal

2,368 results

ImageMultimodalVideo

FollowCam

FollowCam Dataset Description FollowCam is a benchmark for evaluating video generation models on viewpoint transformation and subject following…

10K–100K·MIT
MultimodalTextVideo

fMRI-Shape

fMRI-Shape Dataset: A Component of the fMRI-3D Dataset for MinD-3D++ This repository contains the fMRI-Shape dataset, a component…

1K–10K·Apache-2.0·Text (raw)
ImageMultimodalVideo

LoVoRA

LoVoRA Dataset: Text-guided and Mask-free Video Object Removal and Addition Authors: Zhihan Xiao, Lin Liu, Yixin Gao, Xiaopeng…

10K–100K·Apache-2.0
MultimodalTextVideo

VideoChat2

Video training data of LongVU downloaded from Video Please download the original videos from the provided links: BDD100K:…

100K–1M·MIT·JSON
MultimodalTextVideo

RoboReward

RoboReward Links: Paper · RoboRewardBench Leaderboard RoboReward is a dataset for training and evaluating general-purpose vision-language reward…

10K–100K·CC-BY
ImageMultimodalVideo

PanFlow

PanFlow Dataset The PanFlow dataset supports the research presented in the paper PanFlow: Decoupled Motion Control for Panoramic…

CC-BY
MultimodalTextVideo

phyground

PhyGround Project Page GitHub Paper PhyGround is a criteria-grounded benchmark for evaluating physical reasoning in video generation. The…

<1K·JSON
Advertisement