Skip to content
Advertisement

Video Text To Text

54 results

Text

VideoKR-Train

VideoKR-Train 📄 ArXiv  |  💻 Code  |  🤗 Collection About This repository contains the VideoKR training data presented…

100K–1M·Apache-2.0·JSON
Image

UniVA-Bench

UniVA-Bench Paper: UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist Project Page: Code: UniVA-Bench is a…

1K–10K·Images (folder)
3D / Point CloudMultimodalVideo

DSI-Bench

DSI-Bench: A Benchmark for Dynamic Spatial Intelligence Paper Project Page Code Abstract Reasoning about dynamic spatial relationships is…

1K–10K·CC-BY
Video

omega-multimodal

OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research Introduction The OMEGA Labs Bittensor Subnet Dataset is a…

MIT
Video

AVUTBenchmark

Audio-centric Video Understanding Benchmark (AVUT) This dataset is presented in the paper Audio-centric Video Understanding Benchmark without Text…

1K–10K
Video

OmniMMI

OmniMMI Paper: OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts Code Dataset Description we introduce OmniMMI,…

1K–10K·MIT
ImageMultimodalTabular

MentalBlackboard

MentalBlackboard Benchmark This repository contains the MentalBlackboard benchmark with multiple tasks: - Prediction - Planning Dataset Sources

1K–10K·CC-BY·Parquet
MultimodalTextVideo

VideoChat2

Video training data of LongVU downloaded from Video Please download the original videos from the provided links: BDD100K:…

100K–1M·MIT·JSON
MultimodalTextVideo

finevideo

FineVideo FineVideo Description Dataset Explorer Revisions Dataset Distribution How to download and use FineVideo Using datasets Using huggingface…

10K–100K·Other·Parquet
MultimodalTextVideo

CSL-News

Summary This is the dataset proposed in our paper "Uni-Sign: Toward Unified Sign Language Understanding at Scale". CSL-News…

100K–1M·CC-BY-NC·JSON
MultimodalTextVideo

VideoChat3-LV116k

VideoChat3-LV116K VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video…

1K–10K·Apache-2.0·JSON
Video

ParaVT-Source

ParaVT-Source Source media archives for the ParaVT training corpus. Pair this repository with the annotations in ParaVT/ParaVT-Parquet. Overview…

100K–1M·Apache-2.0
Video

EgoLife

Data cleaning, stay tuned! Please refer to first for general info. Checkout the paper EgoLife ( for more…

10K–100K·MIT
Advertisement