Skip to content
Advertisement

Video

706 results

AudioMultimodalVideo

JavisInst-Omni

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation HomePage Paper GitHub TL;DR We introduce JavisGPT, a…

10K–100K·Apache-2.0
AudioMultimodalVideo

TAVGBench_1m

Installation Download this repo to a local folder, and unzip these .zip files under the TAVGBench 1m/data/. Then,…

1M–10M·MIT
MultimodalTextVideo

FLARE

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries 🤗 About This Benchmark This repository hosts the full…

100K–1M·CC-BY·JSON
MultimodalTabularText

gaming-500-hours

Gaming Dataset (gaming-1) — 494.7 Hours Native PC/console gameplay screen-recordings, organized by game. Each workflow is one play…

<1K·JSON
MultimodalTextVideo

InsViE

InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction Citation If you find this work helpful, please consider…

1M–10M·CC-BY·CSV
MultimodalTextVideo

ProLongVid_data

Dataset Card for ProLongVid-data Uses This dataset is used for the training of the ProLongVid model. We only…

1M–10M·Apache-2.0·JSON
MultimodalTabularText

agibot_alpha_v30

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "robot type": "AgiBot A2D", "total…

10M–100M·Apache-2.0·Parquet
ImageMultimodalVideo

GOKU-2M

Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing GOKU-2M is a large-scale, unified instruction-based…

1M–10M·CC-BY-NC
Advertisement