Skip to content
Advertisement

Datasets

634 results

AudioMultimodalText

meow-10k

Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the…

10K–100K·Apache-2.0·JSON
AudioMultimodalVideo

rh20t_cfg1

rh20t cfg1 — RH20T → LeRobot v3 (flexiv) Unofficial LeRobot Dataset v3 reformatting of an RH20T config. RGB…

1K–10K·Custom / Research-only·Audio (folder)
AudioMultimodalVideo

JavisInst-Omni

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation HomePage Paper GitHub TL;DR We introduce JavisGPT, a…

10K–100K·Apache-2.0
AudioMultimodalVideo

TAVGBench_1m

Installation Download this repo to a local folder, and unzip these .zip files under the TAVGBench 1m/data/. Then,…

1M–10M·MIT
MultimodalTextVideo

FLARE

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries 🤗 About This Benchmark This repository hosts the full…

100K–1M·CC-BY·JSON
MultimodalTabularText

gaming-500-hours

Gaming Dataset (gaming-1) — 494.7 Hours Native PC/console gameplay screen-recordings, organized by game. Each workflow is one play…

<1K·JSON
Advertisement