Skip to content
Advertisement

JSON

248 results

MultimodalTextVideo

VideoChat2

Video training data of LongVU downloaded from Video Please download the original videos from the provided links: BDD100K:…

100K–1M·MIT·JSON
MultimodalTextVideo

phyground

PhyGround Project Page GitHub Paper PhyGround is a criteria-grounded benchmark for evaluating physical reasoning in video generation. The…

<1K·JSON
MultimodalTextVideo

Vript

🎬 Vript: Refine Video Captioning into Video Scripting Github Repo We construct a fine-grained video-text dataset with 12K…

100K–1M·JSON
MultimodalTextVideo

VideoChat3-LV116k

VideoChat3-LV116K VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video…

1K–10K·Apache-2.0·JSON
MultimodalTextVideo

CSL-News

Summary This is the dataset proposed in our paper "Uni-Sign: Toward Unified Sign Language Understanding at Scale". CSL-News…

100K–1M·CC-BY-NC·JSON
MultimodalTextVideo

EyePCR

A dataset for "EyePCR: A Comprehensive Benchmark for Fine-Grained Perception, Knowledge Comprehension and Clinical Reasoning in Ophthalmic Surgery"…

10K–100K·CC-BY·JSON
AudioMultimodalText

meow-10k

Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the…

10K–100K·Apache-2.0·JSON
AudioMultimodalText

synthetic-asr-hi

Synthetic ASR data — hi Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream…

10K–100K·JSON
AudioMultimodalText

synthetic-asr-zh

Synthetic ASR data — zh Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream…

10K–100K·JSON
Text

WideSearch

WideSearch: Benchmarking Agentic Broad Info-Seeking Dataset Summary WideSearch is a benchmark designed to evaluate the capabilities of Large…

<1K·Custom / Research-only·JSON
MultimodalTextVideo

FLARE

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries 🤗 About This Benchmark This repository hosts the full…

100K–1M·CC-BY·JSON
Text

stsbenchmark-sts

STSBenchmark An MTEB dataset Massive Text Embedding Benchmark Semantic Textual Similarity Benchmark (STSbenchmark) dataset. Task category t2t Domains…

1K–10K·Custom / Research-only·JSON
Text

hh-rlhf

Dataset Card for HH-RLHF Dataset Summary This repository provides access to two different kinds of data: Human preference…

100K–1M·MIT·JSON
Advertisement