ProLongVid_data
Dataset Card for ProLongVid-data Uses This dataset is used for the training of the ProLongVid model. We only…
329 results
Dataset Card for ProLongVid-data Uses This dataset is used for the training of the ProLongVid model. We only…
Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in…
🚀 Wikipedia Monthly Last updated: March 14, 2026, 21:06 UTC This repository provides monthly, multilingual dumps of Wikipedia,…
Multilingual Speech Commands Dataset (15 Languages, Augmented)
GLUE (General Language Understanding Evaluation benchmark)
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing GOKU-2M is a large-scale, unified instruction-based…
Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM.
FLUX-Reason-6M FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative…
GSA volc - GSA Embodied Perception Training Dataset Large-scale Grounding-Spatial-Affordance (GSA) training data for embodied perception Teacher…
Dataset Card for Industry Documents Library (IDL) Dataset Summary Industry Documents Library (IDL) is a document dataset filtered…
ChartNet: A Million-Scale Multimodal Dataset for Chart Understanding 🌐 Homepage 📖 arXiv 📝 Changelog June 3, 2026 —…
Dataset Card for Conceptual Captions (CC3M) Dataset Summary Conceptual Captions is a dataset consisting of ~3.3M images annotated…
Dataset Card for CommonCatalog CC-BY-SA This dataset is a large collection of high-resolution Creative Common images (composed of…