L2D
TL;DR of L2D, the world's largest self-driving dataset! Read more about L2D on the official Huggingface blog: LeRobot…
169 results
TL;DR of L2D, the world's largest self-driving dataset! Read more about L2D on the official Huggingface blog: LeRobot…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "robot type": "AgiBot A2D", "total…
S2ORC Full (Semantic Scholar Open Research Corpus)
The GitHub Code clean dataset in a more filtered version of codeparrot/github-code dataset, it consists of 115M code…
I also seperately provide just the prompts in prompts.json keys are the image id, and the values are…
NOTE: I have recaptioned all images here This dataset is the entire 21K ImageNet dataset with about 13…
Chitralekha Dataset Details Dataset Version Some of the fonts do not have proper letters/rendering of different telugu letter…
Dataset Card for CommonCatalog CC-BY This dataset is a large collection of high-resolution Creative Common images (composed of…
Innovator-VL-Instruct-46M Paper Code 🤗🤗 The data is being uploaded continuously Introduction To further enhance the model’s ability to…
Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage…
For more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository. IMPORTANT NOTE!!! The handwritten subset of this dataset…
Dataset Card for Conceptual Captions 12M (CC12M) Dataset Summary Conceptual 12M (CC12M) is a dataset with 12 million…
This dataset has been created by Stability AI and LAION. This dataset contains 12 million 1024x1024 images of…
MapPool - Bubbling up an extremely large corpus of maps for AI MapPool is a dataset of 75…
LLaVA-OneVision-1.5 Instruction Data Paper Code 📌 Introduction This dataset, LLaVA-OneVision-1.5-Instruct, was collected and integrated during the…
Fine Vision FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B…