Skip to content
Advertisement

Fine Vision FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B answer tokens, designed for training state-of-the-art open Vision-Language-Models. More detail can be found in the blog post: The version in this repository concatenated all the configs in the original dataset and then shuffled them. This is done to facilitate streaming the data directly from the hub! Load… See the full description on the dataset page:

Source: Hugging Face Hub (HuggingFaceM4/FineVisionMax). Metadata imported from the dataset’s Hub tags.

Advertisement