Skip to content
Advertisement
Text

LLaVA-CoT

LLaVA-CoT

Dataset Card for LLaVA-CoT The LLaVA-CoT-100k dataset is introduced in the paper LLaVA-CoT: Let Vision Language Models Reason Step-by-Step. This dataset is designed to enable Vision-Language Models (VLMs) to perform autonomous multistage reasoning, integrating samples from various visual question-answering sources with structured reasoning annotations. It aims to address the challenges VLMs face in systematic and structured reasoning for complex visual question-answering tasks.… See the full description on the dataset page:

Source: Hugging Face Hub (Xkev/LLaVA-CoT-100k). Metadata imported from the dataset’s Hub tags.

Advertisement