Skip to content
Advertisement
ImageMultimodalText

VLAA-Thinking

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models 🌐 Project Page • 📄 Arxiv • 💻 Code 🤗 VLAA-Thinker…

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models 🌐 Project Page • 📄 Arxiv • 💻 Code 🤗 VLAA-Thinker Family • 🤔 VLAA-Thinking Dataset 🤗 VLAA-Thinker-Qwen2.5-3B • 🤗 VLAA-Thinker-Qwen2.5-7B Both VLAA-Thinker-Qwen2.5-3B and VLAA-Thinker-Qwen2.5-7Bachieve SOTA performance on OpenCompass Multimodal Reasoning Leaderboard as of April 7th, 2025. Contents Quick Start 🚀… See the full description on the dataset page:

Source: Hugging Face Hub (UCSC-VLAA/VLAA-Thinking). Metadata imported from the dataset’s Hub tags.

Advertisement