Open
2,922 results
Banana-Merged: Multi-page VQA with Hard Negatives
Banana-Merged: Multi-page VQA with Hard Negatives
ChartGalaxy
ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation 🤗 Dataset 🖥️ Code 📄 Paper 🔥 News 2026.02…
RoboBench
RoboBench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain 📋 Overview RoboBench is a…
Vietnamese Medicinal Herb VQA
Vietnamese Medicinal Herb VQA
Cauldron-JA
Dataset Card for The Cauldron-JA Dataset description The Cauldron-JA is a Vision Language Model dataset that translates 'The…
MMMU_Pro
MMMU-Pro (A More Robust Multi-discipline Multimodal Understanding Benchmark) 🌐 Homepage 🏆 Leaderboard 🤗 Dataset 🤗 Paper 📖 arXiv…
Irodori-Ja-Spk2-10k
SynDataLab/Irodori-Ja-Spk2-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker…
French_game_voice
French Game Voice Dataset Dataset of 100k+ cleaned audio samples of French video game voices with transcriptions. Features…
irodori-refs-10k-v2
Irodori TTS Reference Voices v2 (10K) 10,000 reference voices generated with Aratako/Irodori-TTS-500M-v2-VoiceDesign (no ref=True) using a richer…
bengali-tts-missing-v1
Bengali TTS — Missing Rows This dataset contains the rows from rwd51/bengali-tts-combined that are not present in the…
IndicTTS_Bengali
Bengali Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Bengali…
Bambara-ASR-All Audio Dataset
Bambara-ASR-All Audio Dataset