ICDAR2019’s Scanned Receipts OCR and Information Extraction (SROIE)
ICDAR2019's Scanned Receipts OCR and Information Extraction (SROIE)
1,501 results
ICDAR2019's Scanned Receipts OCR and Information Extraction (SROIE)
M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
GenEvolve Data and Bench This repository contains the open-source data release for GenEvolve: Config Directory Records Images Purpose…
DCVLM-Baseline (200B tokens)
QCalEval Dataset Description The dataset contains scientific plots from quantum computing calibration experiments, paired with vision-language…
LVOmniBench (QA repackaged for lmms-eval)
This dataset is composed by 150k examples of English Visual Instruction Data from LLaVA. 150k examples of English…
MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal…
PixMo-Docs We now recommend using CoSyn-400k and CoSyn-point over these datasets. They are improved versions with more images…
PhyX: Does Your Model Have the "Wits" for Physical Reasoning? Dataset for the paper "PhyX: Does Your Model…
Dataset Card for OpenDocVQA This is a training and evaluation QA data file for VDocRAG, a new RAG…
Dataset Card for VideoMarathon VideoMarathon is a large-scale long video instruction-following dataset with a total duration of approximately…
synthvision-validated-qwen-by-kimi Qwen 3.5 annotations validated by Kimi K2.5 (93.1% pass rate) Records: 55,359 About Cross-validated subset from…