Skip to content
Advertisement
ImageMultimodalText

NayanaOCR Corpus 2025

NayanaOCR Corpus 2025

🪷 NayanaOCR Corpus 2025 A 1M-page, 22-language fully-parallel synthetic OCR + VQA corpus for document-centric vision-language models — every page rendered in every language. NayanaOCR Corpus 2025 is one of the largest open-source multilingual, multi-task document datasets for training and evaluating OCR, layout detection, and visual question answering (VQA) in low-resource and underrepresented languages. The headline property: it’s a true parallel corpus. The same ~45,700 source… See the full description on the dataset page: Corpus 2025.

Source: Hugging Face Hub (Cognitive-Lab/NayanaOCR_Corpus_2025). Metadata imported from the dataset’s Hub tags.

Advertisement