Skip to content
Advertisement
Image

OmniDocBench

OmniDocBench English 简体中文 OmniDocBench is an evaluation dataset for diverse document parsing in real-world scenarios, with the following…

OmniDocBench English 简体中文 OmniDocBench is an evaluation dataset for diverse document parsing in real-world scenarios, with the following characteristics: Diverse Document Types: The evaluation set contains 1651 PDF pages, covering 10 document types, 5 layout types and 5 language types. Coverage includes academic literature, research and financial reports, newspapers, textbooks, exam papers, magazines, handwritten notes, historical documents, and more. Rich Annotations:… See the full description on the dataset page:

Source: Hugging Face Hub (opendatalab/OmniDocBench). Metadata imported from the dataset’s Hub tags.

Advertisement