Skip to content
Advertisement

Dataset Card for PubLayNet Dataset Summary PubLayNet is a large document layout analysis dataset built by automatically matching XML representations and PDF content from more than one million PubMed Central Open Access articles. It contains more than 360,000 document images with COCO-style annotations for common layout elements such as text, title, list, table, and figure regions. Supported Tasks and Leaderboards The dataset supports document… See the full description on the dataset page:

Source: Hugging Face Hub (creative-graphic-design/PubLayNet). Metadata imported from the dataset’s Hub tags.

Advertisement