Skip to content
Advertisement
Text

rtm-sgt-ocr-v1

Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic Data from…

Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper “Large Synthetic Data from the ar𝜒iv for OCR Post Correction of Historic Scientific Articles”. Synthetic ground truth (SGT) sentences have been mined from the ar𝜒iv Bulk Downloads source documents, and Optical Character Recognition (OCR) sentences have been generated with the Tesseract OCR engine on the PDF pages generated from compiled source documents.… See the full description on the dataset page:

Source: Hugging Face Hub (ReadingTimeMachine/rtm-sgt-ocr-v1). Metadata imported from the dataset’s Hub tags.

Advertisement