Resources

OCR resources

Resources and data for developing a language-aware OCR document error profiler and PoCoTo tools.

Manuals, lexica, OCR test data for PoCoTo and the profiler

GitHub

15 stars
6 watching
2 forks
Language: Lex
last commit: about 5 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
cisocrgroup/pocotoA Java-based tool for correcting errors in OCR'd historical documents40
ocropus/hocr-toolsTools for manipulating and analyzing multi-lingual OCR results by representing them in a standard HTML format373
lascivaroma/lexicalDevelops OCR models and ground truth data for a Latin lexical resource1
aslez/concorA software package for concordance analysis in R9
lex4all/lex4allSoftware tool to generate pronunciation lexicons for low-resource languages using speech recognition and machine learning algorithms.21
cpitclaudel/alectryonA tool for processing Coq and Lean 4 code embedded in text documents237
ploc-org/cnplA collection of annual reports on domestic programming languages in China.234
talyssonoc/commonregexrubyExtracts common information from text strings in various formats79
chreul/ocr_testdata_earlyprintedbooksProvides test data and models for training Optical Character Recognition (OCR) systems on historical printed books.10
peterc/whatlanguageLanguage detection library using Bloom filters for speed and memory efficiency.685
osrf/osrf_testing_tools_cppProvides common testing tools and utilities for C++ projects33
oncybersec/oscp-enumeration-cheat-sheetA cheat sheet for conducting enumeration during penetration testing and security assessments102
openseg-group/openseg.pytorchProvides a PyTorch implementation of several computer vision tasks including object detection, segmentation and parsing.1,191
mittagessen/krakenAn OCR system optimized for historical and non-Latin scripts, providing layout analysis, character recognition, and support for various formats.757