awesome-ocr

OCR toolkit

A curated list of OCR engines, tools, and formats for extracting text from images and documents.

Links to awesome OCR projects

GitHub

3k stars
128 watching
352 forks
last commit: about 2 years ago
Linked from 4 awesome lists


Awesome OCR / Software / OCR engines

tesseract63,142almost 2 years agoThe definitive Open Source OCR engine
EasyOCR24,876about 2 years agoOCR engine built on PyTorch by JaidedAI,
ocropus3,426over 5 years agoOCR engine based on LSTM,
ocropus 0.417almost 15 years agoOlder v0.4 state of Ocropus, with tesseract 2.04 and iulib, C++
kraken757almost 2 years agoOcropus fork with sane defaults
gocrOCR engine under the GNU Public License led by Joerg Schulenburg
OcradThe GNU OCR
ocular256over 2 years agoMachine-learning OCR for historic documents
SwiftOCR4,623almost 6 years agofast and simple OCR library written in Swift
attention-ocr1,079almost 3 years agoOCR engine using visual attention mechanisms
RWTH-OCRThe RWTH Aachen University Optical Character Recognition System
simple-ocr-opencv525over 2 years agoand its - A simple pythonic OCR engine using opencv and numpy
Calamari1,056almost 2 years agoOCR Engine based on OCRopy and Kraken
doctr4,011almost 2 years agoA seamless & high-performing OCR library powered by Deep Learning

Awesome OCR / Software / Older and possibly abandoned OCR engines

Clara OCROpen source OCR in C
CuneiformCuneiForm OCR was developed by Cognitive Technologies
Eyean experimental Java OCR (image-to-text) application
kognitionAn omnifont OCR software for KDE
OCRchieModular Optical Character Recognition Software
ocreo.c.r. easy
xplabA GTK 2 tool for pattern matching
hebOCR5over 10 years agoHebrew character recognition library (previously named hocr, see )

Awesome OCR / Software / OCR file formats

abby2hocr.xslt XSLT script
ocr-conversion-scripts72over 3 years ago
hocr-tools373about 2 years agoTools for doing various useful things with hOCR files,
hocr-spec74about 2 years agohOCR 1.2 specification
ocr-transform182almost 2 years agoCLI tool to convert between hOCR and ALTO,
hocr-parser13about 11 years agohOCR Specification Python Parser
hOCRTools6about 8 years agohOCR to ALTO conversion XSLT
ALTO XML Schema52about 2 years agoXML Schema and development of the ALTO XML format
ALTO XML Documentation39about 8 years agoDocumentation and use cases for ALTO
alto-tools40almost 3 years agoVarious tools to work with ALTO files, Python
AbbyyToAlto9over 15 years agoPHP script converting from Abbyy 6 to ALTO XML
TEI-OCR1over 10 years agoTEI customization for OCR generated layout and content information
TEI SIG on LibrariesBest Practices for TEI in Libraries
GDZMETS/TEI-based GDZ document format
PAGE-XML Schema66about 5 years agoXML schema of the PAGE XML format along with documentation and examples
omni:us Pages Format (OPF)XML schema very similar to PAGE XML that has some additional features
py-pagexml13almost 2 years agoPython library for handling PAGE XML and OPF files

Awesome OCR / Software / OCR CLI

OCRmyPDF14,363almost 2 years agoOCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Pdf2PdfOCR279over 2 years agoA tool to OCR a PDF (or supported images) and add a text "layer" (a "pdf sandwich") in the original file making it a searchable PDF. GUI included. Tesseract and cuneiform supported
OcrocisProject manager interface for Ocropy, see also
tesseract-recognize44over 2 years agoTesseract-based tool that outputs result in Page XML format ( )

Awesome OCR / Software / OCR GUI

moz-hocr-editor10over 11 years agoFirefox Addon for editing hOCR files
qt-box-editor173almost 2 years agoQT4 editor of tesseract-ocr box files
ocr-gt-tools48almost 6 years agoClient-Server application for editing OCR ground truth
Paperwork2,431over 8 years agoUsing scanners and OCR to grep paper documents the easy way
Paperless7,864over 5 years agoScan, index, and archive all of your paper documents
gImageReader1,653almost 2 years agogImageReader is a simple Gtk/Qt front-end to tesseract-ocr
VietOCRA Java/.NET GUI frontend for Tesseract OCR engine, including a graphical Tesseract editor
PoCoTo40almost 4 years agoFast interactive batch corrections of complete OCR error series in OCR'ed historical documents
OCRFeederGTK graphical user interface that allows the users to correct characters or bounding boxes, ODT export and more
PRImA PAGE Viewer35over 3 years agoJava based viewer for PAGE XML files (layout + text content). Also supports ALTO XML, FineReader XML, and HOCR
LAREX181almost 2 years agoA semi-automatic open-source tool for Layout Analysis and Region EXtraction on early printed books
archiscribe17over 8 years agoWeb application for transcribing OCR ground truth from Archive.org. Deployed instance available at , results are available in
nw-page-editor30over 2 years agoSimple app for visual editing of Page XML files. Provides desktop and versions

Awesome OCR / Software / OCR Preprocessing

NoiseRemove.java in MathOCR168almost 4 years agoJava implementation of Adaptive degraded document image binarization by B. Gatos , I. Pratikakis, S.J. Perantonis
binarize.c in ZBar2,503over 2 years agoC implementations of two binarization algorithms, based on Sauvola
typeface-corpus7almost 12 years agoA repository for typefaces to train Tesseract and OCRopus for natural history collections and digital humanities
binarizewolfjolion30about 9 years agoComparison of binarization algorithms
crop_morphology.py in oldnyc289almost 2 years agoCropping a page to just the text block
Whiteboard Picture CleanerShell one-liner/script to clean up and beautify photos of whiteboards
textcleanerFred's ImageMagick script - Processes a scanned document of text to clean the text background
localcontrastFast O(1) local contrast optimization

Awesome OCR / Software / OCR as a Service

Open OCR1,346about 3 years agoRun Tesseract in Docker containers
tesseract-web-service135over 3 years agoAn implementation of RESTful web service for tesseract-OCR using tornado
docker-ocropy9almost 9 years agoA Docker container for running the
ABBYY Cloud OCR SDK Code samples504over 3 years agoCode samples for using the proprietary commercial ABBYY OCR API
nidaba86almost 9 years agoAn expandable and scalable OCR pipeline
gamera39about 4 years agoA meta-framework for building document processing applications, e.g. OCR
ocr-tools7over 5 years agoProject to provide CLI and web service interfaces to common OCR engines
ocrad-docker2about 10 years agoRun the OCR engine in a docker container
kraken-docker5almost 9 years agoRun the OCR engine in a docker container
KonfuzioFree Online OCR up to 2.000 pages per month and OCR API by [@atraining], see (code is not open)
ocr.spaceFree Online OCR and OCR API by based on Tesseract (code is not open)
OCR4all244over 2 years agoProvides OCR services through web applications. Included Projects: , , and

Awesome OCR / Software / OCR evaluation

ISRI OCR Evaluation Toolswith a

Awesome OCR / Software / OCR evaluation / ISRI OCR Evaluation Tools

isri-ocr-evaluation-tools57over 5 years agofurther development by (2015, 2016)
ancientgreekocr-evaluation-tools22over 8 years agofurther development by (2013, 2014)

Awesome OCR / Software / OCR evaluation

ocrevalUAtion67about 4 years agoCross-format evaluation, CLI and GUI
ngram-ocr-eval1over 12 years agoBrute and simple OCR evaluation using ngrams
quack22almost 4 years agoQuality-Assurance-tool for scans with corresponding ALTO-files

Awesome OCR / Software / OCR libraries by programming language

tesseract-ocr13over 4 years agoA Crystal wrapper for tesseract-ocr
tesseract_ocr55over 4 years agoElixir library wrapping the tesseract executable
gosseract2,751about 2 years agoGolang OCR library, wrapping Tesseract-ocr
Tess4J1,619almost 2 years agoJava Native Access bindings to Tesseract
tess-two3,761over 4 years agoTools for compiling Tesseract on Android and Java API
tesseract for .net2,308over 2 years agoA .Net wrapper for tesseract-ocr
TTesseractOCR4145about 3 years agoObject Pascal binding for tesseract-ocr 4.x
Tesseract OCR for PHP2,897almost 3 years agoTesseract PHP bindings
pytesseract5,919almost 2 years agoA Python wrapper for Google Tesseract
pyocr930over 8 years agoA Python wrapper for Tesseract and Cuneiform
ocrodjvu46almost 4 years agoA library and standalone tool for doing OCR on DjVu documents, wrapping Cuneiform, gocr, ocrad, ocropus and tesseract
tesserocr2,026almost 2 years agoA Python wrapper for the tesseract-ocr API
ocracy37over 11 years agopure javascript lstm rnn implementation based on ocropus
gocr.js98over 12 years agoJavascript port (emscripten) of gocr
ocrad.js3,494about 6 years agoJavascript port (emscripten) of ocrad
tesseract.js35,553almost 2 years agoJavascript port (emscripten) of Tesseract
node-tesseract-ocr308about 3 years agoA simple wrapper for the Tesseract OCR package
node-tesseract-native51almost 8 years agoC++ module for node providing OCR with tesseract and leptonica
rtesseract838almost 3 years agoRuby library wrapping the tesseract and imagemagick executables
ruby-tesseract629about 9 years agoNative Tesseract bindings for Ruby MRI and JRuby
ocr_space70over 7 years agoAPI wrapper for free ocr service ocr.space. Includes CLI
tesseract.rs148over 2 years agoRust bindings for tesseract OCR
leptessProductive and safe Rust bindings/wrappers for tesseract and leptonica
tesseract245almost 2 years agoR bindings for tesseract OCR
Tesseract OCR iOS4,220over 5 years agoSwift and Objective-C wrapper for Tesseract OCR
SwiftOCR4,623almost 6 years agoFast and simple OCR library written in Swift. Optimized for recognizing short, one line long alphanumeric codes

Awesome OCR / Software / OCR training tools

glyph-miner34almost 10 years agoA system for extracting glyphs from early typeset prints
ocrodeg161over 6 years agoDocument image degradation for OCR data augmentation

Awesome OCR / Datasets / Ground Truth

archiscribe-corpus8over 7 years ago>4,200 lines transcribed from 19th Century German prints via
CIS OCR Test Set15about 5 years ago2 example documents each in German/Latin/Greek with ground truth for
Rescribe11almost 4 years agoTranscriptions of Caroline Minuscule Manuscripts
CLTKCorpora from
DIVA-HisDB150 pages of three medieval manuscripts
EarlyPrintedBooks10almost 9 years ago~8,800 lines from several early printed books
EEBO-TCP18over 5 years ago25,363 EEBO documents transcribed by
ECCO-TCP18over 5 years ago2,188 ECCO documents transcribed by
eMOP-TCP3over 10 years ago2,188 ECCO-TCP documents, cleaned up by
Evans-TCP18over 5 years ago4,977 Evans documents transcribed by
FDHNFinnish Digitised Historical Newspapers, , (free) required,
FROC-MSS0over 7 years ago4 Old French Medieval Manuscripts
GERMANA764 Spanish manuscript pages, (free) required
GT4HistOCRGround Truth for German Fraktur and Early Modern Latin
imagessan4about 8 years agoSanskrit images & ground truth (Devanagari script)
IMPACT-BHL2,418 pages from the Biodiversity Heritage Library,
IMPACT-BL294 pages from the British Library, (free) required
IMPACT-BNE215 pages from the National Library of Spain, (free) required,
IMPACT-BNF151 pages from the National Library of France, (free) required
IMPACT-KB142 pages from the National Library of the Netherlands
IMPACT-NKC187 pages from the Czech National Library, (free) required
IMPACT-NLB19 pages from the National Library of Bulgaria, (free) required
IMPACT-NUK209 pages from the National Library of Slovenia, (free) required
IMPACT-PSNC478 pages from four Polish digital libraries,
LascivaRoma/lexical1over 3 years agoTranscription of 19th century lexical resources for Latin learning
MJSynth9m synthetic images covering 90k English words
OCR19thSAC19,000 pages Swiss Alpine Club yearbooks transcribed via
OCR-D180 pages of German historical prints from
OCR_GS_Data15over 3 years agoDouble-checked Arabic Gold Standard from
old-books12about 9 years ago322 old books from
PRImA-ENP528 pages historic newspapers from , (free) required
RODRIGO853 Spanish manuscript pages, (free) required
Toebler-OCR1over 7 years ago(Kraken) Ground Truth transcription of few pages of the Tobler-Lommatzsch: Altfranzösisches Wörterbuch
IMPACT: Tools for text digitisationList of tools software projects related, some related to OCR
OCR-DList of OCR-related academic articles in the context of the project
Mendeley Group "OCR - Optical Character Recognition"Collection of 34 papers on OCR
eadh.org projectsList of Digital Humanities-related projects in Europe, some related to OCR
Wikipedia: Comparison of optical character recognition software
OCR [and Deep Learning]by
Ocropus Wiki: Publications3,426over 5 years ago

Awesome OCR / Literature / Blog Posts and Tutorials

Tesseract Blends Old and New OCR Technology262about 5 years ago(2016)
What You Always Wanted To Know About Tesseract(2014)
Extracting text from an image using Ocropus(2015)
Training an Ocropus OCR model(2015)
Ocropus Wiki: Compute errors and confusions3,426over 5 years ago(2016)
Ocropus Wiki: Working with Ground Truth3,426over 5 years ago(2016)
OCRopus(2016)
10 Tips for making your OCR project succeed(2013)
Overview of LEADTOOLS Image Cleanup and Pre-processing SDK Technology-
Extracting Text from PDFs; Doing OCR; all within R

Awesome OCR / Literature / Blog Posts and Tutorials / Extracting Text from PDFs; Doing OCR; all within R

R programming environmentHow to work with OCR from PDFs in the

Awesome OCR / Literature / Blog Posts and Tutorials

Tutorial: Command-line OCR on a Mac
Practical Expercience with OCRopus Model Training(2016)
Homemade Manuscript OCR (1): OCRopy(2017)
Optimizing Binarization for OCRopus(2017)
Prototype demo for OCR postfix in Danish Newspapers(2016)
How Can I OCR My Dictionary?(2016)
"Needlessly complex" blog(2016) . Several image processing how-tos (Python based), particularly:

Awesome OCR / Literature / Blog Posts and Tutorials / "Needlessly complex" blog

Page dewarping( )
Compressing and enhancing hand-written notes( )
Unprojecting text with ellipses( )

Awesome OCR / Literature / Blog Posts and Tutorials

(Open-Source-)OCR-Workflows(2017) overview of the state of the art in open source OCR and related technologies (binarisation, deskewing, layout recognition, etc.), lots of example images and information on the project
A gentle introduction to OCR(2018)
Worauf kann ich mich verlassen? Arbeiten mit digitalisierten Quellen, Teil 1: OCR(2019) A reflection/criticism on OCR quality, OCR pitfalls in Fraktur fonts

Awesome OCR / Literature / OCR Showcases

abbyy-finereader-ocr-senate129over 10 years agoUsing OCR to parse scanned Senate Financial Disclosure forms
cvOCR18almost 10 years agoAn OCR system for recognizing resume or cv text, implemented in Python and C and based on tesseract
MathOCR168almost 4 years agoA printed scientific document recognition system,

Awesome OCR / Literature / Academic articles

High performance document layout analysis(2003) Breuel
Adaptive degraded document image binarization(2006) Gatos, Pratikakis, Perantonis
[Internship Report](2007) Gupta
OCRopus Addons (Internship Report)(2007) Dantrey
Local Logistic Classifiers for Large Scale Learning(2012) Yousefi, Breuel
High Performance OCR for Printed English and Fraktur using LSTM Networks(2013) Breuel, Ul-Hasan, Mayce Al Azawi. Shafait
Can we build language-independent OCR using LSTM networks?(2013) Ul-Hasan, Breuel
Offline Printed Urdu Nastaleeq Script Recognition with Bidirectional LSTM Networks(2013) Ul-Hasan, Ahmed, Rashid, Shafait, Breuel
OCR of historical printings of Latin texts: Problems, Prospects, Progress.(2014) Springmann, Najock, Morgenroth, Schmid, Gotscharek, Fink
Correcting Noisy OCR: Context beats Confusion(2014) Evershed, Fitch
TypeWright: An Experiment in Participatory Curation(2015) Bilansky
Benchmarking of LSTM Networks(2015) Breuel
Recognition of Historical Greek Polytonic Scripts Using LSTM(2015) Simistira, Ul-Hassan, Papavassiliou, Basilis Gatos, Katsouros, Liwicki
A Segmentation-Free Approach for Printed Devanagari Script Recognition(2015) Karayil, Ul-Hasan, Breuel
A Sequence Learning Approach for Multiple Script Identification(2015) Ul-Hasan, Afzal, Shfait, Liwicki, Breuel
Important New Developments in Arabographic Optical Character Recognition (OCR)(2016) Romanov, Miller, Savant, Kiessling

Awesome OCR / Literature / Academic articles / Important New Developments in Arabographic Optical Character Recognition (OCR)

OpenArabic/OCR_GS_Data13over 9 years agousing for ground truth data

Awesome OCR / Literature / Academic articles

OCR of historical printings with an application to building diachronic corpora: A case study using the RIDGES herbal corpus(2016) Springmann, Lüdeling
Automatic quality evaluation and (semi-) automatic improvement of mixed models for OCR on historical documents(2016) Springmann, Fink, Schulz
Generic Text Recognition using Long Short-Term Memory Networks(2016) Ul-Hasan -- Ph.D Thesis
OCRoRACT: A Sequence Learning OCR System Trained on Isolated Characters(2016) Dengel, Ul-Hasan, Bukhari
Recursive Recurrent Nets with Attention Modeling for OCR in the Wild(2016) Lee, Osindero
Telugu OCR Framework using Deep Learning(2015/2017) , Hastie

Awesome OCR / Literature / Academic articles / Telugu OCR Framework using Deep Learning

TeluguOCRsee also , , ,

Awesome OCR / Literature / Academic articles

A Two-Stage Method for Text Line Detection in Historical Documents(2018) , Leifert, Strauß, Labahn. Code available at

Backlinks from these awesome lists:

More related projects: