pdf2pdfocr

PDF extractor

A tool to extract text from PDFs and add a searchable layer to them

A free tool to OCR a PDF and add a text "layer" in the original file, making a searchable PDF. Use only open source tools. Please tip!

GitHub

279 stars
12 watching
35 forks
Language: Python
last commit: over 2 years ago
Linked from 1 awesome list

dockerocrpdfpdftkpythontesseract

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
steelthread/mimeographA CoffeeScript library for extracting text from PDF files and creating searchable documents with OCR capabilities28
unidoc/unidocA Go library for extracting text from PDF files, particularly invoices.708
tabulapdf/tabula-javaExtracts tables from PDF files using Java1,859
uglytoad/pdfpigA C# library for extracting and analyzing text from PDF files1,794
ckorzen/pdf-text-extraction-benchmarkEvaluates PDF extraction tools' ability to extract meaningful text from scientific articles65
aeksco/aws-pdf-textract-pipelineA data pipeline for extracting structured data from PDFs using AWS Textract and cloud-based services164
jesparza/peepdfA Python tool for analyzing PDF files to identify potential security risks and malicious content.1,319
malfrats/xeuledocA tool to fetch information about public Google documents from various services856
docraptor/docraptor-rubyA Ruby client library for converting HTML to PDF using the DocRaptor API.33
hiddenillusion/analyzepdfA tool to analyze PDF files by examining their characteristics to determine if they are malicious or benign.178
pdf-archiver/pdf-archiverA tool for digitizing and organizing paper documents by scanning and tagging files for easy searching.308
jonmagic/grimA tool for extracting pages from PDFs and converting them to images and text strings.216
bikash/documentunderstandingResearch and development of tools and techniques for extracting information from images and PDFs using deep learning and graph neural networks.96
philsturgeon/codeigniter-unzipA CodeIgniter extension that extracts ZIP files without requiring PECL extensions78
enferex/pdfresurrectAnalyzes and extracts previous versions of a PDF document to reconstruct its modification history81