mimeograph

PDF extractor

A CoffeeScript library for extracting text from PDF files and creating searchable documents with OCR capabilities

CoffeeScript lib for PDF OCR and text extraction

GitHub

28 stars
5 watching
2 forks
Language: CoffeeScript
last commit: almost 14 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
leofcardoso/pdf2pdfocrA tool to extract text from PDFs and add a searchable layer to them279
jonmagic/grimA tool for extracting pages from PDFs and converting them to images and text strings.216
uglytoad/pdfpigA C# library for extracting and analyzing text from PDF files1,794
ckorzen/pdf-text-extraction-benchmarkEvaluates PDF extraction tools' ability to extract meaningful text from scientific articles65
aeksco/aws-pdf-textract-pipelineA data pipeline for extracting structured data from PDFs using AWS Textract and cloud-based services164
gunnarmorling/quarkus-pdf-extractA Quarkus-based microservice to extract text from PDF files24
michaelrsweet/pdfioA C library that provides read and write access to PDF files.204
aymericbeaumet/squeezeA tool to extract relevant information from text17
malfrats/xeuledocA tool to fetch information about public Google documents from various services856
mihaelisaev/wkhtmltopdfA Swift library for generating PDF files from templates and web pages using wkhtmltopdf38
j-f-liu/lopdfA Rust library for working with PDF documents1,680
unidoc/unidocA Go library for extracting text from PDF files, particularly invoices.708
sowcow/blank_slate_pdfA Rust-based framework for generating customizable PDFs with flexible layouts and content structures.18
philipjkim/goreadabilityExtracts readable content from web pages using Open Graph and traditional readability rules.69
bikash/documentunderstandingResearch and development of tools and techniques for extracting information from images and PDFs using deep learning and graph neural networks.96