mimeograph
PDF extractor
A CoffeeScript library for extracting text from PDF files and creating searchable documents with OCR capabilities
CoffeeScript lib for PDF OCR and text extraction
28 stars
5 watching
2 forks
Language: CoffeeScript
last commit: almost 14 years agoLinked from 1 awesome list
Related projects:
| Repository | Description | Stars |
|---|---|---|
| A tool to extract text from PDFs and add a searchable layer to them | 279 | |
| A tool for extracting pages from PDFs and converting them to images and text strings. | 216 | |
| A C# library for extracting and analyzing text from PDF files | 1,794 | |
| Evaluates PDF extraction tools' ability to extract meaningful text from scientific articles | 65 | |
| A data pipeline for extracting structured data from PDFs using AWS Textract and cloud-based services | 164 | |
| A Quarkus-based microservice to extract text from PDF files | 24 | |
| A C library that provides read and write access to PDF files. | 204 | |
| A tool to extract relevant information from text | 17 | |
| A tool to fetch information about public Google documents from various services | 856 | |
| A Swift library for generating PDF files from templates and web pages using wkhtmltopdf | 38 | |
| A Rust library for working with PDF documents | 1,680 | |
| A Go library for extracting text from PDF files, particularly invoices. | 708 | |
| A Rust-based framework for generating customizable PDFs with flexible layouts and content structures. | 18 | |
| Extracts readable content from web pages using Open Graph and traditional readability rules. | 69 | |
| Research and development of tools and techniques for extracting information from images and PDFs using deep learning and graph neural networks. | 96 |