awesome-document-understanding
by tstanislawek
A curated list of resources for Document Understanding (DU) topic
AI summary
DU tech
A curated collection of resources and papers on Document Understanding technology
- stars
- 1.3K
- forks
- 152
- watching
- 37
- awesome lists
- 2
- entries
- 96
- #awesome
- #awesome-list
- #deep-learning
- #document-ai
- #document-analysis
- #document-intelligence
- #document-layout-analysis
- #document-understanding
- #information-extraction
- #intelligent-processing
- #key-information-extraction
- #machine-learning
- #natural-language-processing
- #nlp
- #ocr
- #pdf-documents
- #robotic-process-automation
- #rpa
- #unstructured-data
What's in the list
96 links in 9 sections, with live GitHub stats.activeno commit in 2y
Awesome Document Understanding
Introduction / Papers
Research topics
Research topics / Related
Resources
- The RVL-CDIP Dataset
dataset consists of 400,000 grayscale images in 16 classes, with 25,000 images per class
- The Industry Documents Library
a portal to millions of documents created by industries that influence public health, hosted by the UCSF Library
- Color Document Dataset
from the Intelligent Sensory Information Systems, University of Amsterdam
- The IIT CDIP Collection
dataset consists of documents from the states' lawsuit against the tobacco industry in the 1990s, consists of around 7 million documents
borb
is a pure python library to read, write and manipulate PDF documents. It represents a PDF document as a JSON-like datastructure of nested lists, dictionaries and primitives (numbers, string, booleans, etc)
pawls
PDF Annotations with Labels and Structure is software that makes it easy to collect a series of annotations associated with a PDF document
pdfplumber
Plumb a PDF for detailed information about each text character, rectangle, and line. Plus: Table extraction and visual debugging
Pdfminer.six
Pdfminer.six is a community maintained fork of the original PDFMiner. It is a tool for extracting information from PDF documents. It focuses on getting and analyzing text data
Layout Parser
Layout Parser is a deep learning based tool for document image layout analysis tasks
Tabulo
Table extraction from images
OCRmyPDF
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched or copy-pasted
PDFBox
The Apache PDFBox library is an open source Java tool for working with PDF documents. This project allows creation of new PDF documents, manipulation of existing documents and the ability to extract content from documents
PdfPig
This project allows users to read and extract text and other content from PDF files. In addition the library can be used to create simple PDF documents containing text and geometrical shapes. This project aims to port PDFBox to C#
- parsing-prickly-pdfs
Resources and worksheet for the NICAR 2016 workshop of the same name
pdf-text-extraction-benchmark
PDF tools benchmark
Born digital pdf scanner
checking if pdf is born-digital
OpenContracts
Apache2-licensed, PDF annotating platform for visually-rich documents that preserves the original layout and exports x,y positional data for tokens as well as span starts and stops. Based on PAWLs, but with a Python-based backend and readily deployable on your local machine, company intranet or the web via Docker Compose
deepdoctection
doctection is a Python library that orchestrates document extraction and document layout analysis tasks for images and pdf documents using deep learning models. It does not implement models but enables you to build pipelines using highly acknowledged libraries for object detection, OCR and selected NLP tasks and provides an integrated framework for fine-tuning, evaluating and running models
pydoxtools
Pydoxtools is an AI-composition library for dpocument analysis. It features an extensive toolset for building complex document analysis pipelines and recognizes most document formats out of the box. It supports typical NLP tasks such as keywords, summarization, question_answering out of the box. and features a high quality low-CPU/memory table extraction algorithm and makes NLP batch operations on a cluster easy
Conferences, workshops
- 2021
[ , , ]
- 2021
Workshop on Document Intelligence (DI) [ , ]
- 2021
Financial Narrative Processing Workshop (FNP) [ , , ]
- 2021
Workshop on Economics and Natural Language Processing (ECONLP) [ , , ]
- 2020
INTERNATIONAL WORKSHOP ON DOCUMENT ANALYSIS SYSTEMS (DAS) [ , , ]
- 2020
International Workshop on SCIentific DOCument Analysis (SCIDOCA) [ , , ]
Blogs
- Document Form Extraction
, 2021
Solutions
Inspirations
Nothing in this list matches your filter.
Featured in 2 awesome lists
Each link jumps to the spot where the list mentions awesome-document-understanding.
More related projects
tleyden/open-ocr1.3K
ibm/max-ocr47
dannnylo/tesseract-ocr-elixir55
waitingcheung/artrailer15
iuliaturc/detextify271
namuan/dr-doc-search603
robertmartin8/pyportfolioopt4.6K
arocks/edge841
lyst/lightfm4.8K
graphql-python/graphql-core516
bauerji/flask-pydantic369
orion-ai-lab/kurosiwo44
axiros/terminal_markdown_viewer1.8K
pbkhrv/ulauncher-keepassxc19