Awesome Lists

awesome-document-understanding

by tstanislawek

awesome listpushed over 3 years ago

A curated list of resources for Document Understanding (DU) topic

AI summary

DU tech

A curated collection of resources and papers on Document Understanding technology

stars
1.3K
forks
152
watching
37
awesome lists
2
entries
96
View on GitHub

Embed the badge

Show how many awesome lists link to your project. The count updates automatically.

Awesome Lists badge
Markdown
[![Awesome Lists Badge](https://awesome.facts.dev/shield/tstanislawek/awesome-document-understanding/links.svg)](https://awesome.facts.dev/awesome/tstanislawek/awesome-document-understanding)
HTML
<a href="https://awesome.facts.dev/awesome/tstanislawek/awesome-document-understanding"><img src="https://awesome.facts.dev/shield/tstanislawek/awesome-document-understanding/links.svg" alt="Awesome Lists Badge" /></a>
Image URL
https://awesome.facts.dev/shield/tstanislawek/awesome-document-understanding/links.svg

What's in the list

96 links in 9 sections, with live GitHub stats.activeno commit in 2y

Awesome Document Understanding

Introduction / Papers

Research topics

Resources

  • The RVL-CDIP Dataset

    dataset consists of 400,000 grayscale images in 16 classes, with 25,000 images per class

  • The Industry Documents Library

    a portal to millions of documents created by industries that influence public health, hosted by the UCSF Library

  • Color Document Dataset

    from the Intelligent Sensory Information Systems, University of Amsterdam

  • The IIT CDIP Collection

    dataset consists of documents from the states' lawsuit against the tobacco industry in the 1990s, consists of around 7 million documents

  • borb

    is a pure python library to read, write and manipulate PDF documents. It represents a PDF document as a JSON-like datastructure of nested lists, dictionaries and primitives (numbers, string, booleans, etc)

  • pawls

    PDF Annotations with Labels and Structure is software that makes it easy to collect a series of annotations associated with a PDF document

  • pdfplumber

    Plumb a PDF for detailed information about each text character, rectangle, and line. Plus: Table extraction and visual debugging

  • Pdfminer.six

    Pdfminer.six is a community maintained fork of the original PDFMiner. It is a tool for extracting information from PDF documents. It focuses on getting and analyzing text data

  • Layout Parser

    Layout Parser is a deep learning based tool for document image layout analysis tasks

  • Tabulo

    Table extraction from images

  • OCRmyPDF

    OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched or copy-pasted

  • PDFBox

    The Apache PDFBox library is an open source Java tool for working with PDF documents. This project allows creation of new PDF documents, manipulation of existing documents and the ability to extract content from documents

  • PdfPig

    This project allows users to read and extract text and other content from PDF files. In addition the library can be used to create simple PDF documents containing text and geometrical shapes. This project aims to port PDFBox to C#

  • parsing-prickly-pdfs

    Resources and worksheet for the NICAR 2016 workshop of the same name

  • Born digital pdf scanner

    checking if pdf is born-digital

  • OpenContracts

    Apache2-licensed, PDF annotating platform for visually-rich documents that preserves the original layout and exports x,y positional data for tokens as well as span starts and stops. Based on PAWLs, but with a Python-based backend and readily deployable on your local machine, company intranet or the web via Docker Compose

  • deepdoctection

    doctection is a Python library that orchestrates document extraction and document layout analysis tasks for images and pdf documents using deep learning models. It does not implement models but enables you to build pipelines using highly acknowledged libraries for object detection, OCR and selected NLP tasks and provides an integrated framework for fine-tuning, evaluating and running models

  • pydoxtools

    Pydoxtools is an AI-composition library for dpocument analysis. It features an extensive toolset for building complex document analysis pipelines and recognizes most document formats out of the box. It supports typical NLP tasks such as keywords, summarization, question_answering out of the box. and features a high quality low-CPU/memory table extraction algorithm and makes NLP batch operations on a cluster easy

Conferences, workshops

Blogs

Solutions

Inspirations

More related projects

Add a GitHub project

Missing a project or an awesome list? Paste its GitHub URL and we fetch it right away.