nidaba

OCR pipeline

Automates OCR pipeline for text digitization and conversion of raw images into citable texts.

An expandable and scalable OCR pipeline

GitHub

86 stars
9 watching
12 forks
Language: Python
last commit: almost 9 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
openphilology/tei-ocrCustomizes TEI XML for metadata from OCR processes to capture detailed layout and content information1
openseg-group/openseg.pytorchProvides a PyTorch implementation of several computer vision tasks including object detection, segmentation and parsing.1,191
bandrel/ocyaraPerforms OCR on images and scans them for matches to Yara rules40
allenai/scispacyCustom spaCy models and pipelines for scientific documents1,724
seven45/pdm-ciProvides a base image for creating Python CI pipelines with package manager support11
hhio618/golem-ciA decentralized task pipeline on Golem.network using Python.5
hamdikahloun/windows_ocrAn OCR library allowing developers to embed high-quality character recognition functionality in their products.18
bjpop/rubraA bioinformatics pipeline system that supports running workflow stages on a distributed compute cluster.38
openiti/ocr_gs_dataProvides gold standard data for training and testing optical character recognition (OCR) engines.15
sirfz/tesserocrAn OCR API wrapper that enables concurrent execution using Python's threading module and releases the GIL.2,026
ros-perception/image_pipelineA ROS package providing an image processing pipeline811
osciiart/deepaaGenerates ASCII art from images using deep learning-based convolutional neural networks1,524
calamari-ocr/calamariAn OCR engine with modular design and a command-line interface, providing pre-trained models and a Python API for customization.1,056
druths/xpA tool for creating flexible and self-documenting data science pipelines56