toxy

Document extractor

A .NET framework for extracting text from various document formats across multiple platforms.

.net text extraction framework

GitHub

362 stars
39 watching
107 forks
Language: C#
last commit: almost 2 years ago
Linked from 2 awesome lists


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
felipecsl/wombatA Ruby-based web crawler and data extraction tool with an elegant DSL.1,315
ckorzen/pdf-text-extraction-benchmarkEvaluates PDF extraction tools' ability to extract meaningful text from scientific articles65
dbuenzli/uusegAn OCaml library for segmenting Unicode text into grapheme clusters, words, and sentences.23
eyurtsev/korAn open-source wrapper around LLMs to extract structured data from text1,638
nikolamilosevic86/tabinoutA framework for extracting information from tables in scientific literature using a rule-based approach.42
xyntopia/pydoxtoolsA Python library for extracting information from unstructured documents using AI techniques and customizable pipelines.78
meilisearch/docs-scraperAutomates scraping and indexing of documentation content into a search engine297
sillsdev/standardformatlibA C# library for reading and writing files using standard format markers0
s0rg/crawleyA utility for systematically extracting URLs from web pages and printing them to the console.268
jjelosua/doga_scraperA tool that extracts and converts Galician Official journal documents to different formats based on input year.0
sinairv/yaxlibA flexible XML serialization library for .NET Framework and .NET Core0
feichao93/temmeA lightweight, CSS-based selector for extracting structured data from HTML documents.273
fielddb/multilingualcorporaextractorExtracts and formats multilingual corpora from international bibles into XML, JSON, and HTML files for analysis.0
tjatse/node-readabilityAutomates web page scraping and text extraction to make any webpage readable343
aymericbeaumet/squeezeA tool to extract relevant information from text17