goreadability

Web page summary extractor

Extracts readable content from web pages using Open Graph and traditional readability rules.

Webpage summary extractor using Facebook Open Graph and arc90's readability

GitHub

69 stars
7 watching
8 forks
Language: Go
last commit: over 7 years ago
Linked from 2 awesome lists

opengraphreadabilityscraper

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
keepcosmos/readabilityAn Elixir library that extracts and curates primary readable content from web pages.260
tjatse/node-readabilityAutomates web page scraping and text extraction to make any webpage readable343
philipperemy/stanford-openie-pythonProvides a Python interface to extract structured relation triples from plain text using CoreNLP's open information extraction system.639
jonmagic/grimA tool for extracting pages from PDFs and converting them to images and text strings.216
erikriver/opengraphA Python module to extract and parse metadata from web pages using the Open Graph Protocol.230
cantino/ruby-readabilityA Ruby port of a readability tool that extracts primary content from web pages.927
foolin/pagserA tool for automatically extracting structured data from HTML pages105
neon-jungle/wagtail-readabilityAnalogizes the readability of text content in Wagtail's RichTextField16
s0rg/crawleyA utility for systematically extracting URLs from web pages and printing them to the console.268
peburrows/plotA GraphQL parser and resolver for Elixir that aims to implement the full GraphQL spec.32
vrothberg/vgrepA user-friendly pager for text search and editing669
itteco/iframelyA service that extracts metadata and embeds from web pages1,537
steelthread/mimeographA CoffeeScript library for extracting text from PDF files and creating searchable documents with OCR capabilities28
serpapi/nokolexborA high-performance HTML5 parser for Ruby based on Lexbor with support for CSS selectors and XPath.327
plainas/tqTool that extracts content from HTML documents based on CSS selectors236