node-readability

Web scraper

Automates web page scraping and text extraction to make any webpage readable

Scrape/Crawl article from any site automatically. Make any web page readable, no matter Chinese or English.

GitHub

343 stars
11 watching
36 forks
Language: JavaScript
last commit: about 8 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
zhuyingda/websterA framework for automating web scraping and crawling tasks using Node.js518
joseconstela/webparsyA Node.js library and CLI for scraping websites using Puppeteer and YAML definitions44
philipjkim/goreadabilityExtracts readable content from web pages using Open Graph and traditional readability rules.69
retextjs/retext-readabilityA plugin to assess text readability using various algorithms94
plainas/tqTool that extracts content from HTML documents based on CSS selectors236
chaijs/loupeAn object inspection utility that produces human-readable representations of objects across different platforms and environments.22
jjelosua/doga_scraperA tool that extracts and converts Galician Official journal documents to different formats based on input year.0
litt1e-p/weapp-girlsA Node.js-based web scraping project to extract photos from popular Chinese women's interest websites.247
felipecsl/wombatA Ruby-based web crawler and data extraction tool with an elegant DSL.1,315
nodejs/readable-streamProvides a Node.js implementation of the core streams classes for userland development1,033
amoilanen/js-crawlerA Node.js module for crawling web sites and scraping their content254
miyagawa/web-scraperA Perl toolkit for extracting structured data from HTML documents using a DSL-like interface.104
gmarty/xgettextTools for extracting translatable strings from source code written in template languages.77
disjukr/just-newsA userscript project that parses Korean news site and makes the content more readable191
tj/redsA lightweight search module for Node.js applications using Redis as the backing store.890