web-scraper

HTML scraper

A Perl toolkit for extracting structured data from HTML documents using a DSL-like interface.

Perl web scraping toolkit

GitHub

104 stars
11 watching
31 forks
Language: Perl
last commit: over 9 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
fimad/scalpelA web scraping library providing a declarative interface on top of an HTML parsing library to extract data from HTML pages325
slotix/dataflowkitA framework for extracting structured data from web pages using CSS selectors.667
scrapy/scrapelyA pure-python library for extracting structured data from HTML pages.1,865
benibela/xidelA tool to extract data from web pages using various query languages and selectors.690
rust-scraper/scraperA Rust library for parsing and querying HTML documents using CSS selectors.1,961
jakopako/goskyrA tool to simplify web scraping of list-like structured data from web pages36
medialab/minetA command line tool and Python library for extracting data from various web sources.293
propublica/uptonA web scraping framework that simplifies the process by handling repetitive tasks and provides options for efficient data retrieval1,612
ruippeixotog/scala-scraperA Scala library providing a DSL for loading and extracting content from HTML pages717
jjelosua/doga_scraperA tool that extracts and converts Galician Official journal documents to different formats based on input year.0
the-markup/blacklight-collectorA tool for scraping website content and analyzing browser behavior205
felipecsl/wombatA Ruby-based web crawler and data extraction tool with an elegant DSL.1,315
meilisearch/docs-scraperAutomates scraping and indexing of documentation content into a search engine297
spider-rs/spiderA tool for web data extraction and processing using Rust1,234
zhuyingda/websterA framework for automating web scraping and crawling tasks using Node.js518