wombat

Web scraper library

A Ruby-based web crawler and data extraction tool with an elegant DSL.

Lightweight Ruby web crawler/scraper with an elegant DSL which extracts structured data from pages.

GitHub

1k stars
51 watching
129 forks
Language: Ruby
last commit: over 2 years ago
Linked from 3 awesome lists

crawlerdslrubyscraper

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
bplawler/crawlerA Scala-based DSL for programmatically accessing and interacting with web pages149
benibela/xidelA tool to extract data from web pages using various query languages and selectors.690
postmodern/spidrA Ruby web crawling library that provides flexible and customizable methods to crawl websites809
jaimeiniesta/metainspectorA Ruby gem for web scraping and extracting metadata from web pages.1,038
ruippeixotog/scala-scraperA Scala library providing a DSL for loading and extracting content from HTML pages717
archiveteam/wpullDownloads and crawls web pages, allowing for the archiving of websites.556
miyagawa/web-scraperA Perl toolkit for extracting structured data from HTML documents using a DSL-like interface.104
joseconstela/webparsyA Node.js library and CLI for scraping websites using Puppeteer and YAML definitions44
medialab/minetA command line tool and Python library for extracting data from various web sources.293
oscarotero/embedA PHP library to retrieve metadata and embed code from any web page2,100
slotix/dataflowkitA framework for extracting structured data from web pages using CSS selectors.667
spider-rs/spiderA tool for web data extraction and processing using Rust1,234
jjelosua/doga_scraperA tool that extracts and converts Galician Official journal documents to different formats based on input year.0
s0rg/crawleyA utility for systematically extracting URLs from web pages and printing them to the console.268
fimad/scalpelA web scraping library providing a declarative interface on top of an HTML parsing library to extract data from HTML pages325