rubyretriever

Web crawler

A Ruby-based tool for web crawling and data extraction, aiming to be a replacement for paid software in the SEO space.

Asynchronous Web Crawler & Scraper

GitHub

143 stars
7 watching
26 forks
Language: Ruby
last commit: over 3 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
postmodern/spidrA Ruby web crawling library that provides flexible and customizable methods to crawl websites809
jaimeiniesta/metainspectorA Ruby gem for web scraping and extracting metadata from web pages.1,038
spider-rs/spiderA tool for web data extraction and processing using Rust1,234
webrecorder/browsertrix-crawlerA containerized browser-based crawler system for capturing web content in a high-fidelity and customizable manner.677
stewartmckee/cobwebA flexible web crawler that can be used to extract data from websites in a scalable and efficient manner226
rivermont/spidyA simple command-line web crawler that automatically extracts links from web pages and can be run in parallel for efficient crawling340
internetarchive/brozzlerA distributed web crawler that fetches and extracts links from websites using a real browser.678
brendonboshell/supercrawlerA web crawler designed to crawl websites while obeying robots.txt rules, rate limits and concurrency limits, with customizable content handlers for parsing and processing crawled pages.380
felipecsl/wombatA Ruby-based web crawler and data extraction tool with an elegant DSL.1,315
a11ywatch/crawlerPerforms web page crawling at high performance.51
iamstoxe/urlgrabA tool to crawl websites by exploring links recursively with support for JavaScript rendering.331
archiveteam/grab-siteA web crawler designed to backup websites by recursively crawling and writing WARC files.1,406
amoilanen/js-crawlerA Node.js module for crawling web sites and scraping their content254
pjkelly/robocopA middleware that adds a meta tag to HTTP responses to instruct search engines on how to crawl the content.3
fredwu/crawlerA high-performance web crawling and scraping solution with customizable settings and worker pooling.945