sparkler

Web Crawler

A high-performance web crawler built on Apache Spark that fetches and analyzes web resources in real-time.

Spark-Crawler: Apache Nutch-like crawler that runs on Apache Spark.

GitHub

411 stars
45 watching
141 forks
Language: Java
last commit: over 3 years ago
Linked from 1 awesome list

big-datadistributed-systemsinformation-retrievalnutchsearchsearch-enginesolrsparktikaweb-crawler

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
internetarchive/sparklingA data processing library built on top of Apache Spark to handle temporal web data11
stewartmckee/cobwebA flexible web crawler that can be used to extract data from websites in a scalable and efficient manner226
apache/incubator-stormcrawlerA scalable and versatile web crawling framework based on Apache Storm895
vida-nyu/acheA web crawler designed to efficiently collect and prioritize relevant content from the web459
internetarchive/brozzlerA distributed web crawler that fetches and extracts links from websites using a real browser.678
apache/sparkAn analytics engine designed to handle large-scale data processing and analysis40,170
brendonboshell/supercrawlerA web crawler designed to crawl websites while obeying robots.txt rules, rate limits and concurrency limits, with customizable content handlers for parsing and processing crawled pages.380
archiveteam/grab-siteA web crawler designed to backup websites by recursively crawling and writing WARC files.1,406
tweag/sparkleA tool for creating resilient, scalable analytics applications with Haskell on top of Apache Spark447
hightman/pspiderA parallel web crawler framework built using PHP and MySQLi266
c-sto/recursebusterA tool for recursively querying web servers by sending HTTP requests and analyzing responses to discover hidden content243
postmodern/spidrA Ruby web crawling library that provides flexible and customizable methods to crawl websites809
a11ywatch/crawlerPerforms web page crawling at high performance.51
webrecorder/browsertrix-crawlerA containerized browser-based crawler system for capturing web content in a high-fidelity and customizable manner.677
spider-rs/spiderA tool for web data extraction and processing using Rust1,234