scrapy-cluster

Crawler cluster

A distributed scraping framework that scales crawling and prioritizes sites, utilizing Redis and Kafka for coordination.

This Scrapy project uses Redis and Kafka to create a distributed on demand scraping cluster.

GitHub

1k stars
108 watching
323 forks
Language: Python
last commit: almost 3 years ago
Linked from 1 awesome list

distributedkafkapythonredisscrapingscrapy

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
scrapy/scrapelyA pure-python library for extracting structured data from HTML pages.1,865
dyweb/scralaA web crawling framework written in Scala that allows users to define the start URL and parse response from it113
pjkelly/robocopA middleware that adds a meta tag to HTTP responses to instruct search engines on how to crawl the content.3
cuiweixie/lua-resty-redis-clusterA client library for managing Redis clusters using Lua scripts in an OpenResty configuration.100
efremidze/clusterA map annotation clustering library that efficiently groups and displays geographic pins on an iOS map view.1,274
postmodern/spidrA Ruby web crawling library that provides flexible and customizable methods to crawl websites809
rndinfosecguy/scavengerAn OSINT bot that crawls pastebin sites to search for sensitive data leaks634
rusty1s/pytorch_clusterA PyTorch extension library providing optimized graph cluster algorithms838
elixir-crawly/crawlyA framework for extracting structured data from websites994
holgerd77/django-dynamic-scraperAn app that allows you to manage Scrapy spiders through a Django admin interface.1,155
howie6879/ruiaAn async web scraping micro-framework built with asyncio and aiohttp to simplify URL crawling1,753
needmorecowbell/giggityA tool to scrape and store hierarchical data about GitHub organizations, users, or repositories.127
malfrats/xeuledocA tool to fetch information about public Google documents from various services856
stewartmckee/cobwebA flexible web crawler that can be used to extract data from websites in a scalable and efficient manner226
tidyverse/rvestA package for extracting data from web pages using HTML parsing and CSS/XPath selectors.1,495