PSpider

Web Crawler Framework

A Python web crawler framework with support for multi-threading and proxy usage.

简单易用的Python爬虫框架,QQ交流群:597510560

GitHub

2k stars
114 watching
503 forks
Language: Python
last commit: over 4 years ago
Linked from 1 awesome list

crawlermulti-threadingmultiprocessingproxiespythonpython-spiderspiderweb-crawlerweb-spider

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
hightman/pspiderA parallel web crawler framework built using PHP and MySQLi266
qinxuye/colaA high-level framework for building distributed data extractors from web pages1,501
chenjiandongx/github-spiderA Python-based web crawler for scraping Github user and repository data.264
manning23/mspiderA Python-based tool for web crawling and data collection from various websites348
elliotgao2/gainA Python web crawling framework utilizing asyncio and aiohttp for efficient data extraction from websites.2,037
jmg/crawleyA Pythonic framework for building high-speed web crawlers with flexible data extraction and storage options.188
howie6879/ruiaAn async web scraping micro-framework built with asyncio and aiohttp to simplify URL crawling1,753
feng19/spider_manA high-level web crawling and scraping framework for Elixir.23
kiddyuchina/beanbunA PHP framework for building distributed web crawlers with modular design and extensibility1,249
dyweb/scralaA web crawling framework written in Scala that allows users to define the start URL and parse response from it113
hu17889/go_spiderA modular, concurrent web crawler framework written in Go.1,827
postmodern/spidrA Ruby web crawling library that provides flexible and customizable methods to crawl websites809
wspl/creeperA framework for building cross-platform web crawlers using Go780
zhegexiaohuozi/seimicrawlerA distributed crawler framework that simplifies the process of building crawlers using Spring Boot and Redis1,980
rivermont/spidyA simple command-line web crawler that automatically extracts links from web pages and can be run in parallel for efficient crawling340