gain

crawler

A Python web crawling framework utilizing asyncio and aiohttp for efficient data extraction from websites.

Web crawling framework based on asyncio.

GitHub

2k stars
75 watching
208 forks
Language: Python
last commit: over 7 years ago
Linked from 2 awesome lists

aiohttpasynciocrawlerpythonspideruvloop

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
howie6879/ruiaAn async web scraping micro-framework built with asyncio and aiohttp to simplify URL crawling1,753
jmg/crawleyA Pythonic framework for building high-speed web crawlers with flexible data extraction and storage options.188
chenjiandongx/github-spiderA Python-based web crawler for scraping Github user and repository data.264
feng19/spider_manA high-level web crawling and scraping framework for Elixir.23
xianhu/pspiderA Python web crawler framework with support for multi-threading and proxy usage.1,828
puerkitobio/fetchbotA flexible web crawler that follows robots.txt policies and crawl delays.787
archiveteam/grab-siteA web crawler designed to backup websites by recursively crawling and writing WARC files.1,406
hu17889/go_spiderA modular, concurrent web crawler framework written in Go.1,827
a11ywatch/crawlerPerforms web page crawling at high performance.51
untwisted/sukhoiA minimalist web crawler framework built on top of miners and structure-based data extraction879
fredwu/crawlerA high-performance web crawling and scraping solution with customizable settings and worker pooling.945
cocrawler/cocrawlerA versatile web crawler built with modern tools and concurrency to handle various crawl tasks188
puerkitobio/gocrawlA concurrent web crawler written in Go that allows flexible and polite crawling of websites.2,036
rivermont/spidyA simple command-line web crawler that automatically extracts links from web pages and can be run in parallel for efficient crawling340
qinxuye/colaA high-level framework for building distributed data extractors from web pages1,501