cola

Crawler library

A high-level framework for building distributed data extractors from web pages

A high-level distributed crawling framework.

GitHub

2k stars
166 watching
537 forks
Language: Python
last commit: about 4 years ago
Linked from 2 awesome lists


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
chenjiandongx/github-spiderA Python-based web crawler for scraping Github user and repository data.264
xianhu/pspiderA Python web crawler framework with support for multi-threading and proxy usage.1,828
zhegexiaohuozi/seimicrawlerA distributed crawler framework that simplifies the process of building crawlers using Spring Boot and Redis1,980
crypto-crawler/crypto-crawler-rsA Rust-based library for building and managing cryptocurrency crawlers235
feng19/spider_manA high-level web crawling and scraping framework for Elixir.23
howie6879/ruiaAn async web scraping micro-framework built with asyncio and aiohttp to simplify URL crawling1,753
turnersoftware/infinitycrawlerA web crawling library for .NET that allows customizable crawling and throttling of websites.248
kiddyuchina/beanbunA PHP framework for building distributed web crawlers with modular design and extensibility1,249
jmg/crawleyA Pythonic framework for building high-speed web crawlers with flexible data extraction and storage options.188
elliotgao2/gainA Python web crawling framework utilizing asyncio and aiohttp for efficient data extraction from websites.2,037
hu17889/go_spiderA modular, concurrent web crawler framework written in Go.1,827
elixir-crawly/crawlyA framework for extracting structured data from websites994
puerkitobio/gocrawlA concurrent web crawler written in Go that allows flexible and polite crawling of websites.2,036
felipecsl/wombatA Ruby-based web crawler and data extraction tool with an elegant DSL.1,315
fredwu/crawlerA high-performance web crawling and scraping solution with customizable settings and worker pooling.945