SeimiCrawler

Crawler framework

A distributed crawler framework that simplifies the process of building crawlers using Spring Boot and Redis

一个简单、敏捷、分布式的支持SpringBoot的Java爬虫框架;An agile, distributed crawler framework.

GitHub

2k stars
176 watching
681 forks
Language: Java
last commit: almost 2 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
codesofun/web-beeA Java framework for building web-based crawlers with features like distributed crawling and proxy support.189
wspl/creeperA framework for building cross-platform web crawlers using Go780
crawlzone/crawlzoneA PHP framework for asynchronous internet crawling and web scraping78
turnersoftware/infinitycrawlerA web crawling library for .NET that allows customizable crawling and throttling of websites.248
kiddyuchina/beanbunA PHP framework for building distributed web crawlers with modular design and extensibility1,249
hu17889/go_spiderA modular, concurrent web crawler framework written in Go.1,827
dyweb/scralaA web crawling framework written in Scala that allows users to define the start URL and parse response from it113
apache/incubator-stormcrawlerA scalable and versatile web crawling framework based on Apache Storm895
untwisted/sukhoiA minimalist web crawler framework built on top of miners and structure-based data extraction879
qinxuye/colaA high-level framework for building distributed data extractors from web pages1,501
jmg/crawleyA Pythonic framework for building high-speed web crawlers with flexible data extraction and storage options.188
brendonboshell/supercrawlerA web crawler designed to crawl websites while obeying robots.txt rules, rate limits and concurrency limits, with customizable content handlers for parsing and processing crawled pages.380
howie6879/ruiaAn async web scraping micro-framework built with asyncio and aiohttp to simplify URL crawling1,753
hominee/dyerA fast and flexible web crawling tool with features like asynchronous I/O and event-driven design.135
antchfx/antchA framework for building fast and efficient web crawlers and scrapers in Go.261