gecco

Web Crawler Framework

A lightweight web crawler framework that enables easy extraction of web page data using jQuery-like selectors and supports asynchronous requests and distributed crawling.

Easy to use lightweight web crawler(易用的轻量化网络爬虫)

GitHub

3k stars
144 watching
891 forks
Language: Java
last commit: over 2 years ago
Linked from 1 awesome list

crawlerdynamicfastjsongeccojavajsoup

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
yujiosaka/headless-chrome-crawlerA distributed crawling framework that leverages Headless Chrome to scrape dynamic websites5,534
code4craft/webmagicA framework for building scalable web crawlers in Java.11,456
zhegexiaohuozi/seimicrawlerA distributed crawler framework that simplifies the process of building crawlers using Spring Boot and Redis1,980
yasserg/crawler4jA Java-based web crawler for extracting and processing web page content4,563
geziyor/geziyorA fast and flexible web crawling and scraping framework for extracting structured data from websites.2,646
apache/incubator-stormcrawlerA scalable and versatile web crawling framework based on Apache Storm895
mozilla/geckodriverAn HTTP API proxy for interacting with Gecko-based browsers like Firefox7,223
apify/crawleeA tool for building reliable web scraping and browser automation pipelines in Node.js.16,081
stewartmckee/cobwebA flexible web crawler that can be used to extract data from websites in a scalable and efficient manner226
unclecode/crawl4aiA web crawling tool designed to extract structured data from the web for use in AI applications18,541
matteoredaelli/ebotAn Erlang-based web crawler designed to be scalable and highly configurable330
hakluke/hakrawlerA tool for automatically discovering and crawling web application endpoints and assets4,528
howie6879/ruiaAn async web scraping micro-framework built with asyncio and aiohttp to simplify URL crawling1,753
elliotgao2/gainA Python web crawling framework utilizing asyncio and aiohttp for efficient data extraction from websites.2,037
brendonboshell/supercrawlerA web crawler designed to crawl websites while obeying robots.txt rules, rate limits and concurrency limits, with customizable content handlers for parsing and processing crawled pages.380