Github-spider

Crawler

A Python-based web crawler for scraping Github user and repository data.

Github 仓库及用户分析爬虫

GitHub

264 stars
15 watching
91 forks
Language: Python
last commit: over 9 years ago
Linked from 1 awesome list

crawlergithubscrapy

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
qinxuye/colaA high-level framework for building distributed data extractors from web pages1,501
hu17889/go_spiderA modular, concurrent web crawler framework written in Go.1,827
chenzixinn/spider_reverseA collection of examples demonstrating reverse engineering of web scraping and API interactions in Python617
elliotgao2/gainA Python web crawling framework utilizing asyncio and aiohttp for efficient data extraction from websites.2,037
feng19/spider_manA high-level web crawling and scraping framework for Elixir.23
xianhu/pspiderA Python web crawler framework with support for multi-threading and proxy usage.1,828
puerkitobio/gocrawlA concurrent web crawler written in Go that allows flexible and polite crawling of websites.2,036
fredwu/crawlerA high-performance web crawling and scraping solution with customizable settings and worker pooling.945
holgerd77/django-dynamic-scraperAn app that allows you to manage Scrapy spiders through a Django admin interface.1,155
spider-rs/spiderA tool for web data extraction and processing using Rust1,234
elixir-crawly/crawlyA framework for extracting structured data from websites994
jmg/crawleyA Pythonic framework for building high-speed web crawlers with flexible data extraction and storage options.188
3nock/spidersuiteA cross-platform web spider/crawler tool for analyzing and mapping attack surfaces614
rndinfosecguy/scavengerAn OSINT bot that crawls pastebin sites to search for sensitive data leaks634
puerkitobio/fetchbotA flexible web crawler that follows robots.txt policies and crawl delays.787