MSpider

Web Crawler

A Python-based tool for web crawling and data collection from various websites

Spider

GitHub

348 stars
55 watching
191 forks
Language: Python
last commit: about 4 years ago
Linked from 1 awesome list

mspider

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
rivermont/spidyA simple command-line web crawler that automatically extracts links from web pages and can be run in parallel for efficient crawling340
jmg/crawleyA Pythonic framework for building high-speed web crawlers with flexible data extraction and storage options.188
xianhu/pspiderA Python web crawler framework with support for multi-threading and proxy usage.1,828
hightman/pspiderA parallel web crawler framework built using PHP and MySQLi266
stewartmckee/cobwebA flexible web crawler that can be used to extract data from websites in a scalable and efficient manner226
medialab/minetA command line tool and Python library for extracting data from various web sources.293
mvdbos/php-spiderA flexible PHP web crawler with configurable traversal algorithms and filters.1,336
chenjiandongx/github-spiderA Python-based web crawler for scraping Github user and repository data.264
feng19/spider_manA high-level web crawling and scraping framework for Elixir.23
elliotgao2/gainA Python web crawling framework utilizing asyncio and aiohttp for efficient data extraction from websites.2,037
postmodern/spidrA Ruby web crawling library that provides flexible and customizable methods to crawl websites809
spider-rs/spiderA tool for web data extraction and processing using Rust1,234
fmpwizard/owlcrawlerA distributed web crawler that coordinates crawling tasks across multiple worker processes using a message bus.55
3nock/spidersuiteA cross-platform web spider/crawler tool for analyzing and mapping attack surfaces614
holgerd77/django-dynamic-scraperAn app that allows you to manage Scrapy spiders through a Django admin interface.1,155