crawler

Crawler

A powerful web crawler written in PHP that can execute JavaScript and crawl multiple URLs concurrently.

An easy to use, powerful crawler implemented in PHP. Can execute Javascript.

GitHub

3k stars
66 watching
360 forks
Language: PHP
last commit: almost 2 years ago
Linked from 1 awesome list

concurrencycrawlerguzzlephp

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
apify/crawleeA tool for building reliable web scraping and browser automation pipelines in Node.js.16,081
yujiosaka/headless-chrome-crawlerA distributed crawling framework that leverages Headless Chrome to scrape dynamic websites5,534
jae-jae/querylistA PHP framework for building web scrapers and crawlers with a focus on ease of use and extensibility.2,671
unclecode/crawl4aiA web crawling tool designed to extract structured data from the web for use in AI applications18,541
spatie/laravel-site-searchA package to create a private search index by crawling and indexing a website275
code4craft/webmagicA framework for building scalable web crawlers in Java.11,456
stewartmckee/cobwebA flexible web crawler that can be used to extract data from websites in a scalable and efficient manner226
ruipgil/scraperjsA versatile web scraping module with two scrapers for static and dynamic content extraction.3,714
crawlzone/crawlzoneA PHP framework for asynchronous internet crawling and web scraping78
yasserg/crawler4jA Java-based web crawler for extracting and processing web page content4,563
veliovgroup/spiderable-middleware intercepts requests from web crawlers and proxies them to a prerendering service for rendering HTML39
uscdatascience/sparklerA high-performance web crawler built on Apache Spark that fetches and analyzes web resources in real-time.411
spekulatius/phpscraperA web scraping utility for PHP that simplifies the process of extracting information from websites.544
brendonboshell/supercrawlerA web crawler designed to crawl websites while obeying robots.txt rules, rate limits and concurrency limits, with customizable content handlers for parsing and processing crawled pages.380
hightman/pspiderA parallel web crawler framework built using PHP and MySQLi266