crawler

Crawler

A high-performance web crawling and scraping solution with customizable settings and worker pooling.

A high performance web crawler / scraper in Elixir.

GitHub

945 stars
32 watching
91 forks
Language: Elixir
last commit: over 2 years ago
Linked from 1 awesome list

crawlerelixirfilesofflinescraperscraper-enginespider

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
elixir-crawly/crawlyA framework for extracting structured data from websites994
feng19/spider_manA high-level web crawling and scraping framework for Elixir.23
fmpwizard/owlcrawlerA distributed web crawler that coordinates crawling tasks across multiple worker processes using a message bus.55
spider-rs/spiderA tool for web data extraction and processing using Rust1,234
webrecorder/browsertrix-crawlerA containerized browser-based crawler system for capturing web content in a high-fidelity and customizable manner.677
hu17889/go_spiderA modular, concurrent web crawler framework written in Go.1,827
vida-nyu/acheA web crawler designed to efficiently collect and prioritize relevant content from the web459
turnersoftware/infinitycrawlerA web crawling library for .NET that allows customizable crawling and throttling of websites.248
elliotgao2/gainA Python web crawling framework utilizing asyncio and aiohttp for efficient data extraction from websites.2,037
puerkitobio/gocrawlA concurrent web crawler written in Go that allows flexible and polite crawling of websites.2,036
chenjiandongx/github-spiderA Python-based web crawler for scraping Github user and repository data.264
a11ywatch/crawlerPerforms web page crawling at high performance.51
puerkitobio/fetchbotA flexible web crawler that follows robots.txt policies and crawl delays.787
antchfx/antchA framework for building fast and efficient web crawlers and scrapers in Go.261
helgeho/web2warcA Web crawler that creates custom archives in WARC/CDX format25