grab-site

Web crawler

A web crawler designed to backup websites by recursively crawling and writing WARC files.

The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns

GitHub

1k stars
41 watching
136 forks
Language: Python
last commit: about 2 years ago
Linked from 1 awesome list

archivingcrawlcrawlerspiderwarc

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
archiveteam/wpullDownloads and crawls web pages, allowing for the archiving of websites.556
webrecorder/browsertrix-crawlerA containerized browser-based crawler system for capturing web content in a high-fidelity and customizable manner.677
helgeho/web2warcA Web crawler that creates custom archives in WARC/CDX format25
nla/httrack2warcConverts HTTrack crawls to WARC files by reconstructing requests and responses from logs32
peterk/warcworkerA web archiving tool that archives websites with high-fidelity preservation capabilities.57
n0tan3rd/squidwarcAn archival crawler built on top of Chrome or Chromium to preserve the web in high fidelity and user scriptable manner170
internetarchive/brozzlerA distributed web crawler that fetches and extracts links from websites using a real browser.678
turicas/crauA command-line tool for archiving and playing back websites in WARC format59
internetarchive/warctoolsTools for working with archived web content153
internetarchive/warcproxAn HTTP proxy designed to capture and archive web traffic, including encrypted HTTPS connections.389
vida-nyu/acheA web crawler designed to efficiently collect and prioritize relevant content from the web459
cocrawler/cocrawlerA versatile web crawler built with modern tools and concurrency to handle various crawl tasks188
chfoo/warcatTool for handling Web Archive files152
a11ywatch/crawlerPerforms web page crawling at high performance.51
spider-rs/spiderA tool for web data extraction and processing using Rust1,234