wasp

Web archiver

A containerized web archive and search system using Elastic Search

GitHub

27 stars
13 watching
4 forks
Language: Java
last commit: almost 4 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
sul-dlss/wasapi-downloaderAn application to download archives of web archiving projects6
ukwa/webarchive-discoveryTools for indexing and discovering archived web content117
webrecorder/archiveweb.pageA high-fidelity web archiving system for storing and replaying interactive web pages in browsers.903
derfenix/webarchiveA web-based archive service that allows users to store and manage web pages in various formats.115
webrecorder/pywbA toolkit for archiving and replaying web content accurately and efficiently1,418
vida-nyu/acheA web crawler designed to efficiently collect and prioritize relevant content from the web459
oduwsdl/ipwbA system for dispersing and replaying archived web content using peer-to-peer technology.617
jarofghosts/memento-clientProvides a simple JavaScript interface to access historical web pages via the Wayback Machine14
internetarchive/archA distributed compute analysis system for web archive collections15
florents-tselai/warcdbA library for storing and querying web crawl data in a compact, easily sharable format.397
archiveteam/grab-siteA web crawler designed to backup websites by recursively crawling and writing WARC files.1,406
ikreymer/webarchive-indexingTools for bulk indexing of WARC/ARC files to create a shared url index43
elastic/elasticsearchA distributed search and analytics engine for scalable data storage and real-time search capabilities71,007
netarchivesuite/jwatA toolkit for analyzing and extracting data from legacy web archives in a structured format suitable for further analysis or reuse3
stevepolitodesign/my_site_archiveA simple Rails application for archiving websites27