httrack2warc

WARC crawler

Converts HTTrack crawls to WARC files by reconstructing requests and responses from logs

Converts HTTrack crawls to WARC files

GitHub

32 stars
20 watching
6 forks
Language: Java
last commit: about 2 years ago
Linked from 1 awesome list

web-archiving

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
webrecorder/har2warcConverts HTTP Archive format to Web Archive format48
archiveteam/grab-siteA web crawler designed to backup websites by recursively crawling and writing WARC files.1,406
iipc/warc2htmlConverts WARC files to static HTML with relative link rewriting and renaming41
internetarchive/warctoolsTools for working with archived web content153
iipc/jwarcA Java library for reading and writing WARC files with a typed API48
n0tan3rd/node-warcA tool for parsing and generating Web Archive files in JavaScript using Node.js95
helgeho/web2warcA Web crawler that creates custom archives in WARC/CDX format25
webrecorder/warcioA fast streaming library for working with WARC format web archival data391
chfoo/warcatTool for handling Web Archive files152
steffenfritz/html2warcConverts offline data into a standard archival format18
helgeho/warcpartitionerTool for partitioning and merging Web archive files by MIME type and year1
ukwa/webarchive-discoveryTools for indexing and discovering archived web content117
turicas/crauA command-line tool for archiving and playing back websites in WARC format59
nlnwa/gowarcserverA tool for indexing and serving contents of WARC files.15
ikreymer/webarchive-indexingTools for bulk indexing of WARC/ARC files to create a shared url index43