viewhack

Open a page saved by Safari (.webarchive) or a WARC file on any computer

Drop a .webarchive someone sent you from a Mac, iPhone or iPad, or a .warc or .warc.gz from the Internet Archive, wget or Webrecorder. The saved page is drawn here from the file alone, every file inside is listed and can be saved, and a .webarchive can be turned into one .html file that any browser opens.

The file is read by this tab on your device and is not uploaded. The saved page cannot reach the internet from here: its scripts are removed, and anything it would load from elsewhere is left out and listed.

What it shows for a .webarchive

What it shows for a WARC or WARC.gz

Why a saved page cannot phone home from here

A saved page still holds the addresses of everything it loaded: trackers, ad servers, fonts, analytics. Opened carelessly, it fetches them again and tells those servers when and where it was opened. Here three things stop that:

viewhack's own test drops a page full of absolute http:// and https:// pictures, stylesheets, fonts, scripts and a frame into this page in a headless browser, and checks that not one request leaves the tab.

The two formats

.webarchive (Safari)
An Apple property list, usually binary (it starts with bplist00), whose top dictionary holds WebMainResource (the page), WebSubresources (a list of files) and WebSubframeArchives (one archive per frame). Each file is stored with its bytes, address, MIME type and text encoding. Safari on a Mac writes it with File, Save As, Format: Web Archive; on an iPhone or iPad, Share, Options, Web Archive, then Save to Files. Only Safari opens it, so on Windows, Linux or Android it usually arrives as a file nothing will open. An XML copy made with plutil -convert xml1 opens here too.
WARC (.warc, .warc.gz)
The web archiving standard, ISO 28500 (WARC 1.0 in 2009, 1.1 in 2017), used by the Internet Archive's Heritrix crawler, wget --warc-file, Webrecorder and ArchiveWeb.page, browsertrix-crawler and the warcio library. A WARC is a series of records, each a WARC/1.1 line, named fields such as WARC-Type, WARC-Target-URI and Content-Length, an empty line, then exactly that many bytes. A .warc.gz compresses each record as its own gzip member, so tools can jump to one record; this page unpacks all of them.
A worked example
The file example.warc.gz from the warcio project's tests is 3,816 bytes and holds 6 records: two warcinfo records written by Webrecorder Platform v3.7 in March 2017, a response for http://example.com/ whose 606-byte gzip body unpacks to 1,270 bytes of HTML titled “Example Domain”, the request that fetched it, and a revisit of the same page whose body was not stored again because its digest matched; here it shows the first capture's body.

What this cannot do