Open a page saved by Safari (.webarchive) or a WARC file on any computer
Drop a .webarchive someone sent you from a Mac, iPhone or iPad, or a .warc or .warc.gz from the Internet Archive, wget or Webrecorder. The saved page is drawn here from the file alone, every file inside is listed and can be saved, and a .webarchive can be turned into one .html file that any browser opens.
The file is read by this tab on your device and is not uploaded. The saved page cannot reach the internet from here: its scripts are removed, and anything it would load from elsewhere is left out and listed.
What it shows for a .webarchive
- The page as it was saved. The HTML Safari stored, drawn with the pictures, stylesheets, fonts and frames stored beside it. A stylesheet's own
url()pictures and@imports are taken from the archive too. - Where it came from. The page's address, its title, its type and text encoding, and the date the web server gave when Safari fetched it (from the HTTP response Safari keeps inside the file). The file's own last-modified date is shown too: when it was saved, or, for a copy that was downloaded or mailed, when that copy was made.
- Every file inside. A table of each resource with its address, MIME type and size, and a Save button for each, so a picture or a PDF that was on the page can be taken out on its own. Files the page never uses are marked.
- Frames. Each
<iframe>Safari saved is a small archive inside the archive. It is drawn in its place on the page, and listed so it can be shown alone or saved. - Save page as one .html. The page with every file it uses built in (as
data:addresses), scripts removed. It opens in Chrome, Edge, Firefox or a phone browser, with no folder beside it, and loads nothing from the internet. Save the HTML as stored gives the original markup untouched. - What is missing. Every picture, stylesheet, font or frame the page refers to that is not in the file, with its address. Safari only saves what had loaded, so pictures a site loads late (as you scroll) are often among them.
What it shows for a WARC or WARC.gz
- The file. WARC version, packed and unpacked size, how many records of each type, the date range of the captures, and the fields of the
warcinforecord (the software that wrote it, the operator, the crawl it belongs to). - The web pages in it. Every HTML page captured with status 200, with its title, date and size. Show page draws it from this file alone: its stylesheets and pictures come from other records of the same file, and a redirect recorded in the file (a 301 to a new address) is followed.
- Every record, filtered. Number, type (warcinfo, request, response, resource, metadata, revisit, conversion), target address, date, content type and size. Type in the filter to keep the records whose address, type, date or content type contains every word, or pick one record type.
- Inside each record. The WARC header, the HTTP status line and headers, and the body: a page drawn, a picture shown, text shown (the first 200 KB), or the body saved as a file. A chunked body is put back together and a gzip or deflate
Content-Encodingis undone first. Arevisitrecord, which stores no body because it was identical to an earlier capture, shows the earlier record's body and says which one.
Why a saved page cannot phone home from here
A saved page still holds the addresses of everything it loaded: trackers, ad servers, fonts, analytics. Opened carelessly, it fetches them again and tells those servers when and where it was opened. Here three things stop that:
- the page is rebuilt before it is shown: scripts, event handlers (
onload,onerror…),<base>, refresh and preconnect tags, plug-ins (<object>,<embed>) and form targets are removed; everysrc,srcset,poster, stylesheet link, CSSurl(),@importandimage-set()either points to a file from the archive or is removed; links keep their text but are not clickable, and show where they pointed when you hover; - the rebuilt page starts with its own policy,
default-src 'none'; img-src blob: data:; style-src blob: data: 'unsafe-inline'; font-src blob: data:; media-src blob: data:, so even an address the rebuild missed is refused by the browser; - it is shown in a frame whose
sandboxgrants nothing: no scripts, no access to this site, no forms, no pop-ups.
viewhack's own test drops a page full of absolute http:// and https:// pictures, stylesheets, fonts, scripts and a frame into this page in a headless browser, and checks that not one request leaves the tab.
The two formats
- .webarchive (Safari)
- An Apple property list, usually binary (it starts with
bplist00), whose top dictionary holdsWebMainResource(the page),WebSubresources(a list of files) andWebSubframeArchives(one archive per frame). Each file is stored with its bytes, address, MIME type and text encoding. Safari on a Mac writes it with File, Save As, Format: Web Archive; on an iPhone or iPad, Share, Options, Web Archive, then Save to Files. Only Safari opens it, so on Windows, Linux or Android it usually arrives as a file nothing will open. An XML copy made withplutil -convert xml1opens here too. - WARC (.warc, .warc.gz)
- The web archiving standard, ISO 28500 (WARC 1.0 in 2009, 1.1 in 2017), used by the Internet Archive's Heritrix crawler,
wget --warc-file, Webrecorder and ArchiveWeb.page, browsertrix-crawler and the warcio library. A WARC is a series of records, each aWARC/1.1line, named fields such asWARC-Type,WARC-Target-URIandContent-Length, an empty line, then exactly that many bytes. A .warc.gz compresses each record as its own gzip member, so tools can jump to one record; this page unpacks all of them. - A worked example
- The file
example.warc.gzfrom the warcio project's tests is 3,816 bytes and holds 6 records: twowarcinforecords written by Webrecorder Platform v3.7 in March 2017, aresponseforhttp://example.com/whose 606-byte gzip body unpacks to 1,270 bytes of HTML titled “Example Domain”, therequestthat fetched it, and arevisitof the same page whose body was not stored again because its digest matched; here it shows the first capture's body.
What this cannot do
- Run the page's scripts. The page is drawn with scripts off. A site that builds its content with JavaScript (many shops, social networks and web apps) may look empty or half-built; what Safari or the crawler stored is what you get. Menus, sliders and comments that need scripts do not work.
- Fetch anything that is not in the file. A picture or stylesheet the archive did not store stays missing; it is listed with its address, and never loaded, even if it is still online.
- Replay video streams. Sites such as YouTube send video in many small pieces chosen by a script; crawlers rarely store them all, and without the player script they cannot be put together here. A plain video or audio file stored in a WARC can be saved and played elsewhere.
- Follow links between archived pages. Links are shown as text with their address. To open another page of the same WARC, pick it from the list of pages or records.
- Read every saved-page format. Chrome and Edge's “Webpage, Single File” (.mhtml, .mht), the older Internet Archive .arc format and WACZ packages are not read. A Brotli-compressed body (
Content-Encoding: br) is shown and saved as stored, not unpacked. - Open huge archives on a phone. The whole file is read into the tab's memory, and a .warc.gz is unpacked whole. A few hundred megabytes work on a laptop; a multi-gigabyte crawl does not.