How to archive a website in a future-proof way (involves PDF hybrid)

evenwicht@lemmy.sdf.org · edit-2 4 days ago

How to archive a website in a future-proof way (involves PDF hybrid)

smpl@discuss.tchncs.de · 4 days ago

I archive with the WebScrapBook extension to htz which is also a zip based format. I open it with firefox’s built in method to open jar files like this: jar:file://${file}!/index.html

The fact that it is a zip based format makes it futureproof. I could as well unzip it and open it with any browser. You could probably open the maff files the same way, but I don’t think it is standardized to have an index.html file like htz.

AFAIK WARC is the only standard way of archiving webpages.

evenwicht@lemmy.sdf.org · edit-2 4 days ago

IIUC you are referring to this extension, which is Firefox-only (~~like~~unlike the save page WE, which has a Chromium version).

Indeed the beauty of ZIP is stability. But the contents are not. HTML changes so rapidly, I bet if I unzip an old MAFF file it would not have stood the test of time well. That’s why I like the PDF wrapper. Nonetheless, this WebScrapBook could stand in place of the MHTML from the save page WE extension. In fact, save page WE usually fails to save all objects for some reason. So WebScrapBook is probably more complete.

(edit) Apparently webscrapbook gives a choice between htz and maff. I like that it timestamps the content, which is a good idea for archived docs.

(edit2) Do you know what happens with JavaScript? I think JS can be quite disruptive to archival. If webscrapbook saves the JS, it’s saving an app, in effect, and that language changes. The JS also may depend on being able to access the web, which makes a shitshow of archival because obviously you must be online and all the same external URLs must still be reachable. OTOH, saving the JS is probably desirable if doing the hybrid PDF save because the PDF version would always contain the static result, not the JS. Yet the JS could still be useful to have a copy of.

(edit3) I installed webscrapbook but it had no effect. Right-clicking does not give any new functions.

smpl@discuss.tchncs.de · 4 days ago

It saves the rendered page. It also has a built-in rough DOM editor so that you can edit the document before saving. The way I have it set up it up is to remove all javascript from pages.

evenwicht@lemmy.sdf.org · 4 days ago

In principle the ideal archive would contain the JavaScript for forensic (and similar) use cases, as there is both a document (HTML) and an app (JS) involved. But then we would want the choice whether to run the app (or at least inspect it), while also having the option to offline faithfully restore the original rendering. You seem to imply that saving JS is an option. I wonder if you choose to save the JS, does it then save the stock skeleton of the HTML, or the result in that case?

smpl@discuss.tchncs.de · 2 days ago

I just want to archive content, but if one would want such a thing I think the format to go for is WARC. As far as I understand it record the network requests, but I’m not sure as I’ve got no experience with the format at all. That way it would be possible to keep the original javascript intact and with a little manipulation of network requests to replay the original site.

How to archive a website in a future-proof way (involves PDF hybrid)

How to archive a website in a future-proof way (involves PDF hybrid)

MAFF (a shit-show, unsustained)

MHTML (shit-show due to non-portable browser-dependency)

PDF (lossy)

PDF+MHTML hybrid

We need to evolve

(update) The goals