Comment by genxy
1 month ago
I came here to mention https://github.com/gildas-lormeau/singlefile, I couldn't get it to do what I want so I built my own, for watched domains it pushes a copy of the serialized DOM to a local search database. For structured data, either extract and enrich in the browser or enrich on the server side.
Would Hister support this basic workflow? I'd love to retire my own software.
The next phase was going to move to a recording proxy.
Singlefile based web pack+download is exactly the next feature I need besides indexing URLs / browser tabs with[0], have it on my to-research for a looong time. I do have a server-side script that can be triggered on doc ingestion as a hook "if data/schema/tab is indexed and path is /ml/* then run script "fetch-url.sh" and file the result into {{doc.path}}" but being able to do that in-browser is more user-friendly.
[0] https://github.com/canvas-ui/canvas/tree/main/apps/browser-e...