← Back to context

Comment by pantsforbirds

18 hours ago

We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice!

Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021 when we were doing this.

What if they charged money? Is it something you'd pay for?

  • I was thinking the same thing... paid access for high volume users or scrapers could actually help fund the non-profit. Maybe let website owners decide which scrapers are allowed to use their content, or allow them to get paid for use of it. If news and other sites were getting paid, maybe they could go back to optimizing for good content instead of clicks.

  • I mean we already paid for the article from the source itself. I guess I'd expect a better "diff" source from them, but if they dont even update the article itself, i guess i wouldn't expect a paid service to have those updates either?

    • ah, i think i misunderstood your original post. if you mean the wayback-machine/arkive, then I suspect it'd be hard to justify? You are essentially paying a third-party source to validate that diffs didn't go through on the source material.

      with llms, at some point it probably becomes easier to use your paid api connection to manage your own cached version yourself?

  • I wonder if there would be concern on their part about appearing to be a company that was basically offering paywall circumvention as a product.

    • It wouldnt be a paywall, more like an option for companies to not pay scrappers. At least the payment deviates to the source.