Comment by toomuchtodo

18 hours ago

It is. They will most likely eventually need to move to a walled model for Wayback due to scraper aggressiveness (like Reddit deprecating anonymous old.reddit.com), or behind Cloudflare for aggressive bot and scraping protection. Hard to defend against abuse of a public resource when its intent is public access with as little restriction as possible.

https://en.wikipedia.org/wiki/Tragedy_of_the_commons

(no affiliation)

Reddit has no excuses for the anonymous old.reddit.com removal; they're simply greedy.

On the other hand, the Internet Archive is a non-profit offering a free public resource.

  • Examples provided as technical examples, strong feelings are out of scope for this thread.

    • As someone who operates a large non-profit public data driven website, I have some VERY strong feelings about scrapers. We looked into various commercial solutions (Datadome, HUMAN) and based on our traffic estimates from logs we'd be looking at at least 250k/yr for bot mitigation. Anubis is offering a temporary reprieve, but after reading the recent kernel.org article [0] it's increasingly clear that this is a temporary bandaid.

      The cheapest solution is to require a login and rate limit by API key. I also have strong feelings about the tragedy of the commons.

      [0] https://people.kernel.org/monsieuricon/creepy-crawlies

      2 replies →