Comment by pbronez

13 hours ago

It’s more economical at a compute level, but not at the developer level. The moment you start customizing your crawler to use protocol X for site Y your scale story collapses.

> It’s more economical at a compute level, but not at the developer level.

Developers and compute are interchangeable now.

  • Who's proompting the machine to do it differently without a developer there to ask the right questions?

    • You can have a high level prompt of: make our crawling cheaper and more reliable to run.

It's a good point, but in practice it depends on how easy those customizations are to implement / maintain, and how much money and effort you save. At some point the compute cost can disrupt even the nicest scale story.

I think the path forward is that websites offer one path for humans, and another for scrapers. But the huge catch is the path for scrapers must be _genuinely_ and _reliably_ the more economical and scalable path (either through something like PoW arms races, or through fear of litigation). Otherwise they will continue to ignore instructions and intrude on the human path.

  • Why aren't we litigating against scrapers, anyway? DDoS is a felony.

    • Largely because they're residential botnets in places like Brazil (a real example from one of my sites that was crawled to near-destruction). Someone could probably do something about this, but it's out of reach for individual site owners.

      3 replies →