Comment by doginasuit

8 hours ago

Let's not conflate crawlers with the traffic that bot protection services block. A crawler that respects robots.txt is a good internet citizen and can provide a vital service.

However, so-called AI crawlers are not the same as crawlers of yore. They hit live pages every time a user prompt triggers a web search.

This Web Search API, unlike an AI crawler, only fetches periodically. It feels like a step in the right direction for managing resource strain across the internet. If only the LLM giants could do something similar.

A crawler that respects robots.txt is useless in practice since many sites only allow Googlebot and maaaybe Bing - by name.

And you think all this web AI crawlers will respect robots.txt. That era is gone.

  • The big USA ones do, and it would be madness for them to do otherwise.

    But that's meaningless because 99% of AI crawlers are "bad bots" which ignore robots.txt and use domestic IPs to circumvent blocks.

    • That doesn't match my experience. For example, I recently had Meta scrape robots.txt disallowed paths. On some page they claimed they respect it but they don't. And yes, all the IPs they used to connect to me (they used a different /64 network for each connection) were owned by Meta, it was not someone pretending to be them.

      Some other reports of this: https://github.com/TecharoHQ/anubis/issues/1565

      I ended up just fully closing connections with no response from these assholes on any URL.