Comment by sreekanth850

8 hours ago

And you think all this web AI crawlers will respect robots.txt. That era is gone.

The big USA ones do, and it would be madness for them to do otherwise.

But that's meaningless because 99% of AI crawlers are "bad bots" which ignore robots.txt and use domestic IPs to circumvent blocks.

  • That doesn't match my experience. For example, I recently had Meta scrape robots.txt disallowed paths. On some page they claimed they respect it but they don't. And yes, all the IPs they used to connect to me (they used a different /64 network for each connection) were owned by Meta, it was not someone pretending to be them.

    Some other reports of this: https://github.com/TecharoHQ/anubis/issues/1565

    I ended up just fully closing connections with no response from these assholes on any URL.