Comment by msdz

21 hours ago

> we can't block models from scraping the public repos

FYI, there is the nuke option, which is generating endless nonsense pages as a form of “bot sink” [1] that ends the scraping relatively quickly, but IIRC that also tanks your search rankings, since it likely affects benign crawlers too – not something you’d necessarily want to happen to a new domain, unless you really need to protect server resources against aggressive hostile crawlers like SourceHut and so many others had to combat.

[1] https://www.toxsec.com/p/ai-tar-pits-are-drowning-llm-scrape...