Comment by vkou
5 months ago
> Hey! I'm Nick, and I work on Integrity at OpenAI. These checks are part of how we protect our first-party products from abuse like bots, scraping, fraud, and other attempts to misuse the platform.
How can first-party products protect themselves from abuse by OpenAI's bots and scraping?
This is a completely in-scope question.
How do we defend against your scraping, OpenAI?
I dont want any of my content scraped or seen by you all. Frankly, fuck you all for thinking my content is owned by you.
I use nginx conditionals and useragent checking, then respond with 418 or 410.
Probably too late now but my list needs updating
It's documented here: https://developers.openai.com/api/docs/bots
- which one is to stop you from hammering small servers with hundreds simultaneous connections?
- why don't you just respect existing robots.txt that apply to you already?
- does every LLM scraper seriously think the onus to opt out from the EVERY SINGLE SCRAPER is on the webmasters/owners?
1 reply →
robots.txt bro https://developers.openai.com/api/docs/bots/
"bro": https://www.businessinsider.com/openai-anthropic-ai-ignore-r...
7 replies →