Comment by Imnimo

5 months ago

It's interesting to me that OpenAI considers scraping to be a form of abuse.

It’s funny because the first AI scraper I remember blocking was from OpenAI’s, as it got stuck in a loop somehow and was impacting the performance of a wiki I run. All to violate every clause of the CC BY-NC-SA license of the content it was scraping :)

Quite sure even literal thieves would consider thievery a form of abuse.

  • Engineers working on AI and AI enthusiasts are seemingly incapable of seeing the harm they cause, so I disagree.

    It is difficult to get a man to understand something, when his salary depends on his not understanding it.

  • Yeah, they know it's bad, they just don't think the rules apply to them.

" Integrity at OpenAI .. protect ... abuse like bots, scraping, fraud "

Did you mean to use the word hypocrisy. If not, I'm happy to have said it.

I just want to note, that it is well covered how good the support is for actual malware...

Church, politicians, moralists are all the biggest hypocrites that want to teach you something.

  • I agree on politicians, no idea what a "moralist" is supposed to be but there are good and bad churches and church goers; lumping all church goers into one category calling them hypocrites is wrong. There are many good churches and church goers who help people and their communities.

And have absolutely no reservations about making such an obvious statement on a public forum

I interpreted scraping to mean in the context of this:

> we want to keep free and logged-out access available for more users

I have no doubt that many people see the free ChatGPT access as a convenient target for browser automation to get their own free ChatGPT pseudo-API.

  • > I have no doubt that many people see the free ChatGPT access as a convenient target for browser automation to get their own free ChatGPT pseudo-API.

    Not that hard - ChatGPT itself wrote me a FF extension that opened a websocket to a localhost port, then ChatGPT wrote the Python program to listen on that websocket port, as well as another port for commands.

    Given just a handful of commands implemented in the extension is enough for my bash scripts to open the tab to ChatGPT, target specific elements, like the input, add some text to it, target the relevant chat button, click it, etc.

    I've used it on other pages (mostly for test scripts that don't require me to install the whole jungle just to get a banana, as all the current playright type products do). Too afraid to use it on ChatGPT, Gemini, Claude, etc because if they detect that the browser is being drive by bash scripts they can terminate my account.

    That's an especially high risk for Gemini - I have other google accounts that I won't want to be disabled.

  • This is bad why? Well yeah for openai because all they want it to be is a free teaser to get people hooked and then enshittify.

    Morally I don't see any issues with it really.

[flagged]

  • Very few websites are truly static. Something like a Wordpress website still does a nontrivial amount of compute and DB calls - especially when you don't hit a cache.

    There's also the cost asymmetry to take into account. Running an obscure hobby forum on a $5 / month VPS (or cloud equivalent) is quite doable, having that suddenly balloon to $500 / month is a Really Big Deal. Meanwhile, the LLM company scraping it has hundred of millions of VC funding, they aren't going to notice they are burning a few million because their crappy scraper keeps hammering websites over and over again.

It's not scraping they're concerned about, it's abusing free GPU resources to (anonymously) generate (abusive) content.

Scraping static content from a website at near-zero marginal cost to its server, vs scraping an expensive LLM service provided for free, are different things.

The former relies on fairly controversial ideas about copyright and fair use to qualify as abuse, whereas the latter is direct financial damage – by your own direct competitors no less.

It's fun to poke at a seeming hypocrisy of the big bad, but the similarity in this case is quite superficial.

  • > Scraping static content from a website at near-zero marginal cost to its server, vs scraping an expensive LLM service provided for free, are different things.

    I bet people being fucking DDOSed by AI bots disagree

    Also the fucking ignorance assuming it's "static content" and not something needing code running

    • I think the parent is just pointing out that these things lie on a spectrum. I have a website that consists largely of static content and the (significant) scraping which occurs doesn't impact the site for general users so I don't mind (and means I get good, up to date answers from LLMs on the niche topic my site covers). If it did have an impact on real users, or cost me significant money, I would feel pretty differently.

      9 replies →

    • Also wild that from the tech bro perspective, the cost of journalism is just how much data transfer costs for the finished article. Authors spend their blood, sweat and tears writing and then OpenAI comes to Hoover it up without a care in the world about license, copyright or what constitutes fair use. But don’t you dare scrape their slop.

      5 replies →

    • Off topic, but why is a DoS something considered to act on, often by just shutting down the service altogether? That results in the same DoS just by the operator than due to congestion. Actually it's worse, because now the requests will never actually be responded rather then after some delay. Why is the default not to just don't do anything?

      3 replies →

    • > Also the fucking ignorance assuming it's "static content" and not something needing code running

      Wild eh.

      If it's not ai now, it's by default labelled "static content" and "near-zero marginal cost".

      1 reply →

    • All this reactionary outrage in the comments is funny. And lame.

      Yes, for the vast majority of the internet, serving traffic is near zero marginal cost. Not for LLMs though – those requests are orders of magnitude more expensive.

      This isn't controversial at all, it's a well understood fact, outside of this irrationally angry thread at least. I don't know, maybe you don't understand the economic term "marginal cost", thus not understanding the limited scope of my statement.

      If such DDOSes as you mention were common, such a scraping strategy would not have worked for the scraper at all. But no, they're rare edge cases, from a combination of shoddy scrapers and shoddy website implementations, including the lack of even basic throttling for expensive-to-serve resources.

      The vast majority of websites handle AI traffic fine though, either because they don't have expensive to serve resources, or because they properly protect such resources from abuse.

      If you're an edge case who is harmed by overly aggressive scrapers, take countermeasures. Everyone with that problem should, that's neither new nor controversial.

      8 replies →

  • I understand why OpenAI is trying to reduce its costs, but it simply isn't true that AI crawlers aren't creating very significant load, especially those crawlers that ignore robots.txt and hide their identities. This is direct financial damage and it's particularly hard on nonprofit sites that have been around a long time.

    • > but it simply isn't true that AI crawlers aren't creating very significant load.

      And how much of this is users who are tired of walled gardens and enshitfication. We murdered RSS, API's and the "open web" in the name of profit, and lock in.

      There is a path where "AI" turns into an ouroboros, tech eating itself, before being scaled down to run on end user devices.

    • These are ChatGPT and Claude Desktop crawlers we’re talking about? Or what is it exactly? Are these really creating significant load while not honoring robots.txt?

      Genuinely interested.

      5 replies →

  • That is ridiculous.

    You imply that "an expensive llm service" is harmed by abuse, but, every other service is not? Because their websites are "static" and "near-zero marginal cost"?

    You have no clue what you are talking about.

  • Interesting how other people's cost is "near-zero marginal cost" while yours is "an expensive LLM service". Also, others' rights are "fairly controversial ideas about copyright and fair use" while yours is "direct financial damage". I like how you frame this.

  • Lets not try to qualify the wrongs by picking a metric and evaluating just one side of it. A static website owner could be running with a very small budget and the scraping from bots can bring down their business too. The chances of a static website owner burning through their own life savings are probably higher.

    • If you're truly running a static site, you can run it for free, no matter how much traffic you're getting.

      Github pages is one way, but there are other platforms offering similar services. Static content just isn't that expensive to host.

      THe troubles start when you're actually running something dynamic that pretends to be static, like Wordpress or Mediawiki. You can still reduce costs significantly with CDNs / caching, but many don't bother and then complain.

      3 replies →

  • Have you not seen the multiple posts that have reached the front page of HN with people taking self-hosted Git repos offline or having their personal blogs hammered to hell? Cause if you haven't, they definitely exist and get voted up by the community.

  • The cost is so marginal that many, many websites have been forced to add cloudflare captchas or PoW checks before letting anyone access them, because the server would slow to a crawl from 1000 scrapers hitting it at once otherwise.

  • It's not like those models are expensive because the usefulness that they extracted from scraping others without permission right? You are not even scratching the surface of the hypocrisy

  • It's more ironic because without all the scraping openai has done, there would have been no ChatGPT.

    Also, it's not just the cost of the bandwidth and processing. Information has value too. Otherwise they wouldn't bother scraping it in the first place. They compete directly with the websites featuring their training data and thus they are taking away value from them just as the bots do from ChatGPT.

    In fact the more I think of it, I think it's exactly the same thing.

    • This leads me to thinking: I ask chatGPT a question and they get the answer from gamefaqs.

      But what happens if gamefaqs disappears because of lack of traffic?

      Can LLM actually create or only regurgitate content.

      4 replies →

  • Getting scraped by abusive bots who bring down the website because they overload the DB with unique queries is not marginal. I spent a good half of last year with extra layers of caching, CloudFlare, you name it because our little hobby website kept getting DDoS'd by the bots scraping the web for training data.

    Never in 15 years if running the website did we have such issues, and you can be sure that cache layers were in place already for it to last this long.

  • I don't think a rule along the lines of "Doing $FOO to a corporate is forbidden, but doing $FOO to a charitable initiative is fine" is at all fair.

    What "$FOO" actually is, is irrelevant. I'm curious how you would convince people that this sort of rule is fair.

    The corp can always ban users who break ToS, after all. They don't need any help. The charitable initiative can't actually do that, can they?

  • It is direct financial damage if my servers not on an unmetered connection — after years of bills coming in around $3/mo I got a surprise >$800 bill on a site nobody on earth appears to care about besides AI scrapers.

    It hasn’t even been updated in years so hell if I know why it needs to be fetched constantly and aggressively, - but fuck every single one of these companies now whining about bots scraping and victimizing them, here’s my violin.

  •   > net-zero marginal cost
    

    Lol, you single-handedly created a market for Anubis, and in the past 3 years the cloudflare captchas have multiplied by at least 10-fold, now they are even on websites that were very vocal against it. Many websites are still drowning - gnu family regularly only accessible through wayback machine.

    Spare me your tears.

  • > Scraping static content from a website at near-zero marginal cost to its server

    It's not possible to know in advance what is static and what is not. I have some rather stubborn bots make several requests per second to my server, completely ignoring robots.txt and rel="nofollow", using residential IPs and browser user-agents. It's just a mild annoyance for me, although I did try to block them, but I can imagine it might be a real problem for some people.

    I'm not against my website getting scraped, I believe being able to do that is an important part what the web is, but please have some decency.

  • AI providers also claim to have small marginal costs. The costs of token is supposedly based on pricing in model training, so not that different from eg your server costs being low but the content production costs being high. And in many cases AI companies are direct competitors (artists, musicians etc.)

    (TBH it's not clear to me that their marginal costs are low. They seem to pick based on narrative.)

  • My website serving git that only works from Plan 9 is serving about a terabyte of web traffic monthly. Each page load is about 10 to 30 kilobytes. Do you think there's enough organic, non-scraper interest in the site that scrapers are a near-zero part of the cost?

  • Absolutely not, the former relies on controversial ideas to qualify as legal.

    Stealing the content from the whole planet & actively reducing the incentive to visit the sites without financial restitution is pretty bad.

  • You are, of course, ignoring the production costs of the static content that OpenAi is stealing.

    Stop justifying their anti-social behavior because it lines your pockets.

  • Because you say it is?

    I obviously disagree. I mean, on top of this we are talking about not-open OpenAI.

  • It’s not for techbros to decide at what threshold of theft it’s actually theft. “My GPU time is more valuable than your CPU time” isn’t a thing and Wikipedias latest numbers on scraping show that marginal costs at scale are a valid concern

  • I'm sure the copyright holders would consider your use of their content as direct financial damage

  • The issue is that there are so many awful webmasters that have websites that take hundreds of milliseconds to generate and are brought down by a couple requests a second.

    • OpenAI must be the most awful webmasters of all, then, to need such sophisticated protections.