Comment by holografix
2 days ago
What’s the end goal for Cloudflare and the web here? I don’t think ADOG (anthropic, deepmind, openai, google) is going to pay to crawl.
What would force their hand?
It’s more likely they’ll strike undisclosed agreements with major sources of discussion like reddit etc.
That’s not to say getting new information as a way of context-providing is not going to happen but that’s not scraping.
I think the best answer is, nobody knows. The previous equilibrium for content scraping for search engines on the internet was already at times an uncomfortable one. But I agree that from a game theory perspective, "the AI bots take and give nothing back in return" is not just hyperbole, it's the actual situation. If Google is successful in what seems to be its plans and it becomes a box where you type a question and Google gives you an answer and only a vanishing fraction of the users click through to any underlying website, that instantly eliminates the entire value proposition for vast swathes of the web to actually be on the web.
Something has to happen or Google will end up starved and locked out of everything, by means both technical and legal. Then nobody gets anything.
I don't have the answer as to what happens next, and I doubt anyone else who proclaims one super confidently. But we can do some constraints analysis. There is no world where everyone works for free so Google and other AI engines can get all the value from the content, so we can eliminate those possibilities. I think we can safely discard the world(s) in which all content production just stops. However, off the top of my head, it's hard to get much tighter than that, and that definitely leaves a world where effectively everything everywhere ends up going pay-to-access.
Microtransactions have, to date, failed comprehensively, though, so the constraints on what "everything is pay-to-access" gets weird without them.
And there is never guarantee that there is any solution to any set of constraints. Things can end up overconstrained in reality as easily as a math problem. I don't actually think it'll go that way, but when analyzing this question I think it's important to not let "but $SOMETHING just has to have some way to work, because... uh... it has to!" Let the constraints do the talking. You could end up with a scenario where all content of any value is locked down, and it's fundamentally difficult and expensive to ever access or discover it, and consequently the entire content production industry radically contracts compared to its current size, if there is no pragmatic solution to microtransactions that is low-enough friction to get over the psychological and economic hurdles that have killed it to date. If everything is locked behind "macrotransactions" that's a much smaller commercial web. Probably a much higher quality one, too, but at a pretty stiff cost.
The internet becomes full of free propaganda since it's not the consumer who pays for that?
That's the current state of the internet. In this world it wouldn't be "the Internet" full of propaganda, it would be the AI search engines the propaganda would get concentrated into. For "national security", don't you know. And it would be much easier for them for having an even smaller target. At least the current internet lets you cross-check one source of free propaganda against another today. (Although whether the truth is in any meaningful way "between" any set of them is another question.)
One possible scenario is that that is simply it for the internet as an information source; the search engine's AIs get captured and there becomes effectively no way to discover any of the content already on there.
But then again, people will react to that and do something. Kagi would grow and others too. The more interesting question is whether the governments that captured Google's AI would let them or if suddenly it would be discovered that copyright law doesn't permit search engines to do that. Would that be inconsistent with letting the Approved AIs access whatever they want and chew on it even harder? Yeah, and they wouldn't care.
I get a mixed feeling about all this. Cloudflare is unilaterally making all these decisions which impact the whole internet traffic flow. Taking the lead is one thing, however decisions like this should have the direct involvement of Internet Engineering Task Force (IETF) to account for all stakeholders, otherwise we run into the situation of a fragmented internet
Are they unilateral? Maybe the defaults? For anyone competent they're settings freely chosen.
I'm curious if this is just to pressure Google into separating their crawlers
This is more of a way to create a unauthorized toll tax on highway. There is a problem indeed, however if the proposed direction by cloud flare is to solve it or benefit out of it is a debatable topic.
That wouldn't make any sense from anyone's perspective. They would just pretend to separate them and not.
Universal tax collector of the internet. A penny for every page access. ADOG will not mind as it cements their incumbent status and pulls up the drawbridge by erecting a huge financial barrier for any new entrant.
"I don’t think ADOG (anthropic, deepmind, openai, google) is going to pay to crawl."
You seem to misunderstand. You pay or you die. There's nothing in between. Cloudflare will happily collect the tax. As does Apple (collecting 20B$ yearly from Google for the "tax"). Cloudflare's users also won't mind about how the company handles ADOG as long as they get a chunk of the cake by getting freebies and cheap services.
Pay to crawl is already here.
The usage patterns of how people pay and use AI is basically the same model the web should be using: you pay a small bit of money to access monetized pages, just how you pay a small bit of money to get AI responses.
It just needs people and browsers to get onboard with protocols. Crawlers will have no choice but to pay for content behind these 402 gateways.
*AGOD