Comment by awongh
10 hours ago
It seems like there are some credible rumors that Google is actually winning in terms of actually building models that work and don't lose money- between how they're able to price them, the TPU advantage and their capex advantage (being able to raise debt + just having a lot of cash - well I said not lose money... more like not go bankrupt).
From the outside they look like they're behind in terms of frontier models, but I think they might be the best positioned to not go out of business when the bubble pops.
Also look at the fact that they've been able to deploy AI-assisted search at google scale. It must be another order of magnitude larger (at least) than the model deployments for OpenAI and Anthropic.
Of course unless you're inside Google it's impossible to know for sure.
In terms of open models, Gemma 4 beats the pants off everything else to the point that paying for APIs becomes hard to justify. Qwen has the meme-share for coding, but it feels much less well rounded. I have no doubt that Google have both the infrastructure and the expertise to curb stomp everyone else, should they resolve in earnest to do so.
Lest we forget, "Attention is All You Need" came from Google.
> "Attention is All You Need" came from Google
It also came directly from the university of Toronto, and the university of Toronto seeded all American frontier labs (including Grok (why do you think they could start so fast))
Interesting, glad to hear. We have gemma4 at work, and I was considering localhosting qwen, but gemma4 is so far behind the Opus and Fable I have at home that I've decided to hold off for another model release.
"We have a company provided Toyota at work but it is so far behind the Ferrari I rent at home that I've decided to hold off for another model release."
Are you suggesting Gemma beats GLM 5.2?
At 20x the parameter count I should hope GLM beats Gemma! But is it 20x better? Expertise is demonstrated, not by making big models, but by making small ones. Bigger isn't better if you can't run it at all.
How long until Gemma 5 hits?
It's rumored that Gemini 3.5 flash has a >50% margin, and I'd imagine 3.6 flash is even higher.
I do not think OpenAI or Anthropic are actively chasing margins - though, Anthropic is supposed to be profitable on some form of non-GAAP accounting...
I suspect Google isn't really interested in seeing how far it can get dragged into a race of selling dollars for $0.25, and is more interested to see if it can stay in the race selling $0.50 for a dollar - when everyone else is losing or barely breaking even.
It kind of doesn't make sense though, because typically a large org like Google can afford to crush competitors on pricing. They could probably even go toe to toe with chinese model pricing for years without feeling it.
Maybe they don't want to price war with the other labs so they can comfortably maintain healthy margins on selling them compute?
Maybe they are hoping that when the bottom drops out they will just be able to buy Anthropic or OpenAI for a few tens of billion.
google has to make money. flash is awesome. you can run it free on their infra and the performance and latency is excellent for what you wait and pay for right now, with great perf per watt. every person in the world going to google.com runs it. every query. its far larger than free gpt, localhost qween and what not.
it's their pro that isn't awesome at all. in fact, their pro kinda suck now that everyone else woke up.
That "TPU advantage" might be slowing Google down (though likely not as much as their internal bureaucracy).
Porting CUDA-based research, debugging, and overall experimentation speed is likely slower.
The GPU is still king for training.
But maybe the TPU advantage is in inference? That's what I assume because the number of compute cycles are going to be all in inference vs training. So they could train on GPUs if they want.
lmao, you know all Anthropic models are trained on TPU right?
thats funny because my company sells them nvidia gpu for training. but im happy for the billions, they prolly use them for counterstrike!
They basically don't exist in the currently most profitable LLM market (coding).
Yes, subs like codex are heavily subsidized. But API billing has massive margins and that's what enterprises pay.
On the other hand, the consumer side of the market seems to be less competitive right now.
OpenAI's new Mac app doesn't even have a normal "Chat" option now. OpenAI might be chasing coding and b2b sales more now that they realise very few regular consumers pay for subscriptions.
Does it have "massive" margins? Afaik no one has said publicly what margins there are on an API call?
"As of October [2025], OpenAI's compute margins reached 70%, up from 52% at the end of 2024 and double the rate in January 2024, [The Information] said, citing a person familiar with the figures."
https://www.bloomberg.com/news/articles/2025-12-21/openai-se...
As for Anthropic, the rumors I remember seeing for their API margins were more like 85-90%, but I don't have a reference at hand for those. But once you know the API is wildly profitable and the subscriptions are roughly break-even and not even a big slice of their income, all of the investment makes a lot more sense.
1 reply →