Comment by vineyardmike

10 hours ago

The answer is far more benign.

They getting a higher ROI renting their TPUs to Anthropic et al instead of performing training and serving their own models. Google cloud has insane backlog, and has rapidly expanded to satisfy it. While those DCs get built, they’re cannibalizing their own products for it.

This makes sense because (1) they are investors in Anthropic, so they still win and (2) they can always catch up on model training later when the profit opportunity shifts, or abandon it if there is no way to recapture that value.

I think the reason Google hasn't prioritized larger models that are more intelligent than everyone else's, even though they probably could, is because they have 4 billion active users already. For example, they send AI Overviews for a large portion of Google searches now.

So I think they have to prioritize scaling for their models to a higher degree than other groups. Being within say 5% or so in most cases is probably adequate and matters more overall for their user base than being the absolute best coder. So they may be setting compute constraints for training or inference that are firmer than other teams.

  • This is probably true in a smaller way for OpenAI

    When you have a lot of free users the business demands that you serve them with the best cheap model you can build

    And time spent building that may provide dividends (eg OpenAI has very good RL and reasoning) but it might take resources away from the larger model training

    (I have no inside knowledge, so please consider this to all be speculation)

I don’t think so. If Kimi and Deepseek can launch models better than Gemini with much much lesser resources then it is increasingly looking like an organization issue at Google

  • No I think you misunderstood GP’s comment. The idea (which I personally don’t agree with) was that Google didn’t have to have the best models; it just needed to have the best compute infrastructure, i.e. having TPUs and the software stack to use TPUs. It was a better use of money to develop compute infrastructure than to develop better models. Perhaps Gemini itself was resource-starved because Google liked to rent out TPUs to Anthropic instead. (Second-hand information: I heard that Mythos/Fable were trained on Google TPUs.)

    • > Perhaps Gemini itself was resource-starved because Google liked to rent out TPUs to Anthropic instead.

      This was the core hypothesis.

  • I agree that other organizations can compete with few resources, but my hypothesis is that Gemini training specifically is being given nearly 0 resources, despite Google obviously having lots of resources. The hypothesis is based on an assumption that Google profits more by selling ALL their compute to others training models instead of using it themselves for training.

    They already have good models, so “better” isn’t as profitable.

  • Google does not have to compete at the frontier, they already own a lot of Anthropic. It's not an "issue" for them because it's not one of their goals.

I don't think Google is that freaking short-sighted. Even if you're not trying to be frontier, the value of having domain knowledge via experience for AI is worth saving internal compute alone. By all accounts Google believes in ai as much as everyone else.

If AI turns out 1/100 as important as they seem to think, it would be insane to intentionally be slack on it.

Awfully ironic that Google makes more money (probably an order or magnitude or so) from it's direct competitors than it does from a it's own product.