← Back to context

Comment by scotty79

9 hours ago

I think the moat is going to be compute. So far compute needed to push the frontier is still extremely cheap so the capital can afford to spread its bets. But when further improvement is going to cost in trillions, capital will have to pick a winner and bet only on him. It won't be a matter of finding the best bet, it will be a matter of survival.

This will cause the picked winner to get massively ahead with sheer compute alone used both for training and inference dedicated to recursive self improvement.

I guess all predictions age like milk, but here's one:

There's a law of diminishing returns at play here, and doubling the energy cost of training to wring 2% more performance out of the technology isn't going to be very useful, because most of the problems it is capable of solving will be solvable with the previous-gen 98%-as-good model.

("there's a law of diminishing returns at play here" is an article of faith. But then, so is the belief that these models will keep getting better).

  • As soon as you can demonstrate decent financial returns (ie. the AI can run a company better than humans can), suddenly it makes sense to put a lot more $$$ in even if returns are diminishing - since whoever runs companies the best gets control of a big chunk of the world economy.

    • If that works... why not just jump straight to a planned economy run by LLM? Skip the whole messy "free market" thing altogether?

      (I don't think it will work).

      2 replies →

Surely there is a point where algorithmic improvements will be more cost effective than buying more hardware.

  • If you're actually applying LLMs, all of the things around the LLM that adapt it to coding, for example, that enable it to use existing validation tools for code, and enable it to diagnose and fix tool chain issues that aren't directly coding problems, are what makes the difference between a model that that scores a little higher on a coding benchmark and a model that's useful in a particular code base on a particular platform.

    Are there any use cases that have enabled one customer of a frontier LLM to outperform a competitor using a different frontier LLM? Or is this why we are seeing confected points of comparison like solving challenge problems in mathematics?

  • I'm afraid it might be the other way around. RSI might pick all of the low hanging fruit soon. There must be a physical limit of how much intelligence you can squeeze out of some amount of parameters and compute.

    There are going to still be worthwhile improvements but they are going to be more like not how to make transformers 10x cheaper but how to make next training run cost 9 trillions instead of 10 with a very particular optimization designed at the cost of hundreds of millions for this one specific run.

    • It is just occurring to me that “RSI” expands to recursive self improvement. Thought people were talking about repetitive stress injuries; either in regards to programmers writing too much code/not having to write code anymore, or the frontier AI companies and their tendency to applaud themselves.

But the old models still exist at trivial marginal cost. The frontier models would need to dominate every price point to really take all and so far they haven't been.

  • Don't worry, AI boosters will be in here soon denigrating anyone that uses anything but the latest and greatest models as irrelevant.

Not just compute but energy. Most of Europe has no access to the cost effective power generation needed

  • Build Nuclear, Build Thorium the Chinese are building whatever they can. They’re not locked in by special interest. Is that because they have lots of engineers on the job in government?

  • France is actually pretty cheap in Europe. About 15% more than average USA electric prices (but I know that varies a lot across the states so still likely much more than the cheaper areas)