Comment by baxtr
8 hours ago
I think it might have accelerated things but on a much more basic level, there seems to be no real moat in synthesizing the world’s knowledge into LLMs.
8 hours ago
I think it might have accelerated things but on a much more basic level, there seems to be no real moat in synthesizing the world’s knowledge into LLMs.
There's no question that training leading LLMs requires some serious expertise and know-how, but surely already having advanced LLMs/agents must be helping tremendously not only for software engineers but also for those working on LLMs themselves.
I think one could describe LLM optimization as "hard but not a moat". Years ago, optimizing neural nets was described "graduate student descent" - it's tricky but throw enough conventionally smart people at it and it will happen. It's like tuning a hot rod and finding a reproducible bug in a large code base. It's hard and there are tricks but not absolute hurdles, no problems waiting for a conceptual breakthrough (and at today's scales, are there any problems waiting for an Einstein to solve? That's an open (AI) question).
I also think we're seeing the sigmoid approaching.
I want that to be true (assuming you mean specifically the second half of the sigmoid) just to give me room to adapt to the changes we've already seen; but I've seen comments saying things are slowing down since around when GPT-4 came out.
What is really going on: all the AI labs are doing panicked model releases (and panicked training of new ones) because Qwen4 is rumored to come out end of October and is rumored to be very nice. Question is: is it another "Deepseek-moment" nice? Or just nice?
Btw: with Qwen4 I mean the next large Qwen model that is based on the Qwen4 architecture (Qwen 3.8 flash next was "almost" based on the new arch but obviously was a small model)
It doesn’t matter.
What matters more is if firm’s start using a bundle of American and Chinese models and when they find their feet - how large is the market for frontier?
Frontier has to displace labour one for one at some point or it’s over.
I wondered if there would be a Qwen 4 or we would go straight to 5, re: tetraphobia, but perhaps it's more like an uno reverse card in this case
https://en.wikipedia.org/wiki/Tetraphobia
2 replies →
I'm pretty sure the point of Qwen3.8-Flash-Next was to get the open source engines to integrate the qwen4 architecture.
The fact that it basically broke open the local model supremacy was just a nice side effect.
I'm running: https://github.com/peonist-ai/halogen-server with a quant4, PLE offloaded, and it's resident VRAM is 36GB at 265k context.
Shave 10 more GB off and the TAM openai and anthropic are targeting is a lost cause. Local models are what 90% of people will need.
If the world governments can get a handle on the memory cartel, then there's no more moat for most normal humans.
Awesome 3.8 next runs great on my Framework Desktop so I'm loving more local models.
3 replies →
I think the moat is going to be compute. So far compute needed to push the frontier is still extremely cheap so the capital can afford to spread its bets. But when further improvement is going to cost in trillions, capital will have to pick a winner and bet only on him. It won't be a matter of finding the best bet, it will be a matter of survival.
This will cause the picked winner to get massively ahead with sheer compute alone used both for training and inference dedicated to recursive self improvement.
I guess all predictions age like milk, but here's one:
There's a law of diminishing returns at play here, and doubling the energy cost of training to wring 2% more performance out of the technology isn't going to be very useful, because most of the problems it is capable of solving will be solvable with the previous-gen 98%-as-good model.
("there's a law of diminishing returns at play here" is an article of faith. But then, so is the belief that these models will keep getting better).
As soon as you can demonstrate decent financial returns (ie. the AI can run a company better than humans can), suddenly it makes sense to put a lot more $$$ in even if returns are diminishing - since whoever runs companies the best gets control of a big chunk of the world economy.
10 replies →
Surely there is a point where algorithmic improvements will be more cost effective than buying more hardware.
If you're actually applying LLMs, all of the things around the LLM that adapt it to coding, for example, that enable it to use existing validation tools for code, and enable it to diagnose and fix tool chain issues that aren't directly coding problems, are what makes the difference between a model that that scores a little higher on a coding benchmark and a model that's useful in a particular code base on a particular platform.
Are there any use cases that have enabled one customer of a frontier LLM to outperform a competitor using a different frontier LLM? Or is this why we are seeing confected points of comparison like solving challenge problems in mathematics?
I'm afraid it might be the other way around. RSI might pick all of the low hanging fruit soon. There must be a physical limit of how much intelligence you can squeeze out of some amount of parameters and compute.
There are going to still be worthwhile improvements but they are going to be more like not how to make transformers 10x cheaper but how to make next training run cost 9 trillions instead of 10 with a very particular optimization designed at the cost of hundreds of millions for this one specific run.
4 replies →
But the old models still exist at trivial marginal cost. The frontier models would need to dominate every price point to really take all and so far they haven't been.
Don't worry, AI boosters will be in here soon denigrating anyone that uses anything but the latest and greatest models as irrelevant.
Not just compute but energy. Most of Europe has no access to the cost effective power generation needed
Build Nuclear, Build Thorium the Chinese are building whatever they can. They’re not locked in by special interest. Is that because they have lots of engineers on the job in government?
Training location is flexible. Iceland?
Europe has lots of zero-cost windows for electricity, and areas with cheap prices. The real issue is access to oil and gas.
3 replies →
France is actually pretty cheap in Europe. About 15% more than average USA electric prices (but I know that varies a lot across the states so still likely much more than the cheaper areas)