Comment by cyanydeez
2 days ago
this bodes well for continuing to refine smaller models and open sourcing them.
There's a delusion that what America's AI companies are doing is "best"; the chinese should realize that the forefront is bloated and there's likely hundreds of speed ups viable. Pushing open weights will continue to grind down the bloat.
I was gonna say, this just puts more pressure to deliver ground breaking research with limited resources. And if history teaches us anything it’s that scarcity produces ingenuity.
http://www.incompleteideas.net/IncIdeas/BitterLesson.html
> One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning.
Right, and if you come up with an efficiency gain that makes scaling better, e.g. a 50% reduction in required compute. Or even asymptotic improvements e.g. moving from quadratic to linear. Then you're much much better off.
There is nothing about the bitter lesson that says just be dumb and pour money into a hole, you still have to invent the methods to scale well, and being under immense pressure with constraints seems likely to produce that research.
2 replies →
The implication here is that the only gains left to be had are from scale. That we are already maximally efficient. If that's true, then how has OpenAI repeatedly bragged about reducing the cost of their models by orders of magnitude? (And DeepSeek Flash even more so, of course.)
But we have not been maximally efficient, we keep gaining efficiency. If we keep gaining efficiency, why should we assume it is impossible to gain more?
So what? There are physical and economic ceilings on dumb computation scaling.
americas tech stack always ends up bloated. not everything is worth learning.
endlessly knowing about pokemon is not delivering value proposition
cancer also grows carelessly.
> "There's a delusion that what America's AI companies are doing is "best""
Not sure if the word "delusion" is the correct word here? It has not been proven in either direction. We can all see lots of possible issues with it, but it is also possible that it could be what is needed to unlock key capabilities.
We can see that the Chinese models have been getting better, but OpenAI is out there supporting 10 million active users with their frontier models, and now we know that Deepseek can't even get what they need to properly train models.
They can't get hardware because the US has put restrictions on how much can be sold to China. There is not a technical or know-how limitation, but political. Deepseek could otherwise write some checks to NVidia for what they want.
Thanks to the import restrictions, I expect Chinese GPU hardware to be competitive within a few years.