Comment by novok

19 hours ago

The open models are good because of distillation, which the US labs are actively working against via not revealing CoT ever and now you can see with OpenAI Astra 6 not even having a lot of CoT equivalents being emitted as tokens. Once the anti-distillation stuff is in place the open distillation models will probably start having larger and larger gaps.

If the companies survive the next few years, which they probably will because they represent too much of US economic growth to allow them to fail, this gap will keep on expanding.

Starting from zero without distillation is a lot harder, a lot more expensive and a lot more work. OSS models is what a laggard does to get adoption. China's gov't might keep on sponsoring it as a counter GPU embargo thing, but when gov't get involved, usually the other side gets involved too.

As for people asking where is the evidence for half of this, you will never have public evidence for most of this, but deduce what the partly hidden parts reveal about the whole and it is fairly obvious, especially if you look at the past behaviors of the governments and other actors.

If we accept the premise that the top Chinese labs are simply distilling and can't compete otherwise: why don't US labs simply do the same thing? Distill their own models and slash their costs by 99% while keeping the same quality output. It should be a piece of cake if even the open labs can figure it out, after all.

One way or another they're getting the same results as proprietary labs, with a fraction of the hardware for a fraction of the cost. OpenAI can't keep raising funding rounds of $100billions to subsidize their compute costs and get results by brute forcing parameter count. And if they're having trouble keeping up with Chinese labs' efficiency, maybe they should stop worrying about distilling and instead hire some of the smart people responsible.

  • Because distillation-only is a quick performance shortcut that only lets you get to the level of the thing your distilling for the most part or a little bit worse and does not allow you to actually progress past it. It's like only being able to make VHS copies of videos, and maybe do some basic video editing without being able to actually go out with cameras and make new movies.

    To actually have something competitive and improved within the next 3 months and not be perpetually behind, you need your own independent model creation process. So to extend the metaphor, a complete movie studio with cameras, actors, staff, sets, budgets, etc. It's the right strategic move to do when you are GPU constrained, which the Chinese labs are, but it won't let you get past it.

    A bunch of pedantic people will come out of the wood work citing a bunch of things saying that is not the case because of some detailed mechanics of how model training works and they will get fixated on some of the words I used, but zoom out to the level of what an AI lab is able to produce and this becomes evident.

    • That doesn't really answer the question. What they do internally for training the next model is a separate issue. I'm talking about the models they offer publicly.

      Per the article, companies are dropping OpenAI+Anthropic (partly) because of costs. If distilling is so simple and easy, why doesn't OpenAI take this "quick shortcut" and serve a self-distilled model externally, so they can charge reasonable prices and stop bleeding customers? Surely they can at least match the Chinese labs' efficiency, right? Wouldn't more customers and less opex look good for the IPO?

      6 replies →

If that was the case the major labs would have done that.

They’re burning cash like there’s no tomorrow. They desperately need to show that they have a real business and not just a giant burning pile of cash doing academically interesting things. If they could simply sell models that are 95% as good at 1/10th the price they’d do that. They’re losing the enterprise sector because they’ve not done that.

  • I'm not sure why this style of cocksure Zitronesque comment is so trendy. It seems strange to make this statement, as if one has teleported to September 2026 from the end of the universe. Commenters in this vein aren't looking back to see that Anthropic and OpenAI models were the only truly usable ones for software development at the beginning of the year. Even Deepseek, which was and is revolutionary, was not useful for independent code commits longer than a few dozen lines.

    It certainly is possible that Anthropic and OpenAI are doomed to bankruptcy, but without a known drop in revenue it seems exceptionally confident to make such a certain claim.

    • Did you read the article? It’s not some future predicting theory. Open AI and Anthropic losing these workloads is what’s happening today.

      1 reply →

Basically your argument relies on two claims, both of which must be true.

1. Competitors to OpenAI and Anthropic are good because of distillation.

2. OpenAI and Anthropic will come up with some methods for preventing distillation in the future.

Both of these are dubious imo. For RLVR tasks like coding in particular, you definitely don’t need continuous distillation to improve, otherwise OpenAI and Anthropic themselves would not be able to improve because there is no better model to distill from.

Look up how many Chinese are pursuing a computer science degree.

Chinese AI labs do well because they have the best AI people coming out of a huge talent pool.

  • And yet their models are all behind the US models. I don't think LLM progress is strictly correlated with the number of computer science degree holders in a specific country.

    [There are many examples of distillation attacks.](https://cyberpress.org/anthropic-claude-ai-distillation-atta...) Of course, you could argue this is just a form of "learning" from other intelligent systems, and this is arguably what LLMs have been doing from the beginning. So I don't begrudge the Chinese labs doing it - they all do it. But let us not pretend that it's not happening.

Yes, evidence is needed but especially for the claim that distillation is what makes these open models good. Serious citation needed.

Think about it: even if they distill the shit out of frontier models, the model still gotta learn, right?

If anything, as you can see from the K2 Horizon release, aggressive (self-proclaimed) reliance on distillation does not result in a model that has remotely any frontier capability. Try asking K2 Horizon to write iambic pentameter for instance, or even give it the car wash prompt. I tried both these on the Q8 quant for the 7B model and the results were depressing.

  • To whoever downvoted this, it would be helpful if you actually reply with something substantive. The post I'm responding to makes sweeping characterizations, and I challenge it with a relatively good heuristic and indirect evidence, and only get downvoted?

    • Because it's evident you didn't engage with the last sentence I wrote properly. "Citation needed" in more words is not a sufficient response.

      > As for people asking where is the evidence for half of this, you will never have public evidence for most of this, but deduce what the partly hidden parts reveal about the whole and it is fairly obvious, especially if you look at the past behaviors of the governments and other actors.

      To help you further understand, the Chinese labs will never admit they were distilling until they are better than US labs with their own non-distilling process or some sort of espionage-like public reveal shows it. So you need to look at secondary indicators. Much like how the chinese government lies about their economic stats so 3rd parties use secondary indicators to figure it out.

      1 reply →