← Back to context

Comment by rockinghigh

21 hours ago

Open-weight models were lagging 4 months behind OpenAI/Anthropic at the beginning of the year. They are now just 4-6 weeks behind.

Kimi K3 is still worse than Fable and Fable was trained >4 months ago.

  • Why are you using the product release date for Kimi K3 and the training date for Fable? Either use the release date for both (6 weeks apart) or if you have it the training date for both.

  • To say X is perfectly bad vs Y is false.

    People use these models for diff things.

    Its quite possible for the things they are used for, people do not see much of a difference.

    Do you hold stock in Anthropic?

    • > Profile created 3 days ago

      > Unnecessarily aggressive

      > First ever comment said "Further releases of Chinese models that demonstrate the gap is not growing substantially is a huge problem. The spending will be called into question."

      Yeah I think you have an agenda

    • ah yes, because i said something factually accurate and vaguely positive about Anthropic I must be a shareholder which would mean I either run a venture capital firm or am a current employee of Anthropic...

      I wish.

And given that Chinese models are closing the gap there are basically two thing that could be happening. One is that they are moving faster than US companies developing closed models, and two that we're starting to hit a plateau for model capabilities where all the easy gains have been plucked, and now it's not really possible to move forward at the same rate on the frontier. Of course, both things could be happening at the same time.

  • Probably a little of both.

    Chinese labs have come up with a bunch of genuine innovations: GRPO, auxiliary loss free MoE load balancing, MLA, muon optimizer, and a bunch of other ones. The Deepseek papers are really well written, this isn’t just sneaking a peek at a peer.

    The problems are inherently harder now too, partially because they take longer, so your training pipeline is waiting for long completions.

    Also there probably is some “distillation” (technically pseudo-labeling, which is common in ML). But I wouldn’t put too much weight on it because that was true 18 months ago as well.

    • That's my thinking as well. The whole distillation thing is a distraction from the actual innovation happening in this space. What will be interesting to see going forward is what types of new techniques people manage to come up with to over come the current architecture limits.

  • Or option three is they are drafting hard off the frontier US models via distillation.

    • The process takes time because even when you're distilling answers, you still need to actually do reinforcement training on the model. And given that Fable and GPT 5.6 just came out there simply hasn't been much time to do that. On top of that, Kimi also does better than Fable or GPT on a lot of tasks, distillation alone can't explain that, meaning there is a difference in architecture. You can watch a talk from Kimi founder to see how Kimi was actually trained and why it performs well. https://www.youtube.com/watch?v=5CkCW1P-g88

      Not to mention that US companies models constantly distill each other as Musk was forced to admit under oath. This whole narrative has just been a massive cope.

      4 replies →

  • This happened ages ago.

    But OAI and Anthropic are trying to cash in ahead of their IPO window. I think that window is pretty much gone now.

    • My prediction is that they're going to angle to become a vendor of record for the government and get bailed out. That's the only path at this point because there won't be any competition from China in this niche.