← Back to context

Comment by port3000

21 hours ago

It's amazing how quickly Fable went from 'Game-changing model that needs to be banned' to 'Yeah it's alright, but OpenAI is also just as good and there are a couple of good open weight alternatives that are equivalent for almost everything'

The hype cycles are shortening, perhaps we really are reaching some kind of plateau this time (famous last words)

Open-weight models were lagging 4 months behind OpenAI/Anthropic at the beginning of the year. They are now just 4-6 weeks behind.

  • Kimi K3 is still worse than Fable and Fable was trained >4 months ago.

    • Why are you using the product release date for Kimi K3 and the training date for Fable? Either use the release date for both (6 weeks apart) or if you have it the training date for both.

    • To say X is perfectly bad vs Y is false.

      People use these models for diff things.

      Its quite possible for the things they are used for, people do not see much of a difference.

      Do you hold stock in Anthropic?

      4 replies →

  • And given that Chinese models are closing the gap there are basically two thing that could be happening. One is that they are moving faster than US companies developing closed models, and two that we're starting to hit a plateau for model capabilities where all the easy gains have been plucked, and now it's not really possible to move forward at the same rate on the frontier. Of course, both things could be happening at the same time.

    • Probably a little of both.

      Chinese labs have come up with a bunch of genuine innovations: GRPO, auxiliary loss free MoE load balancing, MLA, muon optimizer, and a bunch of other ones. The Deepseek papers are really well written, this isn’t just sneaking a peek at a peer.

      The problems are inherently harder now too, partially because they take longer, so your training pipeline is waiting for long completions.

      Also there probably is some “distillation” (technically pseudo-labeling, which is common in ML). But I wouldn’t put too much weight on it because that was true 18 months ago as well.

      1 reply →

    • This happened ages ago.

      But OAI and Anthropic are trying to cash in ahead of their IPO window. I think that window is pretty much gone now.

      1 reply →

Fable is still the same model, it’s still a great model, and to be honest all these articles writing and speculating on how the LLM industry is going to evolve are not that insightful nor interesting.

I don’t think one should pay much attention to them.

The plateau is inevitable because their rapacious training methodologies are only viable when there are no defense in place, but information continues to evolve, which means the models will have to be continuously updated, but will be doing so with less and less freely available data.

  • > with less and less freely available data

    My understanding is that the labs ran out of freely available data to train on a while ago, and now primarily rely on human data vendors such as Surge and Mercor to source their data.

  • We are only just starting to model the physical world. There will be more training on empirical data.