Comment by verdverm
5 days ago
We'll have ~Fable level weights in the coming days (K3). Look out to the end of the year and there will be multiple options. There are seemingly more frontier labs than the three US ones.
5 days ago
We'll have ~Fable level weights in the coming days (K3). Look out to the end of the year and there will be multiple options. There are seemingly more frontier labs than the three US ones.
I'll be the first to tell you I love the open models, the fact they exist and my ability to use them. However I doubt K3 is "~Fable" any more that any of the past ones have been Open 4.8 or similar claims. I've tried many of these (not K3 yet, I do want to) and they don't at all feel equivalent to what they score on benchmarks.
I want them to get better, I look forward to each new release, I play around with open models (often quants, but I've used full versions via OpenRouter), but they aren't the same as the offerings from Anthropic/OpenAI in my opinion (yet!).
This is the best side-by-side comparison I've seen so far, we're still waiting for it to become available for download, then we should get much better analyses.
https://news.ycombinator.com/item?id=48999291
I do agree that the benchmarks do not tell much of the story, but this applies to the closed models as well ime.
I’m not kidding when I say bullshit bench and simple bench are the only benchmarks that reflect real world utility in a way that matches my hundreds of hours of experience with frontier models: https://petergpt.github.io/bullshit-benchmark/viewer/index.v...
That said, while I find it hard to trust fireworks given their conflict of interest, their article is pretty good.
1 reply →