Comment by joshstrange

5 days ago

I'll be the first to tell you I love the open models, the fact they exist and my ability to use them. However I doubt K3 is "~Fable" any more that any of the past ones have been Open 4.8 or similar claims. I've tried many of these (not K3 yet, I do want to) and they don't at all feel equivalent to what they score on benchmarks.

I want them to get better, I look forward to each new release, I play around with open models (often quants, but I've used full versions via OpenRouter), but they aren't the same as the offerings from Anthropic/OpenAI in my opinion (yet!).

This is the best side-by-side comparison I've seen so far, we're still waiting for it to become available for download, then we should get much better analyses.

https://news.ycombinator.com/item?id=48999291

I do agree that the benchmarks do not tell much of the story, but this applies to the closed models as well ime.

  • I’m not kidding when I say bullshit bench and simple bench are the only benchmarks that reflect real world utility in a way that matches my hundreds of hours of experience with frontier models: https://petergpt.github.io/bullshit-benchmark/viewer/index.v...

    That said, while I find it hard to trust fireworks given their conflict of interest, their article is pretty good.

    • That is an interesting benchmark! Though I wonder if it is probing at certain behaviors, which may be more revealing to probe at? Will keep it handy, thanks for the share!

      I felt better about Fireworks leadership after hearing some of them more candidly on podcasts (there are 7 founders!)