Comment by typ

6 days ago

There are some benchmarks I personally trust. But the most reliable indicator other than public benchmarks is the quality of products people are working on, and the model they use for the work. Not a benchmark scoreboard or a one~few shot demo, but the product they have been building for weeks. The reality is, even the diehard advocates who build and sell tooling for open weights, still use proprietary frontier models for their jobs.

Sure, the tools and models I use to build real products says a lot about how much I trust them. The closed frontier models are the most familiar, but I'm trying to unlearn my dependence on them. I just don't like them, and I believe AI should be open. In the short term I may lose out on some coding productivity, but in the long term I want to be a better AI engineer anyway. Open model + configurable harness forces me to learn how to get things done more efficiently, and understand how harnesses and models work under the hood.

I'll see what I can ship while not relying on these closed frontier models at all, as unreasonable as that might be. And if I don't get any customers that's on me.

I've used Kimi K3 for my (hobby) Linux kernel work recently, because unlike the offerings from OpenAI and Anthropic it didn't flatly declare everything to be a cybersecurity issue.

(And, yes, I know we should blame C. C turns every bug into a cybersecurity nightmare.)