Comment by rbbydotdev

2 months ago

Looks like we are seeing small but mighty model breakthroughs, outpacing the pure capital firepower of SOTA providers. I love rooting for the little guy, but is it too soon to call it? To play devils advocate, could it just be the benchmarks are not efficient enough to capture success of real developer workflows?

I think people are going to continue to be surprised by the capability of small models.

Now, if you ask this model to have a conversation with you, it's gonna fail and be incoherent. But boy, does it sure reason through math problems well.

I've just started using qwen3.6:35b a couple days ago running on my framework desktop and rather impressed. It runs really well and reminds me of probably the first Claude model I used. It's the first local model that's actually working for me in a coding agent I've tried. Very exciting!

It feels sometimes like optimizations are only starting.

  • I’m beginning to suspect the closed SOTA labs were doing all these optimisations, keeping quiet about it, and just charging us out the yinyang for inference.

    • Also as much landgrab as possible for data centres, infrastructure, that can kepe it running for the next 5-15 years.

      Why does an M1 Max continue to remain capable, if not more capable with every passing year with LM Studio? :)