Comment by AmazingTurtle
12 hours ago
> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?
Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.
Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.
I don’t know what people do with the open models but having tried a lot of them I just can’t make it make sense. they’re too dumb and it effectively makes them useless (to me). it’s probably worth being honest about the low ceiling here.
IMO this comes down to your harness. Any frontier model from a huge shop has an inherent benefit in the system you're using it in. Search, memory, skills, integrations you don't realize even exist make them much more powerful. It is some effort but I recommend trying Hermes Agent and setting it up fully, that's the closest you'll get to a more complete experience.
Really!? Glm5.3 is my daily driver and I feel im having the most productive experience with agentic collaborations so far, by a lot. Using pi with tons of custom extensions, that to be fair I developed since making the jump off of codex and claude about 12 weeks ago. I primarily do not write code for a living. I do a lot of modeling and commercial analysis and a lot of math (related to differentiable simulation)
Same here. Moved from Opus to GLM 5.2 to 5.3 and I've been pretty happy with the result. Mainly, it doesn't hallucinate and convince itself of mistake so it's good at retrieving information or asking the user for it. Opus and Fable always state something, then try to "prove" it but end up convincing themselves of the wrong thing. Having subagents for retrieval and validation helped but were not enough.
3 replies →
What HW are you running this on?
3 replies →
Whenever I see this comment I'd wish they'd preface it with their hardware.
Yeah, expecting the world when all you have is a 8GB graphics card? You're going to be disappointed.
16GB is table stakes (IQ3_XSS). 32 GB is better.
Which models did you try for which tasks?
Qwen 3.8 Flash-Next is not dumb. If you've used it and that was your experience, your workload is either ultra-ultra-sophisticated or you're dealing with a broken quant/buggy chat template/other issue. That model is a smart, reliable workhorse.
Qwen3.8-Flash-Next seems pretty much auto pilot when I get it the right context.
Perhaps reverse the question: Are your build/construct requirements just really counter-productive to how LLMs need to understand things?
I've found constructing the code, writing the tests, adding the docs; then running through them gets most of the way there.
I've also found that making a simple obvious edit is a useless endevour when the LLM is primed for the long context tasks.
So, again, the question is reversed: are you over reliant on the LLM to do even stupid simple likes like editting a css variable?
if the answer to 'the model is bad at X' is "you're over-reliant on it" - then yes, the model is bad at X in comparison to alternatives.
Last time I estimated it was like 30 years to pay back. I doubt the hardware will even last that long.
I have 2 x ChatGPT Pro 20x, Claude Max 20x, and Kimi Vivace. It's about ~12 months payback for two units and the cable.
The problem is they can't fit any frontier level open models.
Is Kimi really competitive enough to have it in your mix?
3 replies →
how do you get 2 cgpt pro?
2 replies →
Last time I estimated, it would only take 3 months to pay back because the 1TB Mac Mini running Qwen RSIingly developed ASI and made infinity dollars off of crypto and I got put in jail by the SEC.
Where'd you get 30 years from? Show your work.
I would like to subscribe to your newsletter.
flash next is good, I've been running it for like 2 weeks now and it's pretty solid, hope you like it and it meets your needs. I still lean on Claude and codex a fair bit for harder stuff, but I'm rapidly moving towards 2x $20 plans instead of 2x $200 plans
I assume you’ve calculated expected cost vs subscription.
How do the numbers pan out? Ack that it isn’t always juts about cost, so even if it is pricier to self/host it might still be better for you for other reasons.
reliability and self reliance is worth a lot to most. Heck, you could be the best in the world at what you do, but if you're unreliable you wont find stable employment. So, not having some amoral shady company errode model quality out from under you constantly is also worth a lot more than simple cost balancing calculations can capture.
I'm getting really sick of the constant rot and "magic breakthrough" cycle, so im going full local, at expense on paper but being able to trust something which I need to understand the reliability of is priceless.
I like predictable. I'll take slightly less capable over unreliably capable, since with reliable i can calibrate my expectations and learn what aspects of my workflows to entrust and trust it will work. You simply cannot do that with models you don't control and in my experience they will all errode after the initual marketing wave passes, likely you eventually get fed heavily quantized versions and are expected to accept degraded service when what convinced you to pay was a far superior product. No such issues with local.
Thanks. If I may ask a follow-up: how did you decide what amount of compute to buy?
$10k - $12,000 I estimate, for reference.
Serving compute is their main value prop
Yet… even Altman called out Anthropic for serving dumbed down models.
Shits weird man
> Yet… even Altman called out Anthropic for serving dumbed down models.
Even Altman called out Anthropic? Isn't Anthropic the biggest competitor Sam Altman has?
> Isn't Anthropic the biggest competitor Sam Altman has?
Not by a mile. They're even (probably illegally and there's apparently a class action lawsuit oncoming: at least something to that extent was posted on HN today) teaming up, as a duopoly, to push for the same bullshit regulations / "we need to slow down AI research".
The reason they're teaming up is the real competition is, as in many other domains, China.
Yeah, and? They share a business model
>Also I will likely save some money on subscriptions.
Unlikely. The $200 Claude subscription allows for billions of tokens/month, and that kind of hardware will take years to amortize.
I wouldn't be so sure. The generosity of the subscription plans has declined GREATLY over the past 6 months or so. They are likely trending towards api pricing parity. In which case, having your own hardware makes sense if you can utilize it well.
I max out my Claude Max plan every week, and I can measure the output, and for me it's stayed fairly constant, subject to the various "bonuses" whenever Anthropic is feeling the competitive pressure.
Nah, I used $1500 in api prices last week, after caching (8900 ignoring caching) for what amounts to $50 a week. They’re nowhere near api prices yet
There could be gym logic at play. Hundreds signed up, 20 people actually exercising. Though it's probably more likely in the lower tiers.