← Back to context

Comment by ramon156

2 days ago

People will claim it's not comparable to Opus despite it beating the score. I'm not sure I disagree, but I'm also unsure whether I care. Most new models nowadays are "good enough". I cannot complain because I'd rather spend that time improving my prompts and docs. Opus might be a _slight bit better_ at picking up vague hints, but it's also extremely expensive, and I hit the 5 hour limit way too quick.

I care a lot about speed and efficiency right now. For my setup I would like to have 2-3 different model families. I've settled on GLM-5.3 (formerly Deepseek v4 pro 0813) for architecting, Deepseek V4 Pro 0813 for developing, and Gemini flash lite (any recent cheap model) for repo scouting. I'll add another one in the mix for reviewing (in this case Gemini 3.7) and that's all I need.

I've tried most models except Grok.

Qwen is too expensive IMO (Alibaba Cloud subscriptions are hard to come by and I'm not spending 50 euros a month for a tool, so 18 euros it is). If it ever becomes efficient enough to run locally I will definitely look back.

Claude is slow and expensive (the cache hit prices are absurd).

OAI is pretty good, I might add it to my arsenal seeing how cheap it is.

These opinions change every day. Last week I would've never picked Deepseek until I read about the pricing. even post aug 16 it's worth it (although it's getting close to gemini pricing).

Right now my costs are 12 euros a month (z.ai) + whatever deepseek consumes. This typically isn't more than 8 euros a week. 44 euros a month and I have a setup that is doing pretty well.

I'm convinced a lot of the anti-open-weight model comments at this point are inorganic traffic - there's trillions in investor money riding on a world where these models aren't cheap commodities. Having actually used things like the recent GLM, Kimi, and Qwen I think any edge the labs have is marginal at most and actually prefer the open weight models in most day to day usage.

Anthropic's recent releases are wordy to the point of exhaustion. Every time I use opus recently I find myself wanting to yell "GET TO THE POINT" at a terminal, which is exacerbated by it being slow.

  • Why are you convinced of that? Pretty much every time I see statements like that online, I can find plenty of organic traffic supporting it, not everything is a bot.

    I just wouldn't bias myself that way, most people haven't really used local models. This stuff is pretty much all subjective evaluation, there's plenty of reasons for people to favor certain models or disfavor others.

> I've settled on GLM-5.3 (formerly Deepseek v4 pro 0813) for architecting

Dude, GLM-5.3 released _today_.

The phrasing "I've settled on" is incorrect for this context.

  • I'm generally extremely skeptical about a lot of the model hype that show up in comments. Except when there is an extreme mismatch the performance, quirks and quality of these things are difficult to nail down. You wouldn't know that from the comment section of every single release.

    I think some of these are excited, eager users always ready to hype up the new thing. The same crowd that previously would constantly push for a rewrite from angular->react->svelt->god knows what. Instead now it is on a 6 week cycle and about models/harnesses.

    I think most of it is bot driven spam by the various labs. It's hard not to notice 3 month old accounts with very strong opinions about various frontier labs and little else.

    Ultimately, I think some of it is legitimate shifts in who's in lead and what is the best. You gotta dig through a lot of crap to get to that, and I don't really know how to to.

    Ultimately I'm saying is that I always applied a fair amount of skepticism about what I see in comment sections but these day it is extreme amounts.

    • > I think most of it is bot driven spam by the various labs. It's hard not to notice 3 month old accounts with very strong opinions about various frontier labs and little else.

      I've assumed the same as well.

      I also assume that many of the companies developing these models engage in benchmaxxing.

      At my company we've developed our own internal benchmarks for evaluating LLM models as they become available. The benchmarks are tailored to our particular use cases but the utility and knowledge our benchmarks assess is still fairly universally applicable. I see wide differences between what our internal benchmarks report and what the major benchmarks do.

      There was a whole lot of fanfare about how amazing GLM 5.2 was when it was released, but it was pure rubbish on our internal benchmark -- far behind OpenAI, Anthropic, Gemini, DeepSeek, etc. I don't know how to reconcile the fact that GLM 5.2 performed very well on some of the major public benchmarks, but consistently performs so poorly on ours. OpenAI models tend to dominate our internal benchmarks.

      3 replies →

    • I think most of it is bot driven spam by the various labs. It's hard not to notice 3 month old accounts with very strong opinions about various frontier labs and little else.

      Really no different than when there was suddenly online personas everywhere hyping up TSLA out of the blue. You can see the same thing going on with the BoringCompany subreddit. Crazy that the botnet master isn't able to convince us that Grok is also the best model. I don't think buying twitter was an accident it was probably just literally covering up the evidence.

      https://www.rhsmith.umd.edu/research/twitter-bots-boost-tesl...

    • IMO this is getting hyped because the 27b version runs on a decent gaming GPU. This is NOT a thread for their largest model, this model will run on a mid-high end gaming PC, which you probably have in your household. Mine is 6 years old and it runs quite well.

    • I strongly suspect that many of these accounts you think might be bots from the labs are just people who only have one interest.

      One thing I have noted a lot more of is that there is comparative fanboying going on. Like "this model has done badly, my favoured competition has a model out soon that will beat this in every way" — comparing a released product to unverifiable hopey claims about an unreleased product.

      It's tempting to assume that is bot stuff, but if you've been around any other "hot" technical hobby online (cameras, phones, 3d printers, whatever) you will know it's not. It's just fans aligning into teams, some of them laconic and amusing, some of them overkeen and toxic.

  • hence the "former deepseek v4 pro". I tried it out this morning and have had no complaints. I already liked glm 5.2

    • > Deepseek v4 pro 0813

      Which itself released yesterday? You're writing, reading, and evaluating enough software in a ~36 hour period to form, reject, and form another opinion about which model makes better _architectural_ choices?

      2 replies →

    • The sentence still doesn't make sense, because "settled on" implies a long testing phase with a verdict eventually emerging out of that.

      What you're currently doing is "testing out"

    • Honest question, how do you assess models this quickly? What metrics are you using? Would love to get my suite from multiple days and hundreds of prompts down to minutes. Got a few first pass tasks I run upon release for an initial experience, but those only work because even Fable and Sol fail despite objectively correct solutions existing, so it works because most models fail, but then, those are consciously not enough for coding, tool use, adherence or task specific inference and assessment…

      4 replies →

Grok 4.6 is a game-changer. I have yet to go back to other models after starting to use it. You just can't beat the price + output quality (even K3 is more expensive)

  • Supporting a far-right megalomaniac, whilst helping them to train their ML, and giving them all your data ... what could go wrong.

> Alibaba Cloud subscriptions are hard to come by

no they aren't. they discontinued their always-sold-out coding plan and launched QwenCloud (basically a friendly frontend with Alibaba Cloud as the hidden backend) and launched typical subscription plans for Qwen & co alongside it.

i think you will like luna if you haven't tried it yet

  • Luna is twice the price of Deepseek V4 Flash 0731, and less capable :/

    • With max reasoning, Luna is actually less than half the actual cost to run compared to DeepSeek V4 Flash 0731 with updated prices (based on Artificial Analysis Cost per Task)

    • No way is it less capable. When deepseek can't get its shit together, I give Luna a go, then Terra, then Sol.

      Most of the time Luna figures it out where Deepseek was failing. I rarely have to go to Sol levels.

You should check out Grok, it's quite a good deal from the Cursor subscription side but it's cheap even by API prices.

  • I think it's very clear that someone who has checked out all the models but the one that called itself mechahitler and is explicitly being fine tuned to support far-right politics is making the choice for reasons other than performance and cost. It's not like all the other models even had plausible claims to those metrics.

    • As I said to a dead reply, for coding all of that is immaterial, as long as it codes well then that's all that matters to most people, except it seems those who have an idelogical issue in which case the other model companies also have issues.

      6 replies →

  • [flagged]

    • Yeah I'm sure everyone on r/cursor or in previous HN threads about Grok 4.5 or 4.6 are all unserious and insane.

      No one actually cares about the politics as long as the model codes well.

      Edit, quite interesting to see the reception to this comment compared to essentially the same type of comment I made on a Grok 4.6 benchmark HN post: https://news.ycombinator.com/item?id=49275385#49275571

      It's true that Cursor gives a lot of usage with Grok, most users of Cursor don't care about Musk.

      12 replies →