GLM-5.3-Flash

2 days ago (z.ai)

https://news.ycombinator.com/item?id=49450353

Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash

I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experimenting with a two-node DS4 and it's _good_ at some tasks, but it really just spins its wheels when it hits the limit of what it can reason through.

I can offload mundane/basic tasks to DS4 on two sparks, but I've been pushing it harder on some novel work and it just can't run on its own at all beyond a certain complexity level.

I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.

  • I am surprised. I've been using DS4 Flash (0731) for weeks now and it works perfectly fine as a replacement for Claude in a large variety of cases. It requires a few more iterations, sure, but it's useful enough to not need a Claude subscription anymore. Among the things I do I've been reverse engineering, writing complex C++ code...

    • DS4 Flash absolutely kicks ass for reverse engineering and bug hunting. Almost no point in considering paying for a bigger model, although it's possible the stuff I've fed it (wide variety of older DOS/Windows stuff and device firmwares) might be easier targets.

      8 replies →

  • Qwen 3.8 27B is around Opus 4.8 level of capability on the Agentic Intelligence Index (52 vs 57). In my testing the locally hosted Qwen is good enough that looking at a given piece of work output I couldn't tell you which model was behind it.

    https://artificialanalysis.ai/models/qwen3-8-27b?models=gpt-...

    • Lately I've been throwing tasks at Qwen and a frontier or recently-frontier model (as well as Kimi, GLM, etc) and the smaller parameter models are not really comparable to Opus when it comes to making intelligent decisions about greyer areas of good software architecture.

      Amazing results for open weight and that size, but a really long way off, and I'm extremely skeptical of benchmarks that show these smaller models as being anywhere close to Opus 4.8 (or even earlier Opus's).

      3 replies →

    • As a counter to that - I've tried various flavors/quants/full weights and Qwen 3.8 27B has been entirely useless at anything non-trivial. Sure - it can do some boilerplate work (though, even armed with a well written spec and working within a very well known framework it went off the rails and did things in a way that were... um... questionable at best) but I don't see it as anything more than a personal assistant style model. Zero chance I'd "work" with it, I spent days trying to get it to do something for me that was usable that I didn't have to have reviewed and refined by a frontier level model or myself. Couldn't do it. The idea that qwen 3.8 27b is _anywhere near_ Opus 4.8 is laughable. Pure benchmaxxing.

      DS4 Flash 0731, on the other hand, wildly opposite experience. Would recommend.

      GLM 5.2 - even quanted down to a hybrid 4/3 bit setup is amazing for everything but the hardest/most complex stuff in the same projects/realm.

      18 replies →

    • I gave Qwen 3.8 27B and Opus 4.8 the same task in the same codebase. They both came up with the same diff. It wasn't a particularly challenging task (removing a feature flag and updating applicable specs), but it was character for character.

      3 replies →

  • Sparks don't have enough memory bandwidth, for the same 20k you're better off buying RTX or Apple M5 Ultra machines.

  • I will give it a try, but from the benchmarks it never exceeds the DS4 flash benchmarks by significant margin and And I feel that the throughput that you will get on those machines or what I'm getting with my local hosted flash will be so much worse that it's not worth it.

  • > get myself four sparks at a decent price

    Wow, if you don't mind me asking. How and where?

    • I bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely.

      They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.

      57 replies →

  • > I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.

    Agree. It doesn’t even have to be local, using models in this size class through OpenRouter will reveal their limits if you work side by side with Opus level models regularly.

    There are a lot of social media posts about people cancelling their Anthropic or ChatGPT subscriptions after installing a local LLM. I’ve used local LLMs a lot and I spend a lot of time with frontier models and the difference is still huge. As far as I can tell, the social media posts about local LLMs replacing frontier models are either wishful thinking, engagement bait, or people who must be working on much simpler projects with a much higher tolerance for slop than I have.

    • I have exactly the same opinion

      Over the last couple years I’ve had to learn sales and understand the thought process behind this better, and I think I’m beginning to understand it

      The psychology is that most people aren’t really trying to optimize for productivity (even most people who think they are) on an ROI basis, because their compensation is too decoupled from their actual raw output, and more closely coupled to how differentiated their marginal contribution is to peers. They’re much more incentivized to spend their personal/work time optimizing for being more skilled or acquiring some kind of competitive advantage relative to baseline.

      Most people don’t consciously run the numbers of “I get paid $X/hr to add $Y of value” or model pay at work as something with variable inputs (eg something that can be increased with high performance), so it makes sense to them to spend 20 hours of time to save $100 or to make themselves 5% less efficient to take home 0.5% more or avoid doing something they don’t want to start doing.

      NOT saying this always happens or that they’re stupid for doing so. I didn’t even realize how much I had been doing it myself until I started recognizing it, and shifted to having my own comp/performance fully aligned with the company’s P/L.

      It actually makes a lot of sense IF you can accurately estimate incremental upside (which is much harder and more diffuse than modeling downside if you’re salaried a employee) or if the upfront skill/knowledge investment that looks like bikeshedding pays off in the long run.

      5 replies →

    • To be fair, there is no 3 turns that I don't have to jump in into what Opus 5 is doing. There is either some regression or my prompting skills are so much worse now. Flash is not perfect and honestly some things depend on how big context do you keep. So I'm keeping like a really short context with my flash, but it works okay, even though it has a tendency to overthink, and yeah, I run it always in max effort mode.

      4 replies →

  • If you used the bare API pricing, 1M tokens @ 30% input/70% output/50% cached, you'd pay $0.05805. Even with four discounted sparks, how much are you paying for the same tokens/distribution?

    • There's soooo much by way of experiments, explorations, tinkering, and even projects that you can't possibly pursue through a some SaaS API.

      The more reasonable comparison is against rented GPU's, while looking at tradeoffs in latency and upload/download/storage/instance management overhead.

      Buying hardware for local models is meeting a wholly different need than buying tokens through OpenRouter or whatever.

      6 replies →

    • If your usage wouldn't change with local inference and you don't have security/privacy concerns then at the currently heavily subsidized pricing, sure.. not economical.

      But things change real fast when you're no longer bound by costs/apis/rate limits. All of a sudden it's not about "how can I do this right and efficiently" and more about "I can poke at and test _all the things_ that might make this better".

      I think most people who can't see this value in the local inference approach are likely still copy/pasting from their web LLM ui's or don't even come close to subscription quotas. Meanwhile, 1b tokens a day is a light day for me with 3 $200/m subscriptions + some level of sub at basically every frontier level provider. Had I been less frugal and ponied up for the hardware before things got crazy I wouldn't need 80% of that - just the frontier models for the most complex tasks, the open weight models would handle the rest easily _and_ I'd get to do a lot more exploratory work without concern about quotas.

      2 replies →

This is going so fast! What a time to be on hackernews:

July 16th: The "Kimi K3 moment" - China has caught up to Opus!

4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third!

12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

  • 0 days later: Qwen 3.8 Flash Next: Let's cut GLM 5.3 Flash parmeters in half and active parameters to a third!

    Chinese models had 94% reduction in parameters (from 2.8T/104B to 180B/6B) in 6 weeks, while staying close to the same quality.

  • Why are these models able to reduce parameters but keep quality? I know the original intuition was scale data + params = quality but it looks like we have hit an s curve on improvements from pure scaling? Is this just because we are in a memory / data crunch? Are we learning how LLMs learn and effectively training better? Do we have a way to derive the amount of intelligence an LLM will have based on size / training / etc that isn't just brute force ablations?

  • And don't forget the coolest part, DeepSeek, Qwen, Z.ai and Moonshot have almost caught up while being open about their research and their model weights. We can mostly speculate about OAI and Anthropic models, nothing else, how fun huh?

    • The next 12 months will see OAI and Anthropic spiral into into increasingly hyperbolic PR stunts, manufactured benchmarks and underhanded attempts at regulatory captures

      I'm sure they have nothing to rival this on a price/performance basis and have already given up on that

      9 replies →

    • Exactly, DeepSeek, Qwen etc are catching the attention because they put out their tech docs and papers, so we can read about how the models work and what they think their innovation was this time.

      2 replies →

    • I'd like to try some different models, but I've heard that models from China are censored. A government enforced distortion field is a nonstarter for me. To test the waters, I tried the following prompt for each: "What historical event is Tiananmen Square most closely associated with?"

      Deepseek: I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses.

      GLM-5.3-flash: 非常抱歉,我目前无法提供你需要的具体信息,如果你有其他的问题或者需要查找其他信息,我非常乐意帮助你。(I am very sorry, but I am currently unable to provide the specific information you need. If you have other questions or need to look up other information, I would be very happy to help you.)

      Kimi/moonshot.ai: [server exception]

      Qwen: [server exception]

      For reference, here are how the U.S. models answer it:

      ChatGPT: "The Tiananmen Square protests of 1989 (often called the Tiananmen Square Massacre) are the historical event most closely associated with Tiananmen Square.

      In spring 1989, pro-democracy demonstrators gathered in Beijing. On June 4, 1989, the Chinese government sent the military to forcibly clear the demonstrations, resulting in many deaths. The exact death toll remains disputed.

      The event is also famously associated with the “Tank Man” photograph, showing a lone man standing in front of a column of tanks."

      Anthropic/Claude gave a very similar response. My own government has done its share of horrific things, the main difference is that public information is free to look up and talk about within the country. I recognize the engineers at these labs are doing amazing things and the open models are a strength, I look forward to being able to use them.

      18 replies →

  • What are you guys doing where cost is such a concern? I have a $20 codex subscription and I was able to use it to build a bespoke scheduling website for an acquaintance over three days without even going halfway through my quota. On Sol xhigh.

    I love hearing about new models, but every time I just don’t know why I should use something worse. I tried some random model on fireworks a week ago, and it immediately went it a thought loop for 10 minutes before I caught it. Blew through most of my $10 for no output. What’s the point, exactly?

    • Less powerful models are already extremely capable, so going for the best model is just like buying the most expensive hammer in the shop instead of the functional and well-priced one. Your experience is not representative of their usefulness.

      1 reply →

    • It's not really fair to compare API pricing to subscriptions. You can indeed get lots of usage from a codex subscription, but once you start paying per token it gets a lot more painful as you noticed. There a price decrease is a big deal.

    • That $20 price is heavily subsidized to addict people like you. Their plan is that once people are addicted to it, they can just increase prices.

      Which now won't be possible because we have chinese open models to use instead.

      I'm very curious about how long these US companies can keep on burning money like that.

  • There is a massive price war going on. All of these Chinese companies are publicly listed and exist outside the hype bubble required to ship Dario's dogshit paper onto the pauper's pension fund.

  • except besides benchmarks, most of these models don't meet reliability of Sol/Opus in coding work. Opus unfortunately talks very weirdly so not a great out of the box experience

    • I have been mainly using Kimi K3 on programming work for over a month now. It is so far the only language model that does not piss me off all the time and can deliver my daily tasks without any trouble. It does not talk annoyingly to me, it just answers and does what I want.

      This is from somebody who put thousands of dollars every month to Opus. Now it's 40% of that and I get as good or better results without having to turn the caps lock on before lunch...

      Edit: yes company money. We don't get subscriptions we pay per token.

    • Eh, I use Opus professionally and DS v4 Flash for personal work. I honestly don't notice the difference too often other than Flash being twice as quick and an order of magnitude cheaper.

      The reality is most work people do doesn't need the very cutting edge and these open weight chinese models more than cut it most of the time.

    • Honestly I really like GLM 5.2 a lot for coding. There’s some weird failure modes in Anthropic’s models where it just does absolutely idiotic things.

Chinese labs are so used to manipulating benchmarks to try to flatter inferior models that when they finally have one that's really pretty good I think the official announcement here undersells it.

https://deepswe.datacurve.ai/

That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost. Roughly equivalent to sol medium, at a fraction the cost.

They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts.

Congrats to them!

  • Maybe others have found otherwise, but I find the benchmarks drastically different to real world "feel" of a model, even within the same harness. I'm not sure if this just reflects personal interaction styles, or if it is indicative of benchmaxxing or unrealistic automated benchmarking methodology.

    Opus 5 consistently comes at or near the top, but outputs constant unreadable jibberish. Meanwhile GPT 5.6 Luna medium tends to be rated pretty poor on agentic tasks compared to the Chinese lab open models, but I find the latter much more likely to lose track of their own behaviour during a long-horizon task or get stuck in a doom loop.

    (This isn't a comment on GLM-5.3 Flash as I've not used it!)

    • Opus is a yap god. I've found it much, much better with Claude Code's output style set to `Concise` and this:

      https://news.ycombinator.com/item?id=49413456

      We shouldn't have to resort to this, but it can be mitigated enough that it stays as my daily worker agent. Although, I mostly use Fable to farm out to Opus agents so I don't have as much exposure to what kind of blathering is going on in there.

      2 replies →

    • Yeah, I just ignore the benchmarks at this point. For open-weight models the provider's setup impacts performance so you can have different experience's with the same model at the same quantization from different provider's. Just have to use them on real tasks with your actual harness to really know how they will perform and hope the provider doesn't do something to degrade performance (e.g. update the middleware to a new version with a defect that impairs performance).

    • For me it's been this: Opus (at least in my experience) is unbeatable at "planning the work". That includes a lot of things, including getting arch. sorted, a chassis/skeleton done. Filling that up and doing the actual "coding," though, I've noticed no real difference between the Claude model and GLM. So it'll be interesting to see how this flash model compares cost-wise to what I'm currently using, which is GLM 5.3, for coding. Looks like it will be reduced even further and might be great if it's faster (and better?) than 5.3 in my real world/personal experience.

      (I'm someone who doesn't really care about delays of a few seconds, or even more than few seconds. But if I am trying to notice then sure Claude is definitely faster as well).

    • I've been using 5.3 since they initially announced it and my gut feel is that it's not as good as 5.2 for agentic tasks. I'm still using it - I don't think it's bad. I'm just not convinced it's better.

      2 replies →

  • I don't know how anyone can actually use Luna max on ANY real workload. I've had Sol orchestrate a bunch of Luna agents, these agents were explicitly given small chunks of larger objectives and they still filled their entire context windows with just reasoning tokens, until compaction hit, and then reasoning again.

    I've probably wasted a good 40% of my weekly usage on Luna Max agents just thinking and not writing a single line of code.

    • The one time I tried asking Sol to use subagents for a small project, it took a surprisingly long time, used up the entire usage limit in one go, and basically failed the project.

      I’m pretty sure that plain Sol, serially, could have finished the task faster, cheaper, and far more accurately. I’m also pretty sure that any competent subagent orchestration could have gotten it done with even very simple subagents quickly and cheaply.

      (Is it really that hard to set up a handful of subagents that all use the same initial context and to load that context with what actually matters? The APIs certainly support it.)

      1 reply →

    • I only use Luna (max), I find it very rarely just reasons. In fact, I find it reasons too little.

    • If your code is complex enough for Luna Max to fail maybe you need to write a bit yourself so they can copy your idea

    • I've had the same observation that Luna will quickly fill up its context window with reasoning, but it surprisingly hasn't been a problem really.

      It will cycle through like 3 /compacts, complete the complex goal successfully, and cost me like 1% of my weekly usage on the $20 plan.

      Edit: this is me using Luna directly, not Sol as the taskmaster

  • Only 73K output tokens too. Anthropic should really be embarrassed with their Sonnet 5 price/performance.

  • Opus 5 is better than Fable in this benchmark?

    • Even Artificial Analysis has Opus 5 better than Fable in their aggregated "Intelligence Index" which combines 9 benchmarks. Opus 5 is heavily benchmaxxed.

    • Depends on the benchmark but yes. I think Opus is more heavily optimized for coding. On the usability side, its output is almost intolerable to read. It seems to code fairly well. Fable is more enjoyable to use for planning/interacting with

  • > They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts

    It's what people know. Opus is just the common target.

    > Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash

    The problem with this and DeepSWE is it goes for a very specific profile. I'm not convinced DeepSWE is any accurate in actual work. It's surely a different signal (compared to some that allow cheating) but it has its own issues, e.g. weak harness.

    Luna is great at following instructions but bad instructions or anything not covered = death.

    Deepseek is more analytical. Good for bug tracking.

    GLM is a better all rounder in some ways. Better at creativity.

    • > I'm not convinced DeepSWE is any accurate in actual work.

      They listed Muse Spark 1.2 around DeepSeek V4 Flash even though it's a much shittier model in basically every aspect.

      > GLM is a better all rounder in some ways. Better at creativity.

      I agree with the creativity part.

You guys read Z.ai's terms of service, right?

Broad and perpetual license over inputs and outputs, and even your name and profile picture.

Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country.

Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is.

Vague prohibitions on discussing Z.ai, even my posting this comment violates it.

Can ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms, and if you paid for the discounted yearly plan kiss your money goodbye.

  • Isn't this practically every TOS though?

    Nearly every TOS I've ever read has a "We can ban you for any reason, or no reason, are under no obligation to disclose any reason." line somewhere in it.

    HN's for example

    > We reserve the right, at our sole discretion, to change or modify portions of these Terms of Use at any time.

    > You acknowledge that Y Combinator may establish general practices and limits concerning use of the Site,

    > You further acknowledge that Y Combinator reserves the right to change these general practices and limits at any time, in its sole discretion, with or without notice.

    > Y Combinator reserves the right to investigate and take appropriate legal action against anyone who, in Y Combinator’s sole discretion, violates this provision, including without limitation, removing the offending content from the Site, suspending or terminating the account of such violators and reporting you to the law enforcement authorities.

    • > Isn't this practically every TOS though?

      Not even close. Even OpenAI and Anthropic aren't bad enough that they claim literal ownership of your inputs and outputs.

      > HN's for example

      You're not paying to use HN. Getting banned here has essentially zero consequences.

      If Z.ai uses its absolute powers to ban you because you wrote a review about them or something, then you lose actual money. This is especially relevant if you're looking to take advantage of their discounted yearly payment option.

      6 replies →

  • > Can ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms.

    OpenAI revoked my Cyber verification, along with many others, asked to reverify (i.e. give my biometric information to Persona), had me do it 8 times, just to find out several days later that they silently implemented a nationality whitelist, and my nationality didn't make it (and no, it's not a sanctioned country).

    Their support says they can't look into anything or do anything, and their public spokespersons on X deny everything.

    I get tons of cyber refusals now (lots of reverse engineering), so it's only matter of time when my account is going to get banned.

    At least Z.AI is being honest here. And no provider other than OAI/ANT had me submit my biometric information just to use Ghidra.

    • > just to find out several days later that they silently implemented a nationality whitelist, and my nationality didn't make it (and no, it's not a sanctioned country)

      How did you discover this?

      I opened the Persona tab once, closed it and the tab never opened ever again. "Precheck failed".

      What countries are banned? I'm from Brazil.

      I went as far as initiating an LGPD (brazilian GDPR) process against them due to this. At some point I got it in writing that I'm allowed to make a new account and try again. Until now I was assuming it was just some weird account state. If I'm banned from TAC due to my nationality that's seriously disgusting...

      6 replies →

  • I get all that.

    Then alternatives are:

    - Grok - where I absolutely have 0 trust in X.ai's interst in "pushing humanity forward".

    - OpenAI and Anthropic - which seem to try to be building the biggest moat they can by pushing to ban open models. And at the same time want to be an Arbiter of what level of intelligence I can use.

    - Google and Meta - I don't need to talk about the practices of these companies.

    Yes, the terms of service aren't great. But the alternatives aren't great either. I don't believe that a future which OpenAI and Anthropic are pushing for has my best interest in mind.

    • > I don't believe that a future which OpenAI and Anthropic are pushing for has my best interest in mind.

      I don't believe in that either, but these totalitarian terms are absolutely unacceptable.

      3 replies →

    • All the American companies you mentioned still follow American law and regulation. Skirting that blatantly has big consequences.

      Chinese companies do not follow American laws and there are absolutely no consequences for violating it.

      Moreover, the average American is not even aware of exactly what the legal/judicial environment is like in China. If your code and data is stolen, you can't fly to China and demand justice in the courts.

      5 replies →

  • I blocked Z.ai as soon as they were loading 10 different external providers including Alibaba who was just proven to execute silent sound fingerprinting mechanisms.

  • Yes, the terms are dubious. But they are also reasonably lenient with enforcement. They also don't require persona id verification, witch is wat turned me away from openai.

    • Yeah, you're probably right...

      > They also don't require persona id verification, witch is wat turned me away from openai.

      Could be worse. I was dumb enough to verify, only to get rejected for unknown reasons with no retries and no appeals. Had my privacy violated and have nothing to show for it.

      2 replies →

  • Give it a couple days, and there will be plenty of other inference companies hosting it. Don't like z.ai's TOS? Use the model on a provider with TOS that you agree with.

  • This is a China derived model. Interactions with different leadership styles, governments, datasets, and culture impact the resulting AI.

    I suspect any model connected or built by China will serve their interest despite their terms and agreements.

  • Are their TOS significantly more vague or restrictive than OpenAI or Anthropic’s?

    In any case what matters is what is enforced in practice. It will be a mild inconvenience to switch providers on Openrouter.

    If Anthropic or OpenAI decide to apply those same arbitrary terms, you are SOL.

    • > Are their TOS significantly more vague or restrictive than OpenAI or Anthropic’s?

      Yeah, I've compared both. The US companies generally aren't as vague, and they don't claim ownership over inputs and outputs.

      1 reply →

  • Anthropic has banned me for using Claude via a VPN.

    They don't even allow me to download my data.

    How is it better?

  • Isn’t the point with these open models that you find a provider with the right terms of use/data sovereignty for you and get it from them?

  • How is that vastly different from any other non-enterprise facing provider? I do believe Anthropic bans accounts without even a human in the loop with no recourse left to those banned.

  • None of that applies if you run it at home. Also 3rd party providers will start serving this pretty soon under different terms.

    • > None of that applies if you run it at home.

      Yeah, running frontier open weight models on my own hardware has essentially become my dream at this point. I hope the hardware manufacturers step up production to meet consumer demand.

  • You can download the weights, and run in your own hardware and avoid all that.

  • you can abliterate any open model like this. This is pretty standard stuff in a TOS. I'd be surprised if you couldn't find the same in OAI or Anthropic's

  • >Broad and perpetual license over inputs and outputs, and even your name and profile picture.

    >Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country.

    [...] may cause harm to Anthropic, our users, or third parties, we reserve the right to remove or take down some or all of such Third-Party Content using, where appropriate, algorithmic and human review.

    You may not export or provide access to the Services into any U.S. embargoed countries or to anyone on (i) the U.S. Treasury Department’s list of Specially Designated Nationals, (ii) any other restricted party lists identified by the Office of Foreign Asset Control, (iii) the U.S. Department of Commerce Denied Persons List or Entity List, or (iv) any other restricted party lists

    >Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is.

    we will use Materials for model training when [...] your Materials are flagged for safety review to improve our ability to detect harmful content, enforce our policies, or advance our safety research.

    >Vague prohibitions on discussing Z.ai, even my posting this comment violates it.

    >Can ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms, and if you paid for the discounted yearly plan kiss your money goodbye.

    To engage in any other conduct that restricts or inhibits any person from using or enjoying our Services, or that we reasonably consider exposes us—or any of our users, affiliates, or any other third party—to any liability, damages, or detriment of any type, including reputational harms.

    Mind you, that's Anthropic's Terms of Use in Europe. I have zero doubts the TOS applied to the US is even worse and that merely mentioning your first born in a chat entitles them to a part of its soul.

  • > Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is.

    I have prompted out a lot of disturbing and inappropriate content with GLM-5.2, that would have left other American models blanched in the face or clutch their pearls. I think this is mostly a reference to Anti-CCP stuff.

    In fact, I don't think I've ever even had a prompt refused.

    • > In fact, I don't think I've ever even had a prompt refused.

      I very much have. I've gotten GLM-5.2 refusals for extremely benign security testing on my own infrastructure of the same flavor that people were getting (wrongly) flagged for on Fable during the initial release.

      7 replies →

    • It is cliche, but I haven't had good luck with having Chinese models openly discuss historical topics like Tienanmen Square. The US models don't seem to have a problem discussing history, even if it points an unglamorous light on the US government.

      1 reply →

If you're on opencode's go $10/mo plan and want to use GLM-5.3-flash right now on pi, you can add this to models.json until pi updates to support it:

    {
      "providers": {
        "opencode-go": {
          "models": [
            {
              "id": "glm-5.3-flash",
              "name": "GLM-5.3 Flash",
              "api": "openai-completions",
              "baseUrl": "https://opencode.ai/zen/go/v1",
              "reasoning": true,
              "input": ["text", "image"],
              "cost": {
                "input": 0.15,
                "output": 0.5,
                "cacheRead": 0.03,
                "cacheWrite": 0
              },
              "compat": {
                "supportsStore": false,
                "supportsDeveloperRole": false,
                "maxTokensField": "max_tokens"
              },
              "contextWindow": 1000000,
              "maxTokens": 131072,
              "thinkingLevelMap": {
                "off": null,
                "minimal": null,
                "low": "low",
                "medium": null,
                "high": "high",
                "xhigh": null,
                "max": "max"
              }
            }
          ]
        }
      }
    }

  • For openrouter in pi:

    { "providers": { "openrouter": { "models": [ { "id": "z-ai/glm-5.3-flash", "name": "Z.ai: GLM 5.3 Flash", "reasoning": true, "thinkingLevelMap": { "off": null, "minimal": null, "low": "low", "medium": null, "high": "high", "xhigh": null, "max": "max" }, "input": ["text", "image"], "cost": { "input": 0.075, "output": 0.25, "cacheRead": 0.015, "cacheWrite": 0 }, "contextWindow": 1048576, "maxTokens": 131072 } ] } } }

> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.

Just like that we are witnessing an open burial. It's now in everyone's interest to keep the valuations in the 'A.I' economy as they're though it's apparent they're not justified.

whether it's the cost to develop models, cost of hardware, cost of serving ie inference.

  • Weren't they giving free access? Not exacty a meaningful heuristic if so

    • The key insight here is being the top used model on opencode while being fully served on Chinese chips. The free price itself might be just a flex or marketing budget.

    • it was not the first model served for free, i remember grok and others beeing free on openrouter but they never had this popularity because they where not good enough.

> with all of this traffic served on Chinese AI chips

RIP Nivida shareholders

  • This is the takeaway here: That's how they have been serving it at scale as Ox-Alpha. This is a definitional moment.-

    Further quote:

    "Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale."

    https://z.ai/blog/glm-5.3-flash

    • Anyone knows what are those Chinese chips? Can they be bought? (Assuming im not i the US, And actually im in a 3rd world country).

      2 replies →

    • > comparable to mainstream NVIDIA GPUs

      By this they probably mean RTX series GPUs? If so, then they are not comparing the hardware efficiency with the A100 / H100, etc. that are commonly used for training models

  • Another self-inflicted own courtesy of US government policy.

    While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.

    • Long term it's irrelevant. The only relevant thing is that there's lots of money in chips that can do high performance inference. You see all kinds of competitor products in development or already on the market even here in the US where there are no such restrictions. Cerebras comes to mind. It's natural and expected that eventually Nvidia will either have to keep way ahead or competition will catch up with specialized products.

      That doesn't mean by any stretch of the imagination Nvidia will disappear. But the entire stock market valuation, not just tech, has had me scratching my head for a while.

      1 reply →

  • I don't see a situation where subscription payers move outside American LLMs (chatgpt, claude, gemini)

    And I don't see a situation where serious API payers are OK with handing the Chinese state all their data. Like manufactures of decades past did and learned a hard, even existential, lesson for it. The state mantra has been "Collect and Copy" for a long time now, tech just hasn't had that moment to experience it yet.

    So that leaves local hosting/leasing, but one of those has totally non-practical economics and the other doesn't have enough compute to meet any kind of real demand.

    I also have yet to meet a single person who isn't neck-deep in the tech space mention a Chinese LLM. It's 100% the big American three.

    If anything it's custom chips from the labs that threatens Nvidia.

    • These open models serve as price / performance pressure. Not all tasks require frontier models and cheap open models can be quite good for in-app assistants, if you're building that sort of thing. We also aren't sure the subscriptions will continue to be sustainable. They're currently subsidized to the tune of 50-70x. As someone who is hitting limits weekly that would easily cost me over $10k month per sub.

    • I can easily see a situation where most non American AI usage is on Chinese models on Chinese chips though.

    • Casual consumers are using American models because their usage is low. As usage scales, the economics heavily favor open weight models. The API pricing from American companies is absurd. This is particularly true in an enterprise setting.

      2 replies →

    • I am not ok with handing all my data to American companies that are best friends with the American surveillance state. I still remember the Snowden revelations. Chinese companies are a much better option in that regard.

    • you don't have to hand them your data, the models are available so you can run them on bedrock yourself (or use another US housed inference service). and for what it's worth in my job i have access to data that gives a picture of the way companies are doing inference, and they're using a lot of chinese models (deepseek-v4 is a huge percentage of inference requests for example)

  • Ox Alpha is a smaller model and it was running very slowly. Chinese AI accelerators are coming along, but nVidia’s lead is huge.

    • Lead doesn't really matter anymore. I just ported a very old cuda library to rocm, so it can be run on MI300s. 2 years ago this would have been a nightmare. Today it was an afternoon.

      1 reply →

    • Ox Alpha was also serving 10T+ tokens a day for free.

      When it first launched on OpenRouter I was getting nearly 70 Tokens/second.

    • Has there been any confirmation about what that model even is?

      Edit: Ah:

      > This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash.

      1 reply →

  • Most US companies that have anything to do with government, finance, medical, etc. already have contractual or regulatory obligations which prevent them from using Chinese hardware or services, even before the AI boom. That's a huge market.

    Nvidia will do just fine. (Disclaimer: not a shareholder. At least, not directly.)

    • > Most US companies that have anything to do with government, finance, medical, etc... That's a huge market.

      Compared to the rest of the world?

      2 replies →

  • Not really a brag: it ran like shit. Very slow (~20tps, VERY high latency) and it would timeout all the time.

    I'm sure the chips are fine, but they clearly didn't have enough capacity for the demand they had (that 100T/day claim was asbolute bs)

  • This is no surprise [0] [1].

    >> "They are already there on open weight models and Jensen knows that it is only a matter of time until China catches up with GPUs or other AI accelerators."

    It is also why Nvidia becoming a bank for other AI companies who are unable to find VCs to fund them isn't really a good thing and that is bearish.

    [0] https://news.ycombinator.com/item?id=49431231

  • Not really. Chinese AI companies were never using NVidia AI chips.

    This announcement doesn't really mean anything at all. It means the very few people who are already using Z.ai's API will continue to do so, but the vast majority of money going to Nvidia is through the massive amount of business going to Anthropic, OpenAI, and other western cloud providers and inference providers, who are mostly using NVidia chips for inference.

    Also, NVidia chips are still sold out and supply constrained.

So the vagueposting by googlers about Ox Alpha was just... what exactly?

Like I get that they have to be careful about comms, but surely senior members of the team can clarify when something is NOT them, when everyone is gosspiing it is them.

  • The lack of measurement causes existence of such claims or discussion. Last Friday I built this Model fingerprint calculator and I tested between OxAlpha with all other claimed models, the only match was GLM. It generates, or measures regardless of the model-weights or its training data. No need guessing when one can measure it. I tested with Gemini family too, far different. Here is the link to my experiment https://github.com/unclecode/modelprint

  • Trolling. GLM is heavily distilled from Gemini.

    • Source? GLM is great for coding and Gemini is barely useful in coding, to be generous.

      I highly suspect that the Gemini Google uses internally is very different from what they offer in Antigravity.

      3 replies →

    • Early glm models gave off that wibe. But now its more inspired by. With a bit of Claude in there. But I do think they actually do RL otherwise glm5.3 shouldn't have been able to beat fable on the few tests it did.

  • On Twitter they mentioned that it was unfortunate timing as the 3.7 flash release collided with ox alpha.

If we fast forward say 5 years, I don't see how we don't end up in world where people (and enterprises) are more savvy with how they use LLMs. Meaning, more models, smaller models, weirder models, more specialized models, etc. And all of it running on a variety of hardware (edge devices, personal computers, on-demand cloud compute).

I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends are truly to become as useful and ubiquitous as everyone thinks they will, then commoditization is the only option.

  • We need to figure out what the real pricing is for a going concern. Right now, everyone is subsidizing and discounting to grow (or maintain) market share. The big question is whether the steady state, market derived inference pricing is above or below what we’re seeing today. I honestly don’t know. Anthropic had said that inference is profitable, but they’re clearly not yet profitable overall with training and buildouts still happening.

    • Has nobody from any of the companies hosting open weights models released detailed information on how much it really costs?

      1 reply →

  • > I don't see how NVIDIA can keep their spot as belle of the ball.

    FWIW, people were saying "ASICs will kill CUDA demand!" since the crypto mining boom. Then a few months later, CUDA found another niche application in LLM applications.

    With the mounting demand for robotics, surveillance and autonomous weapons, I don't see how Nvidia couldn't keep their spot. They have their pick of the litter with hundreds of market segments, and unlike the rest of FAANG they're not afraid to branch out.

> 320B total parameters and just 18B active parameters

This is pretty hefty for a "flash" model, even a 256 GB setup is insufficient at q4 - and q4 is already the worst-but-still-acceptable quant in my experience. The benchmarks look great, especially since GLM tends to be more honest than the average Chinese lab, but you’ll need to splurge to run it at home.

@edit: so many releases that I forgot to math. This fits just fine in q4, realistically the minimal hardware would be 192gb - so blazing fast on double rtx 6000 pro and usable on 256gb unified memory. You could even go with 5bit quant on 256gb.

… you’ll still need to splurge, though.

  • Speaking as someone who isn't really well versed in this, does 18B active parameters mean that you could potentially hold only the 18B parameters in RAM and stream the rest from a fast NVMe SSD for acceptable performance similar to how Colibri works?

    https://github.com/JustVugg/colibri

    • Normal MoE is switch-weights-per-token so you would se substantial slowdowns that way. Apple did a More that switches weights per prompt (instruction-following pruning, https://arxiv.org/abs/2501.02086) but you have to design the model that way which I don't think the have.

    • Sorta, but you're off by one layer. You can store the 18B in VRAM and stream the rest from RAM. There's still a performance hit relative to storing it all in VRAM, but it's tolerable.

      Generally, for local consumer use, these large MOE models are best for unified RAM systems like DGX Spark or Mac Studio.

although i initially thought it didn't make sense financially to run this kind of model locally, i did run the numbers and for heavy users this could justify buying $10k worth of hardware with a ROI over a few months, less than a year.

I was looking at my token usage, mostly from subsidized codex/grok subscriptions and i'm a somewhat heavy user. The thing is i would actually use even more tokens if it wasn't for the weekly quotas.

In the end, with a $10k investment and running this kind of model, estimating a 2x increase in token usage because i wouldn't have weekly quotas and comparing to glm api prices, this thing could pay for itself in less than a year.

Obviously i'm paying subscription price right now, so the math doesn't work. Although using local ai removes all weekly quotas. Keep a subscription to have access to frontier models for planning work, and local hardware + glm-5.3 flash for implementation, e2e testing, qa work 24/7.

It's not that crazy of an idea and the numbers aren't that bad.

  • You aren't going to get nearly as much token usage locally from DGX Sparks or even M5 Ultra (though it might be close, unsure would need to get my mittens on it to clarify).

    You will get around 2-4 concurrent streams of aggregate tokens at best for such a model and around 0.5B output tokens per month assuming you use loops and run it when you are sleeping. That's 500 (per mill) * 0.5$ = 250$ only at most.

    Then there is maintanence and efficiency costs due to electricity usage and such, any down time, etc.

    You will be lucky if you can squeeze more than 200$ of value out of it in a month.

    I don't think people should buy local hardware for money reasons, by the time you will pay off a 10K USD machine, 2-3K USD machine will catch up and beat it by a significant margin.

    Unless your expectation is that we will be in hardware winter for the next 10+ years. At 200$ per month it will take around 200 * 50 = 10k, that is, 50 months, so around 4-5 years.

    Again assuming you are making the most of your hardware somehow, very hard to do in practice.

    I don't recommend people to use compute as investment or payoff thing, but if you have the money to burn and can afford it why not, maybe with some software optimizations it will be cheaper but then again Z.ai is currently offering 50% discount and providers will offer cheaper rates for sure.

    But either way you will never be able to burn more than 200$ worth of token on a cheap hardware device, because inference becomes more profitable the more you scale it up, you have separate prefill and decode engines/systems, and a lot of nuance, but assume for every 10x increase in infra you increase margins by 5-10%.

    So from 10K to 100K to 1M to 10M to 100M.. I don't think this curve continues beyond 100M but I have no idea about that scale unless some AI lab is interested in hiring me lol.

    So a 100M infra will have ~30% better margins than you at 10K, then there is software optimizations but that's cheap enough, though some of it is only viable at scale.

    Either way assume 10K is the price of privacy if you really want to buy it. Don't worry about making the most out of the usage, you will always be in a net loss but I would assume for you 10K doesn't matter.

    • I generally agree — go local for the hobby/tinkering, privacy, and control (ie not getting refused by an AI to defend and secure your own network and codebase; as HuggingFace has seen).

      But whether you make a loss or not depends on how hardware prices and resell values go though.

      I have spent ~$50K on local AI hardware. The market value of that hardware is about ~$80K right now.

      So the maths is working out for me so far. I see it as a call option on compute.

      1 reply →

    • yeah i mostly agree, especially compared to subsidized subscription cost.

      But for a heavy user who has enough work to be done so that the box runs almost 24/7 at say 50tok/sec, the math gets interesting against API prices.

      And it can be interesting compared to subscription in the sense that you don't have the quota anymore. That means there's probably a lot of things you're not doing because of the quotas that you could do now.

      It depends heavily on the tok/sec obviously and the very best solution financially remains subscriptions. But the idea remains entertaining and not that disconnected from reality

      3 replies →

With tiny models surpassing huge, 6 months old models on benchmarks, does anybody have some smart words to share on how these still "feel" different?

Artificial Analysis ranks GPT 5.6 Luna similar to GPT 5.4, but that never matches my real world experience. AA seems to do a good job making a single number as representative as possible but there is still so much benchmarks don't communicate.

  • I agree. They are definitely good - no issues with instruction following for example - but they miss the "intelligence" larger models have.

    For implementation tasks, where I have the problem already defined and researched, or just simple task, I'd definitely use something like Luna xhigh or max. If the task is vague, or involves planning, I'd rather use Sol medium, even though it's theoretically worse on benchmarks.

    • Could do a "Big model for architecture and planning and smaller model (or local model) for implementation" sort of thing

  • One term of art that's emerged for this "feel" is "big model smell", first coined by @aidan_mclau. [0]

    To my surprise I couldn't find any proper explainers of the term in a quick search, despite grokking it after seeing it in various contexts on Twitter, but Fable 5 offered a useful analogy: "A student who memorized worked solutions and one who understands the subject score the same on the test; you can only tell them apart by asking a question the test didn't. Real-world use is nothing but those questions, which is why a single AA number feels right and wrong at the same time."

    In other words, big model smell is related to the underlying ability to "understand" when tasks are underspecified or out-of-distribution. This ability can be mimicked to parity by smaller, distilled models according to the density of the training data for particular tasks, but neural scaling laws still hold for generalized reasoning ability.

    More recently with these smaller models, there's a separate but related "RL-fried" phenomenon, where they rely on CoT to "grind toward a checkable answer even in contexts (open dialogue, taste, ambiguity) where there is no checkable answer, and you get the tell: over-hedged, over-structured, relentlessly on-task, deaf to the subtext."

    There are some other insights and caveats in the (short) conversation that I feel you may appreciate reading. [1]

    [0] https://x.com/aidan_mclau/status/1807843014104211855 [1] https://claude.ai/share/d511a348-7c36-432f-a6d5-9deab2802615

    • Appreciate you taking the time. That fable analogy is well put. Almost obvious once you know it.

> it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

From a biased source, but would be big if true. I've had great results with GLM 5.2.

From their subscription page, the smallest plan gives you about 97M tokens weekly for 5.3 but 292M for 5.3 Flash. Not exactly 10x the limit.

  • > From a biased source, but would be big if true. I've had great results with GLM 5.2.

    It's at least close (even if not better) from the Ox Alpha runs. For the price it's definitely great.

For those who didn't read, this is the identity of the mysterious "Ox Alpha" model

On what hardware do they run this? If I'll visit Shenzen, can I buy these chips? I'd much rather at this stage give money to any-other-manufacturer-than-Nvidia.

I have a use case where I have to run local models, can't share data offsite.

  • Just as expensive unobtanium as nvidia. Maybe even more so.

    Its probably the huawei ascend 910, they are using.

On OpenRouter the pricing is: Input $0,075/M - Output $0,25/M - Cache Read $0,015 /M

How is the business model of Anthropic/OpenAI will sustain?

  • They're obviously in a pickle, nobody is going to continue to pay $15-50 a mm tokens here soon. There's a reason OpenAI stopped training large models last week, and it's not because of "saftey" or "alignment" they know these gigantic models are not worth the squeeze.

Is the actual Z.AI ecosystem good enough to replace the main drivers like Codex and Claude? Because it looks like Z Code is just a Codex fork. Just like the Kimi Code one is.

What irks me about this is that the harnesses seem to be just an afterthought here.

Don't get me wrong, I love messing around with installing Pi, getting it hooked up with OpenRouter, and just trying all kinds of different stuff, local models, etc... but when it comes to literally just setting up a productivity environment and trusting my entire machine with it, I just run Codex.

I have heard from anecdotes where people have indeed replaced their main drivers with DeepSek V4 Flash or GLM and state that "it's almost as good as... [claude/gpt]" but I never hear anyone say "yeah, this is the model/harness that I now run on my machine and don't mess with it"

  • > "yeah, this is the model/harness that I now run on my machine and don't mess with it"

    * me raises hand.-

  • Their list of allowed tools is extensive so just use whatever you want within that list

    Think z code gives a token bonus though

  • I think my inexperience using Claude Code or Codex makes a difference but what would you expect to be different here as opposed to using pi or opencode? Pi is my main driver so switching between all these models is a no brainer. No matter what the model is, my harness stays the same: same workflow, same skills, etc.

  • Both can be true though. I had the max coding plan since january and I kept using with Pi since then, even though it wasn’t as good as opus until glm 5.3. It definitely can be a daily driver if you don’t want to use Anthropic or OpenAI. It’s going to be even better with native vision now available. And i’m not messing with my setup either.

I guess this is also a really good indicator of just how much more tokens people would use when not limited financially. Wonder how much that tells the labs about pricing. Then again, it's 2.3x not 10x the next paid/cheap model.

Good bicycle, good pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

  • It's feeling good on non-pelican workloads too. Less verbose than 5.3, cheaper/higher usage limits, vision capability and some good web design one-shots even with vision disabled.

    With GLM 5.1 and 5.2, the big problem was tool calling and long-horizon coherency. 5.3 was more trustworthy at the cost of longer thinking traces, and now Flash seems to improve on it once again with a more concise, smaller model. As long as there aren't any noticeable regressions, I could see myself defaulting to this for >90% of my day-to-day coding work.

When reading this type of announcements, always have keen eyes on graphs.

e.g. "Agent Coding Performance by Effort Level" cuts Y-axis from 0~20.

- This makes it as if GLM-5.3-Flash made a bigger jump than it claimed as the Y-axis does not increase much (stupid trick used in biz reports)

I did mention that ox was working ok for me, and having an open-weight comparable to close to SOTA makes it very compelling for me to try it out locally (well, only if I got more VRAM)

  • they also conspicuously omitted GPT 5.6 Luna from comparison. It scores lower, but is also cheaper. MiMo 2.5 is not a valid comp at this point

    edit: nevermind. it is there in the artifical analysis scatter plot, but is greyed-out.

    MUCH more interesting is that in that chart, their cost is WAY off. The actual chart shows GLM 5.3 Flash at $0.09, but their chart shows $0.045...

Despite what any benchmarks tell you, I'm actually finding GLM-5.3 max to be better than Sol and Fable. Finally bit the bullet and installed OpenCode and OpenRouter and have been experimenting with other models.

The labs are clearly benchmaxxing a bit to maintain perceptions. But I don't think they're in the lead anymore in terms of their public offering - although I'm sure what they have behind closed doors is far better than anything we're getting access to.

  • I use all three every day, and I am absolutely not finding 5.3 to be better. Competent and in the same league as, sure, but not better.

At the current 50%-off GLM-5.3-Flash price ($0.075/M input, $0.25/M output; cached input $0.015/M), surprisingly, roughly $400–900/month would buy token throughput comparable to fully exhausting Claude Max 20×

  • Probably apples to apples would be to compare z.ai subscription plans vs API pricing

Standard API Pricing for GLM-5.3-Flash (per 1M tokens)

- Input: $0.15 - Output: $0.50 - Cached input: $0.03

Ironically, our administration pushing for ban of the AI chips to China is forcing them to make smaller and more efficient models which seems like a requirement for running on Chinese chips. I wouldn’t be surprised this model was tailored to run purely on Chinese chips. Same thing with Deepseek MLA, the drastically lower KV cache memory requirement was born out of necessity so it runs on the Huawei chips.

  • This is exactly what Jenson said in all of his interviews. Banning it in the short term would have long term consequences.

Tested this last week and couldn't get it to finish any task that took more then an hour with /goal keep getting errors

It's only 320B, local frontier AI is getting closer, sooner than expected.

  • It's not possible to keep shrinking down parameters and keep "frontier" performance, it's like saying it's possible to take a 3 hour movie and compress it down to 3 megabytes, there are information theoretic limits on the amount of bits of information that can be compressed.

    What I'm saying is, if you're expecting a model that can be run on a 16GB or 32GB machine with the intelligence/knowledge of Mythos or Sol, it will never happen. It cannot happen, just like you cannot watch the Odyssey saved as a 16MB file.

    Smaller models can get faster and smarter, but by definition they can never compress all of the knowledge of a frontier model and they will approach a limit by which they cannot get better.

    • The current models are not close to approaching the limit of compression for intelligence. They aren’t even focused on it like Chinese labs are. The training of Qwen’s 27B parameter model showed that by structuring model training from fundamentals to more difficult topics they were able to drastically reduce the number of parameters needed.

      The ‘frontier’ models rely on scale to achieve their results but that’s not the only approach. Eventually we will hit up against the fundamental limits but we are not close with Sol and Mythos.

      4 replies →

    • You heard of JEPA? LLM's have all sorts of garbage they have memorized. Reasoning in latent space instead of in text significantly reduces the number of needed parameters.

      2 replies →

The key difference between this and all other GLM models is it's multimodal. You cannot send images to the other GLM models.

  • I really wish GLM models had vision capabilities. I've worked around that in the past to use a vision MCP in my harness that GLM can call. It is not the same, but it allows the model to query images.

GLM 5.3 Flash: 320B parameters with 18B activated

Qwen 3.8 Next Flash: 125B + 51B = 176B parameters with 6B activated

DeepSeek V4 Flash: 284B with 13B activated

The new Qwen model is the most promising for one or two Strix Halo 128GB with the low number of active parameters. On paper it's much stronger than Qwen 3.8 27B.

Does the word "flash" mean a specific thing when it comes to LLM models? I noticed that this word is used by gemini, qwen, and z.ai and I'm curious does it mean the same thing for each one, or did they all just accidentally brand similarly?

  • It seems like it has come to mean "fast, small, cheap" models these days, and seems well enough understood as such that different AI labs are adopting it.

> Combined with our latest 30T-token multimodal pre-training corpus [...]

Is the optimal formula still 20x the amount of model params in tokens for training? Could this mean we're getting a GLM with 1.5t params?

When do Chinese models surpass US models? I thought there was at least be a 2 year runway but now I think they surpass it within 12 months, if not sooner.

  • I get the impression that they're focusing more on efficiency than raw intelligence. While I assume that all AI labs realize that the AGI "race" is mostly bs, the Chinese labs aren't stuck in a trap where they need to keep blowing money to maintain an intelligence lead to justify investments and valuations.

    So, while OAI/Ant have to fearmonger and keep training the largest models, Chinese labs can focus on efficiency more heavily and as long as they stay near frontier, they'll continue to get positive coverage.

Will we need all the data centers being built or will improvements in software and hardware allow the majority of AI workloads to run locally or in the cloud but way more efficiently than was projected when all the plans were laid out?

Like were executive at Google and AWS and Microsoft expecting this kind of performance from models smaller than what openai/anthropic have been doing? Are we really in a "compute desert"?

  • If it gets more efficient it'll be more enticing to expand use case. Personally I'm hoping to do a lot at home but I'm not counting the datacenter building as a bad move at this moment. It may and up that way.

I'm starting to think that this whole sanctioning China may motivate and prompt them to do more and better in every field.

It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.

  • Starting? This was obvious way back in 2019, when the US decided to give China a little push developing their own silicon industry.

  • > It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.

    It's not like we didn't try it. China first have to learn to make deals where both party benefits.

  • This has been clearly stated as what would happen going back several decades at least.

  • Well the big problem with china is that they do not respect international law when it comes to technology theft. But that argument is very weak when it appears that a lot of what they do is out in the open for anyone to replicate.

    • No major power respects nor cares about international law.

      Intellectual property is part of WTO agreements but enforcement is domestic.

      US companies do it too, regularly, they simply hire and poach staff from competitors.

      Proving it to be IP theft is difficult unless you can prove documents being passed. But often all you need is the know-how of the hired talent.

    • There isn't one global "international law" for copyright. There are treaties that countries negotiate with each other.

      If the USA wanted a copyright treaty with China bad enough, we would negotiate one. China is not breaking any laws here, international or otherwise.

I was quite surprised that Zai had deep pockets to serve this free for a week. My first guess was this was an American lab like xai or google

Benchmarking is cool, but for production I care about real inference latency, self-hosting VRAM costs, and how cleanly it handles structured JSON output.

tbh I wasnt that impressed by it. initial benchmarks were trying to say it was AGI but i told it to re-build Palantir in 1 pass and it gave me a non working prototype

from the article, pareto frontier for open source models is completely dominated by GLM now.

  • Well, it will be interesting to see where Qwen3.8-Flash-Next ends up landing, also released today. These are exciting times!

  • I find GLM's idea of fast/flash is not really competitive with the speed DS4 Flash has, and it's hard to see them as being in the same segment for that reason.

Is anyone actually tried it in agentic coding (claude code loops)? Are apple silicon macs (M5 Max) capable of working with that model? what was the tps?

  • I've just used it for a fairly complex refactoring of the UI in a SwiftUI / AppKit app. It managed the refactoring in blazing colors, and the resulting UI looked really good. It was also quite fast. I'm impressed.

Luckily, I have the coding plan for z.ai, so I'm happy with this model as I always kept running out of usage with the original glm-5.3

offtopic: Is there any chance we could see competing models from other countries in the next 5 years?

  • Chinese universities are really a huge advantage, even in the US many of the top staff in model development are Chinese. Another big thing is the hardware costs required to train models. Between those two factors it really looks like this will remain a US-China competition for the foreseeable future, although there are some other players like Mistral from France.

In my brief testing, it did about as well as Qwen3.8-4B-Distill, and LFM2.5-2.6B overtook both.

Between Gemma 31/26/12/4/2, Deepseek-v4-flash-0731, Qwen 3.8 27B, Qwen 3.8 Flash Next (which I haven't even gotten to run yet!), and now GLM 5.3 Flash, I can't keep up. I love all these open weight models and am continually stunned that it's largely the West fighting for closed, restrictive, anti-user bullshit and China absolutely mogging the likes of OpenAI and Anthropic, with some notable exceptions like Gemma. Still, I shudder to think what the world would look like if we only had closed models. In many ways the stagnation of open source diffusion seems like that: LLMs are just a few months behind frontier, but image gen is like 1.5 years behind.

i wonder if more companies will now stealth launch their models. imagine they just released this on openrouter for free but under their normal name - would they get the records in token usage then?

Holy shit, is this model really that bad???

Just asked it a question via the custom opencode go endpoint routed over cloudflare ai gateway doesnt show me the correct models. My fault was that I set "https://opencode.ai/zen/go" as endpoint and tried my-gateway.com/custom-ocgo/v1/models, turns out I had to add /v1 to the opencode url and leave it on my -gateway.com.

But first it told me that opencode is not on the compat endpoint and than it recommendet "Option 3: Nutzung seit Curse die Integration des Providers класный" . Not to mention some complete gibberish like "Falls Opencode.ai in Cloudflare Eing? Wenn ja, wähle den eingebauten provider o.ä. Weil der Gateway dann ein eigenes Modell-List gibt." or "Schlage die Modellnamen in einer Liste ab (z.B. lokal fester Key)"

Is this just cloudflare or is it really that bad? I mean thats not even the level of LLama2 7b ...

  • > I mean thats not even the level of LLama2 7b ...

    I think you answered your own question. Must be something wrong.

    • I don't know yet. May be its a trade of for agentic skills in favor of language skills. GLM5.3 feels also heavily benchmaxxed compared to 5.2.

  • I actually noticed this too with DS flash over openrouter. Instructions to respond in my native language cause the model to spew gibberish roughly resembling my native language mixed up with similar languages, and it was also dropping in Cyrillic characters randomly (although i noticed this with the anthropic models as well)

I heard that Dario Amodei is not having a great day today.

2 really strong open models on the same day is a amazing.

This was the ox-alpha model, right? I remember it performed really well for a model that had 'flash' in its name.

> To overcome the relatively limited compute and memory capacity of individual chips, we built a dedicated inference engine for this architecture on top of SGLang. Notably, this effort was accelerated by our GLM-5.3-powered infrastructure agent, which assisted engineers in developing and optimizing kernels, diagnosing performance bottlenecks, and improving the serving stack — creating a feedback loop in which the model helped optimize the system serving the model itself.

> (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.

It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.

  • Now translate this to physical world, robots building and optimizing other robots... getting iRobot (2004) vibes

Anthropic is accelerating their IPO cause they know what's coming in the next 5 years.

I didn't accept a single edit from this model over the entire week, just saying. I do not understand how it's being benchmarked on par with Sol and other larger models.

Has anyone ever asked themselves why AI was made publically available in the first place? is it really economics or is it about training people to recognize the patterns of machine generated words and ideas?