← Back to context

Comment by Aurornis

2 days ago

Ox Alpha is a smaller model and it was running very slowly. Chinese AI accelerators are coming along, but nVidia’s lead is huge.

Lead doesn't really matter anymore. I just ported a very old cuda library to rocm, so it can be run on MI300s. 2 years ago this would have been a nightmare. Today it was an afternoon.

It was being served for free. They were almost certainly being overloaded.

  • Presumably the efficiency numbers they're quoting are for the high concurrency state they were serving.

    RAM was probably the bottleneck for the amount of context they were offering.

    I assume it would run a little faster with lower concurrency but "RIP nVidia" is a little premature. The cutting edge inference hardware is amazingly powerful

Ox Alpha was also serving 10T+ tokens a day for free.

When it first launched on OpenRouter I was getting nearly 70 Tokens/second.

> and it was running very slowly

... I'm at a loss for words here. It was being served for free. To the entire world.

  • GPT-5.6 Luna is also served for free to the entire world with a tokens per second rate nearly 10X higher.

    > ... I'm at a loss for words here

    No need to be so dramatic. I think it's great that they're developing chips, but the whole "RIP nVidia" claim was overly dramatic.

Has there been any confirmation about what that model even is?

Edit: Ah:

> This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash.

  • It's also in this very announcement, in the first paragraph:

    > Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.