← Back to context

Comment by johndough

1 hour ago

> There is a larger 120B that you can't realistically run on consumer hardware at reasonable tok/s too.

gpt-oss-120b runs at 30+ tps on Strix Halo and +75 tps on a MacBook Pro M5 Max 128GB.

> I wish OpenAI updated these models more frequently though.

I think the spiritual successor is the Nemotron 3 series, although they also are getting a bit long in the tooth: https://research.nvidia.com/labs/nemotron/Nemotron-3/

The Gemma 4 models are a bit more up-to-date: https://huggingface.co/collections/google/gemma-4

Or Qwen3.6: https://huggingface.co/collections/Qwen/qwen36