Comment by johndough
2 hours ago
> There is a larger 120B that you can't realistically run on consumer hardware at reasonable tok/s too.
gpt-oss-120b runs at 30+ tps on Strix Halo and +75 tps on a MacBook Pro M5 Max 128GB.
> I wish OpenAI updated these models more frequently though.
I think the spiritual successor is the Nemotron 3 series, although they also are getting a bit long in the tooth: https://research.nvidia.com/labs/nemotron/Nemotron-3/
The Gemma 4 models are a bit more up-to-date: https://huggingface.co/collections/google/gemma-4
Or Qwen3.6: https://huggingface.co/collections/Qwen/qwen36
No comments yet
Contribute on Hacker News ↗