← Back to context

Comment by andai

24 days ago

Halfway thru the article it shows a comparison with several frontier-ish LLMs. But they're all from half a year ago. "Our new model is better than all these Chinese models from 3 generations ago" is pretty funny to me.

It’s a 6bn model. Totally different class. I’m more excited about “frontier small language models” tbh.

  • It's a 119B model, 6B active.

    That's still 3-10x smaller than the other models in that graph though (400B, 1T, 1.5T).