Comment by Dylan16807
2 days ago
Modern laptops and desktops are so powerful though. That's not a great way to measure tiny.
With some patience you can run huge models directly out of flash. Edge0–35B-A3B, based on Qwen, has 35 billion parameters and will do 15 tokens per second on pretty boring hardware. If you treat Kimi K3 with 2.8 trillion parameters similarly (importantly, cutting the number of simultaneous experts in half) you can get the active weights under 12GB and get just over 1 token per second on a good laptop. Models in between have speeds in between.
The terminology does seem to be all over the place. GPT-2 is definitely a "small" LLM at best, but definitions of "tiny" might cap out at 100M or at 3B...
I think if I wanted to involve computer capability, I would say tiny is what you can train from scratch on a laptop. Not just run.
No comments yet
Contribute on Hacker News ↗