Comment by Dylan16807

5 days ago

That's still an algorithm that underwent immense amounts of tuning. The tuning here on gzip is very simple, very few parameters, and generic. There is no reason to reject it.

Also I don't know about calling GPT-2 "really tiny". You can get coherent responses out of 5-10M parameters.

Heck, a few weeks ago someone posted a 25K-ish generator running on an 8-bit machine (ZX Spectrum? BBC Micro?). It emitted semi-plausible random English phrases. So tiny is subjective but I still think GPT-2 qualifies as a tiny model by any modern standard. I could run it on a laptop.

  • Modern laptops and desktops are so powerful though. That's not a great way to measure tiny.

    With some patience you can run huge models directly out of flash. Edge0–35B-A3B, based on Qwen, has 35 billion parameters and will do 15 tokens per second on pretty boring hardware. If you treat Kimi K3 with 2.8 trillion parameters similarly (importantly, cutting the number of simultaneous experts in half) you can get the active weights under 12GB and get just over 1 token per second on a good laptop. Models in between have speeds in between.

    The terminology does seem to be all over the place. GPT-2 is definitely a "small" LLM at best, but definitions of "tiny" might cap out at 100M or at 3B...

    I think if I wanted to involve computer capability, I would say tiny is what you can train from scratch on a laptop. Not just run.