Comment by seemaze
19 hours ago
As they say, time is money.
In the age of the rampocalypse, the peasants may not have a choice between the two.. time it is!
19 hours ago
As they say, time is money.
In the age of the rampocalypse, the peasants may not have a choice between the two.. time it is!
Smaller models are also generally faster, so thinking "more" may not matter and may even come out ahead.
If Q4 takes less than 1.3x as many tokens as bf16 or q8, it could still end up being faster, given how decode tends to be bandwidth bound. The kv cache was still bf16, so a few ops are the same between quants.
Computer science has known the tradeoffs between memory and compute since ages ago. The same could be reflected here.