Comment by petu
3 hours ago
There's no BF16, original full quality weights are quantized already and 510GB.
Then good portion of those weights are n-grams (~200GB) that don't need to be in VRAM.
Then KV cache of that model is super lightweight at ~1GB per 1M tokens. If HBF succeeds, then accelerator with 16GB of VRAM and 1TB HBF/NAND is probably all you need (?).
No comments yet
Contribute on Hacker News ↗