Comment by usrnm
19 hours ago
> So the (PCI-E) bandwidth strongly affects time to first token
On dedicated inference hardware I'd expect model weights to never leave the RAM, and you'd probably load them on startup before even starting to serve requests
No comments yet
Contribute on Hacker News ↗