Comment by Iolaum
12 hours ago
As long as they are transparent about what quant they serve the model and any other optimization they do that also affects performance of inferred tokens.
12 hours ago
As long as they are transparent about what quant they serve the model and any other optimization they do that also affects performance of inferred tokens.
No comments yet
Contribute on Hacker News ↗