Comment by Tade0
4 hours ago
To add to the other comment, there's also Ridge quantisation - the majority of weights are indeed Q3_x, but the most sensitive layers are FP8.
4 hours ago
To add to the other comment, there's also Ridge quantisation - the majority of weights are indeed Q3_x, but the most sensitive layers are FP8.
No comments yet
Contribute on Hacker News ↗