Comment by DiabloD3
4 hours ago
A lot of this is managed by the inference engine, and has nothing to do with the model.
Models that use, for example, sparse attention mechanisms are just trying to make the bad situation slightly less bad, such as using less RAM for context (thus requiring less context quantization) or using less bandwidth (thus running faster).
If people keep using temp, top-k, top-p, and min-p, and nothing else for samplers, we're ignoring ~3 years of sampling research that virtually eliminates the worst of context rot issues.
No comments yet
Contribute on Hacker News ↗