Comment by SillyUsername
4 hours ago
Yep more hops from the lower Q is likely going to skew the vectors further over time.
I wonder if there's a way to mitigate this by running it through an original Q8 draft model, attuned somehow for the PTQ1 quant, but giving it a higher threshold for the acceptance linear with the context length itself?
The longer the context, the higher the multiplier on the threshold, and more likely the draft result is used. Not ideal but it may extend the usable max context.
This model might, even without this, be amazing for short lived agents that work via generations / have changing tasks.
No comments yet
Contribute on Hacker News ↗