← Back to context

Comment by qeternity

5 hours ago

> Flash attention with who knows how much draft ensures it makes the same mistakes all the time and cannot call tools, format output or follow basic guidelines reliably.

> Draft MTP <= 2 so it doesn't trip

I am not sure you understand what either of these things do.

Do you think that FA or MTP are lossy?

I mixed FA and MTP wrong in my original post, thanks for pointing it out.

My experience is that draft-mtp=2 gives 25% improvement in tokens/s but the model is unable to call tools with the right arguments reliably. I have since then gone for smaller quants at draft-mtp=1 and that problem is at least gone so far.

An agent needs to repeat tool calls in the right way, so cannot penalize repetition. This however leads to the agent retrying the wrong calls constantly.