← Back to context

Comment by danielklnstein

12 hours ago

Works much better now! Got 103.9 tok/s, not quite 200 - but still amazing! Thanks for sharing

Something a lot of model providers don't talk about: any time an engine uses speculative decoding the throughput will depend on how much your output token distribution matches what the draft model was trained on.

The DFlash2 draft model we're using was trained on a lot of code, so if you use it in a coding agent you'll probably notice it run a lot faster (we've seen it break 300 tok/s).