Comment by EnPissant
2 days ago
It's not a bug. It's the reality of token generation. It's bottlenecked by memory bandwidth.
Please publish your own benchmarks proving me wrong.
2 days ago
It's not a bug. It's the reality of token generation. It's bottlenecked by memory bandwidth.
Please publish your own benchmarks proving me wrong.
I cannot reproduce your bug on AMD. I'm going to have to conclude this is a vendor issue.