Comment by lebovic
2 hours ago
> all frontier models benefit from more tokens not just Kimi K3
Past a point, that doesn't hold and the score plateaus.
Token hungry models tend to plateau at a much higher token count. Because Kimi K3 is a token hungry model – and 100M tokens (including cache hits) seems at the edge of the plateau for these evals – it could disproportionately benefit from a higher token budget.
For Kimi K3 specificially, policymakers are interested in whether it can find and exploit the same scope of vulnerabilities as models like Mythos. In that context, an answer of "yes, but with quintuple the token budget" is materially different from "no, it performs significantly below the most recent frontier cyber-capable models".
(As an aside, I like the UK AISI and think they're the best example of that kind of group!)
No comments yet
Contribute on Hacker News ↗