Comment by igravious
11 hours ago
I call bullshit on the diverging nature of the dashed lines in this info-chart https://www.nist.gov/sites/default/files/styles/1400_x_1400_...
these two claims can't be true at one and the same time:
(a) they're distilling our secret sauce!
(b) they'll never catch us!
I would read the divergence as mostly evidence that Mythos et al, had offensive cyber capabilities as part of their RL training. I.e. they’re models specifically trained to be good at offensive cyber, rather than being general purpose models that happened to become good at offensive cyber via emergent behaviour from sheer scale.
The fact the Opus 5 seems to be as capable as Fable/Mythos on everything except cyber, and Anthropic explicitly say they removed all offensive cyber training data, I think lends further credence to the idea that Mythos was designed from day zero to excel at offensive cyber capabilities.
If that’s true, then we would expect divergence in open models of their capabilities come from distillation, no frontier class cyber capable model has seen significant public availability. Which means there simply isn’t data to distill from.
It also calls into question the entire narrative around Mythos capabilities being a complete surprise for Anthropic, and an inevitable outcome of scaling up LLMs.
Statements of the form “X can never happen” tend to be weak, so that’s a bit of a strawman. But, in spirit, of course they can both be true. There are matters of degree and the intervention of countermeasures.