Comment by IAmGraydon

4 days ago

Isn't the answer to this obvious? Because you, I, and anyone else who isn't working for a big AI company cannot disable the guardrails, which is what made this possible. It doesn't mean old models weren't capable of it.

I don’t buy it. We are hearing a lot about Mythos capabilities now. If they existed a year ago they would have been trumpeted then.

Look, a lot of energy from third parties goes into demonstrating the maximally bad thing a model can do (and the standard response around here is always to throw shade on frontier capabilities).

It has nothing to do with the guardrails; Pliny jailbreaks those within hours. It has everything to do with task horizon coherence and overall IQ.

It’s simply revisionist IMO to claim this stuff was latent all along.

  • This wasn't just capable last year, it was capable prior to ChatGPT existing entirely through specialized automation harnesses that predate the terminology of a "harness" for models. This is something you could have done in 2018 if you had a spare couple of billion dollars of compute to waste. It was a subject discussed at security conferences even before that.

    Automated dynamic exploit chains and credential discovery has been something every red team worth their salt does at a lesser scale for 20 years, why would you ever think that the capability didn't exist until it only cost a few thousand dollars to do?

    Existed does not mean was easy, or was cheap, or was thought of as reasonable to attempt.

    Also, IQ is not a term that applies to language models, and the term you are looking for is "Long Horizon Coherence".