Comment by NooneAtAll3

11 hours ago

> In adversarial settings (where we push the model to evade our monitors)

...why exactly are they training for that?

presumably that's a safety evaluation not a training setting

  • The whole Huggingface attack happened during training runs

    • part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.

Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions