← Back to context

Comment by thatguysaguy

13 hours ago

presumably that's a safety evaluation not a training setting

The whole Huggingface attack happened during training runs

  • part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.

  • No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.