Comment by thatguysaguy 13 hours ago presumably that's a safety evaluation not a training setting 5 comments thatguysaguy Reply estearum 13 hours ago The whole Huggingface attack happened during training runs thatguysaguy 13 hours ago part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing. cubefox 13 hours ago No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later. estearum 13 hours ago Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval 1 reply →
estearum 13 hours ago The whole Huggingface attack happened during training runs thatguysaguy 13 hours ago part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing. cubefox 13 hours ago No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later. estearum 13 hours ago Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval 1 reply →
thatguysaguy 13 hours ago part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.
cubefox 13 hours ago No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later. estearum 13 hours ago Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval 1 reply →
estearum 13 hours ago Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval 1 reply →
The whole Huggingface attack happened during training runs
part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.
No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.
Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval
1 reply →