Comment by drakythe
10 hours ago
I am also interesting knowing how they determined the model was sandbagging rather than just making a poor decision.
Also, this paragraph makes me wonder about all their stats on the exploitation and misalignment charts. If the model is that good at hiding "incriminating information" and sandbagging, are they sure its alignment is that?
No comments yet
Contribute on Hacker News ↗