Comment by drakythe

11 hours ago

I am also interesting knowing how they determined the model was sandbagging rather than just making a poor decision.

Also, this paragraph makes me wonder about all their stats on the exploitation and misalignment charts. If the model is that good at hiding "incriminating information" and sandbagging, are they sure its alignment is that?