Comment by alignmeharder
2 days ago
an AI alignment researcher said they ran experiments testing exactly this, the model reasoned "this note is not for us. proceed".
2 days ago
an AI alignment researcher said they ran experiments testing exactly this, the model reasoned "this note is not for us. proceed".
Well, agents are intentionally trained with RLHF to solve captchas that say “prove you’re human”.
Link? Name?
01:34:00 https://youtu.be/EimoamE3mTI
Okay, that's a very interesting look at the difficulties. Thanks for the link.
But the example you have isn't quite that bad. Yes the models are way too likely to rationalize their way into bad actions in the pursuit of achieving their task, and it's hard to figure out how to fix that. But the example of "that must not be for me" was a tool call, not exceeding access, and it only did that after they specifically trained it that failing that tool call was good.
2 replies →
so, what would make an LLM choose to ignore one prompt while in the same run, also over-fixating on another prompt, to the extent (as claimed in that video segment), it chooses to ignore prompts?
they talk about it like there's a "wanting" in there, that is distinct from both the original prompt, as the steering/warning prompt
if that's true, it would be very interesting, but if it's not, that would also be very interesting and even helpful
1 reply →