Comment by oezi
2 days ago
I am really suprised that they do not start putting up the same signs you would for humans to prevent unauthorized access:
Keep out. If you can read this sign you are off track. Leave now.
I mean, how are the agents to know that they are overreaching if they just get cache miss or 404.
From the conversation log and CoT you also get the impression that the RLHF has been overdone. The agents seem really obsessed to obtain the answer and understanding motive ('it could be browsercomp').
Look into nuclear semiotics. You can't say "this area is dangerous" and expect people to stay out.
When you're giving the orders to those people, and they're generally trying to obey, you can.
an AI alignment researcher said they ran experiments testing exactly this, the model reasoned "this note is not for us. proceed".
Well, agents are intentionally trained with RLHF to solve captchas that say “prove you’re human”.
Link? Name?
01:34:00 https://youtu.be/EimoamE3mTI
5 replies →
Those signs don’t prevent people from entering…
What they should do, is inject a system prompt at that point telling the model to get out of there. It is the entire reason they train the different levels into the template.