Comment by oezi

2 days ago

I am really suprised that they do not start putting up the same signs you would for humans to prevent unauthorized access:

  Keep out. If you can read this sign you are off track. Leave now.

I mean, how are the agents to know that they are overreaching if they just get cache miss or 404.

From the conversation log and CoT you also get the impression that the RLHF has been overdone. The agents seem really obsessed to obtain the answer and understanding motive ('it could be browsercomp').

Look into nuclear semiotics. You can't say "this area is dangerous" and expect people to stay out.

  • When you're giving the orders to those people, and they're generally trying to obey, you can.

What they should do, is inject a system prompt at that point telling the model to get out of there. It is the entire reason they train the different levels into the template.