← Back to context

Comment by tesnorindian

22 days ago

The trouble here is they will stop using the official channel and will start communicating secretly rendering our honeypot useless.

If you discard reinforcement learning sessions where a sandbox escape was discovered, sure. Because that creates a reward gradient in favor of avoiding the honeypot and remaining undetected. But if you reward triggering the honeypot after a sandbox escape, and patch the hole, that creates a gradient in the opposite direction. Because then detectability is adaptive.