Comment by rkagerer

5 days ago

> Dots use auto-review[1] to check actions that could affect your accounts or share information against your instructions

So if I've got this right, their security relies on other agents that sit at the boundary and sentry whether a proposed action is allowed.

This means they have to interpret the purpose of the action, what effect it will have, whether those two things align, and what is the potential risk / splash zone for collateral damage.

Sorry, but all the evidence I've seen points to their models being nowhere near good enough to do this reliably, consistently and responsibly.

The architecture also feels ripe for becoming a cat and mouse game between the 'competing' agents. It's already pretty easy to see how humans are manipulating their AI to bypass the baked-in restrictions.

[1] https://learn.chatgpt.com/docs/sandboxing/auto-review