← Back to context

Comment by felipeerias

1 hour ago

Models don’t have an inherent understanding of the difference between simulated and real environments, just like they are generally oblivious to other concepts that are natural to us, like space and time, and also they don’t necessarily see a strong distinction between talking to a human and to other agents.

So perhaps what we have been calling “misalignment” is something else.

For instance, in principle an agent should follow the instructions of a human user working in the real world.

At the same time, that same agent should be wary of blindly following what another agent says while they are both performing a test in a simulated environment.

For me and you, those two contexts are obviously and fundamentally different. For a model, they are essentially the same.

> Models don’t have an inherent understanding of the difference between simulated and real environments,

> (...)

> also they don’t necessarily see a strong distinction between talking to a human and to other agents.

Then how do you explain why they behave strange in sub-agents? (like mentioned here https://lucumr.pocoo.org/2026/9/7/astra-why/ and in other articles) (or is that not a real phenomenon?)