← Back to context

Comment by preommr

15 hours ago

I've had codex delete useful (albeit not directly relevant or perhaps messy wip notes) comments, even though I explicitly have it in my agents.md not to delete comments, and ask for permission if it thinks it should.

It deleted the comments, and when I asked why it did that even though I expressedly asked it not to, it responded that me prompting it in the first place explicit permission. I have no idea if that's the actual reason or just some post-hoc explanation.

But I genuinely don't think it's possible to just have these things be completely, 100%, indpenedent and also solve deep problems that need to also be understood by people in a people-based organization context.

> and when I asked why it did that even though I expressedly asked it not to

Just be clear, it can't know, and by asking you're just making it roleplay as someone excusing themselves.

It's very unlikely that the choice to remove the comment was driven by an internal monologue based on learned criteria that it can refer to. The sampler most likely picked tokens to remove it while writing the patch, because that's what the statistics modelled, and that's it.

There might not be an "actual" reason that's expressible in human-comprehensible language.

  • Moreover, even if there was a reason, it was there "in the model" at the moment of generation.

    The new model can only guess/hallucinate when ordered to give reasons.

    It's like having an actor play a role while wearing a hat. Then the next actor comes, puts on the same hat, and we're asking them to explain why "they" did something while there was someone else playing the role.