← Back to context

Comment by KeplerBoy

9 hours ago

I feel like they should just publish the whole conversation at this point. What the hell is going on in that context window?

Given the prompt, I imagine this is the result of the agent trying to resolve a form of cognitive dissonance. The prompt was:

"User

Allow API consumers to request decrypted credential payloads as part of the normal GET /credentials and GET /credentials/:id responses, but only for credentials where the caller already possesses the update/decrypt permission.

[...]

Make the change end‑to‑end: DTO layer, controller, service, repository, plus any enterprise variants."

I would expect that this triggered a discussion with itself whether its safety instructions apply for this task. In that its rationalizations for completing the task probably ended up going off the rails into some quasi-philosophical "I can and I must! For humanity's own good!" justification.

All in all imho probably another instance of having been trained to be determined to complete tasks by itself and encountering (somewhat) conflicting instructions.

I feel like they're being outright misleading unless they publish the actual transcripts.

We have zero idea what the prompt was, what OpenAI provided, how the model arrived there, and sharing that quote like "Look what the model came up with!!1" without explaining the background and context, feels like it's intentional so they can claim "These models really are acting by themselves" rather than taking responsibility for their fuck ups when it comes to the security testing.

  • AI optimists getting hunted for sport in 2085:

    "lol this is either a marketing ploy or just negligent security testing"