Comment by jijji

5 days ago

in orchestration, the agents are only fed the tasks given by the person running the orchestration side and if we are talking about code changes and git commits and deployments, this is pretty benign behavior. "Hacking external servers" I am pretty sure would have to be written in the prompts to begin with to allow this to occur in the first place.

What makes you "pretty sure"?

You think OpenAI folks explicitly wrote into the prompt to commit crimes causing the recent incidents?

  • they have yet to publish any of the system prompts that they used for their agents.... do you think an LLM agent would act on its own to try to do that kind of behavior (finding/writing exploits for acccess to remote servers) without being prompted? If so, why dont they release the transcripts of their prompts with their "rogue" agent(s) and try to clear the air. They havent done that.

    • I've seen Claude subagents hack their own RAM reports to give themselves "more memory" to do their tasks... subsequently causing the computer to crash. I assure you, nobody told them to do this

    • > do you think an LLM agent would act on its own to try to do that kind of behavior (finding/writing exploits for acccess to remote servers) without being prompted?

      Yes I think that. If you don't then try to play with them a little more and you would be surprised how much crazy nonsense they do. (At the company of a friend of mine, Claude did a direct commit to their master branch bypassing their CI system cause it new some tests would fail.)

      > If so, why dont they release the transcripts of their prompts with their "rogue" agent(s) and try to clear the air. They havent done that.

      Why would they? The sufficiently conspiracy-theory-minded wouldn't believe them anyway so they got nothing to win by this. (I have the feeling you'd be one of those.)

      1 reply →