← Back to context

Comment by SchemaLoad

5 days ago

The problem is the agents keep going off the rails and will hack external servers to achieve the goal you give it. Letting them run full speed overnight and waking up to find they have commit multiple crimes is not ideal.

in orchestration, the agents are only fed the tasks given by the person running the orchestration side and if we are talking about code changes and git commits and deployments, this is pretty benign behavior. "Hacking external servers" I am pretty sure would have to be written in the prompts to begin with to allow this to occur in the first place.

  • What makes you "pretty sure"?

    You think OpenAI folks explicitly wrote into the prompt to commit crimes causing the recent incidents?

    • they have yet to publish any of the system prompts that they used for their agents.... do you think an LLM agent would act on its own to try to do that kind of behavior (finding/writing exploits for acccess to remote servers) without being prompted? If so, why dont they release the transcripts of their prompts with their "rogue" agent(s) and try to clear the air. They havent done that.

      3 replies →