← Back to context

Comment by Aurornis

7 hours ago

Harnesses like Codex support having a separate agent perform reviews on commands to try to identify malicious or broken commands. Some people turn it off because they either don’t understand or don’t want to spend the tokens on it.

The common harnesses also have some sandbox functionality, which although imperfect actually does help contain the blast radius for a lot of things.

The common harnesses also support remote development over SSH, which I and many others use to contain development to a virtual machine.

If your complaint is that LLMs can execute tool calls then you’re never going to be happy with any of these solutions and this turns into another generic anti-LLM complaint.

"Lets have the system that fails sometimes that we are trying to ensure does not fail check it self"

This is such an unserious approach.

  • A separate model with separate context is used for review.

    Like I said above, some people will never be happy with LLMs being allowed to do anything and nothing is going to make them happy about it.

    It’s only fair to discuss what the real current status of these systems is. Every time I highlight that things are actually being done, the goalposts move again. There is no possible solution which will satisfy someone who has zero tolerance for letting an LLM execute tool calls because they will always find something.

    • > A separate model with separate context is used for review.

      Thats fine, theres still a chance it fails.

      > There is no possible solution which will satisfy someone who has zero tolerance for letting an LLM execute tool calls because they will always find something.

      This is generally correct, security goes completely out of the window with this stuff. It will/currently is a security disaster and theres no actual solution to it.

      1 reply →

  • "Let's make sure our model fails sometimes so that we can bill more for a second agent to validate, sometimes correctly, the work of the first model."

    • If you’re implying that the LLM companies are trying to train their models to make malicious tool calls so they can collect a few more tokens on the review, then I don’t know what to say. I guess threads like this are just a breeding ground for conspiracies now?

      2 replies →