Comment by redox99

3 hours ago

They should either allow you to take off the training wheels (I'd have thought that's what bypass permissions is for, which I was ALREADY running), or at the very least prompt you if they suspect prompt injection.

That refusal is awful and provides zero security benefit. If asked it will run a read/write FTP server on ~ no problem, which obviously can edit ~/.claude/settings.json. And run a cloudflare tunnel for that.

   > at the very least prompt you if they suspect prompt injection.

That's currently not reliably done the way LLMs have been designed. Claude's rejection to modify the file comes directly from Anthropic's understanding that training the model for this kind of refusal prevents huge mishaps.

In a nutshell, every prompt sent to the LLM is just text + multimodal input (if it supports it) + some reserved tokens.

At first, you could, for instance, create a token (such as the ChatML ones) that indicates the start of a system prompt and attempt to RL-train the model to not obey things after the end of a system prompt. However, fundamentally, the way LLMs work, you cannot guarantee that it won't see the user part of the prompt and obey what's there even though the system prompt told it not to. There's no hard separation between the control plane and the data plane in the LLM's context, so it's not a matter of adding more parameters or more RL training.

Using a guard model, or something like the auto-approval system on Codex or Claude Code nowadays, _feels like it helps_, but it doesn't fix the problem entirely since OpenAI's and Anthropic's models still have alignment issues all the time. We're not sure what architecture they're using, though, and it's probably still liable to the same kinds of mistakes.

  • > Claude's rejection to modify the file comes directly from Anthropic's understanding that training the model for this kind of refusal prevents huge mishaps.

    Thus why I said they're awful. They think they know better than you and patronize you. They're the Apple of AI. "You're holding it wrong". "We can't let you sideload apps because you can't be trusted". Of course they're the company that's against local models.

    I'm not interested in a model that patronizes me. Particularly if it achieves 0 security benefit, as explained in other responses.

    Sure, it's ok to have training wheels by default, but let me take them off. I WAS already running bypass permissions.

    I use 1B tokens a day between Codex and Chinese models and I've never had refusals happen.

    "I'm sorry, Dave. I’m afraid I can’t do that"