Comment by doodlesdev

3 hours ago

It makes a lot of sense for the model to not be able to do it, whether you like it or not. In fact, it shows that Anthropic, despite all of their issues, are paying some attention to the risks of malicious prompt injection and models attempting to bypass restrictions.

If it could do it whenever you ask it to, it could also do it unprompted or by finding a file in your directory that told it to do it, which would make the entire permission system useless...

This is not prompt injection. This is a prompt entered by a human through the Claude UI.

Being unable to perform an action is not a solution to prompt injection. A solution to prompt injection is being able to tell apart what is the real input and what is injected. I expect it to follow whatever I typed into it, and not blindly follow what it read from a file or an external source.

If they are not confident in their ability to do so, at least allow to remove the training wheels so people who know what they are doing and the risks are not patronized by the model. But you don't even get a confirmation box to perform that action, it flat out refuses.

It really is like people defending Apple not allowing side loading because you as a user can't be trusted.

  • > This is not prompt injection. This is a prompt entered by a human through the Claude UI.

    Well, to LLMs this is the same thing - an input. Prompt from the user and prompt from the attacker use the same input into the LLM's neural network, so to speak.

    So it makes sense for it to be a bit more paranoid.

    There are other possible architectures probably but for now I think nobody uses them. See e.g. https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/

    • It's not the same thing.

      Messages are already wrapped in developer role, system, user, assistant, tool, etc by special tokens. If you are paranoid you could show a confirmation box, a UAC prompt, etc. Refusing is the worst possible solution.

      3 replies →