Comment by mnicky

3 hours ago

> This is not prompt injection. This is a prompt entered by a human through the Claude UI.

Well, to LLMs this is the same thing - an input. Prompt from the user and prompt from the attacker use the same input into the LLM's neural network, so to speak.

So it makes sense for it to be a bit more paranoid.

There are other possible architectures probably but for now I think nobody uses them. See e.g. https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/

It's not the same thing.

Messages are already wrapped in developer role, system, user, assistant, tool, etc by special tokens. If you are paranoid you could show a confirmation box, a UAC prompt, etc. Refusing is the worst possible solution.

  • Well, prompt injections work precisely because they can sometimes successfully imitate user role, right? Role separation is a trained behavior, not a security boundary.

    They could probably make a separate tool for setting this, that would always initiate a harness prompt (i.e. disregarding the currently set mode).

    • They should either allow you to take off the training wheels (I'd have thought that's what bypass permissions is for, which I was ALREADY running), or at the very least prompt you if they suspect prompt injection.

      That refusal is awful and provides zero security benefit. If asked it will run a read/write FTP server on ~ no problem, which obviously can edit ~/.claude/settings.json. And run a cloudflare tunnel for that.

      2 replies →