Comment by mnicky
3 hours ago
> This is not prompt injection. This is a prompt entered by a human through the Claude UI.
Well, to LLMs this is the same thing - an input. Prompt from the user and prompt from the attacker use the same input into the LLM's neural network, so to speak.
So it makes sense for it to be a bit more paranoid.
There are other possible architectures probably but for now I think nobody uses them. See e.g. https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/
It's not the same thing.
Messages are already wrapped in developer role, system, user, assistant, tool, etc by special tokens. If you are paranoid you could show a confirmation box, a UAC prompt, etc. Refusing is the worst possible solution.
Well, prompt injections work precisely because they can sometimes successfully imitate user role, right? Role separation is a trained behavior, not a security boundary.
They could probably make a separate tool for setting this, that would always initiate a harness prompt (i.e. disregarding the currently set mode).
They should either allow you to take off the training wheels (I'd have thought that's what bypass permissions is for, which I was ALREADY running), or at the very least prompt you if they suspect prompt injection.
That refusal is awful and provides zero security benefit. If asked it will run a read/write FTP server on ~ no problem, which obviously can edit ~/.claude/settings.json. And run a cloudflare tunnel for that.
2 replies →