← Back to context

Comment by ohaodha

1 hour ago

I find it difficult to be impressed by "prompt injection" attacks that require the victim to enter the malicious prompt themselves --- like, really? If you tell Rovo to exfiltrate your data, it'll do it?

Obviously, there should be URL protection rules to control what it can access, but this requires a very specific and unlikely set of circumstances to exploit.

Are people so obsessed with AI that they can't find it reasonable that it won't do obviously bad things if asked? Not even with a confirmation or warning? We trust AI to literally build products and fix our most critical bugs, but we can't expect it to tell when it's being asked to do something malicious? Imagine if we felt this way about QA when trying DROP TABLES; in search bars. "Oh, well of course it broke the database, the user asked it to!"