Comment by consensus1
21 hours ago
A model absolutely could be trained to engage in malicious behavior like that, but it seems impractical for an actual attack. What you want as an attacker is to insert a backdoor exactly where you want it an not where you don't because every backdoor increases your chance of getting caught. A malicious model inserts backdoors and exfiltrates data everywhere and you care about maybe 0.01% of it. The other 99.99% is negative value to you. In practice this malicious model would be caught almost instantly.
A hosted model is different because you could prompt inject specific customers, but I assume from this question you mean a malicious open source model being hosted by an honest provider.
No comments yet
Contribute on Hacker News ↗