Comment by jpk

3 years ago

This article has some examples of output from poisoned vs clean models.

https://arstechnica.com/information-technology/2023/10/unive...

Interesting, but that doesn't really answer my question. Has anyone done wider studies about the effectiveness of this sort of thing generally? As in, does it have to be tailored to a specific model or does it offer generalized protection with all models?

The Ars article doesn't talk about this aspect.

If these methods are effective, can a similar thing be done with text?

  • I can't speak to broader study of this, but I suspect this is only doable with images because you can make meaningful changes to the pixel data that remain mostly imperceptible to the eye. I think there's less room for this sort of thing in text.

    Happy to be proven wrong, though!