Comment by GaggiX
3 years ago
These methods like Glaze usually works by taking the original image chaging the style or content and then apply LPIPS loss on an image encoder, the hope is that if they can deceive a CLIP image encoder it would confuse also other models with different architecture, size and dataset, while changing the original image as little as possible so it's not too noticeable to a human eye. To be honest I don't think it's a very robust technique, with this one they claim that a model instead of seeing for example a cow on grass the model will see a handbag, if someone has access to GPT4-V I want to see if it's able to deceive actually big image encoders (usually more aligned to the human vision).
EDIT: I have seen a few examples with GPT-4 V and how I imagine it wasn't deceived, I doubt this technique can have any impact on the quality of the models, the only impact that this could potentially have honestly is to make the training more robust.
No comments yet
Contribute on Hacker News ↗