The rumblings I'm hearing are that this a) barely works with last-gen training processes b) does not work at all with more modern training processes (GPT-4V, LLaVA, even BLIP2 labelling [1]) and c) would not be especially challenging to mitigate against even should it become more effective and popular. The Authors' previous work, Glaze, also does not seem to be very effective despite dramatic proclamations to the contrary, so I think this might be a case of overhyping an academically interesting but real-world-impractical result.
The screenshots you sent in [1] are inference, not training. You need to get a Nightshaded image into the training set of an image generator in order for this to have any effect. When you give an image to GPT-4V, Stable Diffusion img2img, or anything else, you're not training the AI - the model is completely frozen and does not change at all[0].
I don't know if anyone else is still scraping new images into the generators. I've heard somewhere that OpenAI stopped scraping around 2021 because they're worried about training on the output of their own models[1]. Adobe Firefly claims to have been trained on Adobe Stock images, but we don't know if Adobe has any particular cutoffs of their own[2].
If you want an image that screws up inference - i.e. one that GPT-4V or Stable Diffusion will choke on - you want an adversarial image. I don't know if you can adversarially train on a model you don't have weights for, though I've heard you can generalize adversarial training against multiple independent models to really screw shit up[3].
[0] All learning capability of text generators come from the fact that they have a context window; but that only provides a short term memory of 2048 tokens. They have no other memory capability.
[1] The scenario of what happens when you do this is fancifully called Habsburg AI. The model learns from it's own biases, reinforcing them into stronger biases, while forgetting everything else.
[2] It'd be particularly ironic if the only thing Nightshade harms is the one AI generator that tried to be even slightly ethical.
[3] At the extremes, these adversarial images fool humans. Though, the study that did this intentionally only showed the images for a small period of time, the idea being that short exposures are akin to a feed-forward neural network with no recurrent computation pathways. If you look at them longer, it's obvious that it's a picture of one thing edited to look like another.
Hey you know what might not be AI generated post-2021? Almost everything run through Nightshade. So given it's defeated, which is pretty likely, artists have effectively tagged their own work for inclusion.
Correct me if I'm wrong but I understand image generators as relying on auto-labeled images to understand what means what, and the point of this attack to make the auto-labelers mislabel the image, but as the top-level comment said it's seemingly not tricking newer auto-labelers.
Even if no new images are being scraped to train the foundation text-to-image models, you can be certain that there is a small horde of folk still scraping to create datasets for training fine-tuned models, LoRAs, Textual Inversions, and all the new hotness training methods still being created each day.
If it doesn't work during inference I really doubt it will have any intended effect during training, there is simply too much signal and the added adversarial noise works on the frozen and small proxy model they used (CLIP image encoder I think) but it doesn't work on a larger model and trained on a different dataset, if there is any effect during training it will probably just be the model learning that it can't take shortcuts (the artifacts working on the proxy model showcase gaps in its visual knowledge).
Generative models like text-to-image have an encoder part (it could be explicit or not) that extract the semantic from the noised image, if the auto-labelers can correctly label the samples then the encoded trained on both actual and adversarial images will learn to not take the same shortcuts that the proxy model has taken making the model more robust, I cannot see an argument where this should be a negative thing for the model.
The context windows of LLMs are now significantly larger than 2048 tokens, and there are clever ways to autopopulate context window to remind it of things.
The animation when you change images makes it harder to see the difference, I opened the three images each in its own tab and the differences are more apparent when you change between each other instantly.
I didn't see it immediately either, but there's a ton of added noise. The most noticeable bit for me was near the standing person's bent elbow, but there's a lot more that becomes obvious when flipping back and forth between browser tabs instead of swiping on Twitter.
Maybe it's more about "protecting" images that artists want to publicly share to advertise work, but it's not appropriate for final digital media, etc.
Seems obvious that the people stealing would be adjusting their process to negate these kinds of countermeasures all the time. I don't see this as an arms race the artists are going to win. Not like the LLM folks can consider actually paying their way...the business plan pretty much has "...by stealing everything we can get our hands on..." in the executive summary.
Huge market for snake oil here. There is no way that such tools will ever win, given the requirements the art remain viewable to human perception, so even if you made something that worked (which this sounds like it doesn’t) from first principles it will be worked around immediately.
The only real way for artists or anyone really to try to hold back models from training on human outputs is through the law, ie, leveraging state backed violence to deter the things they don’t want. This too won’t be a perfect solution, if anything it will just put more incentives for people to develop decentralized training networks that “launder” the copyright violations that would allow for prosecutions.
All in all it’s a losing battle at a minimum and a stupid battle at worst. We know these models can be created easily and so they will, eventually, since you can’t prevent a computer from observing images you want humans to be able to observe freely.
The level of claims accompanied by enthusiastic reception from a technically illiterate audience make it sound, smell, and sound like snake oil without much deep investigation.
There is another alternative to the law. Provide your art for private viewing only, and ensure your in person audience does not bring recording devices with them. That may sound absurd, but it's a common practice during activities like having sex.
That doesn't sound like a viable business model. There seems to be a non-trivial bootstrap problem involved -- how do you become well-known enough to attract audiences to private venues in sufficient volume to make a living? -- and would in no way diminish demand for AI-generated artwork which would still continue to draw attention away from you.
The thing is people want the benefits of having their stuff public but not bear the costs. Scraping has been mostly a solved problem especially when it comes to broad crawling. Put it under a login, there, no more AI "stealing" your work.
This would just create a new market for art paparazzis who would find any and all means to inflitrate such private viewings with futuristic miniature cameras and other sensors and selling it for a premium. Less than 24 hours later the files end up on hundreds or thousands of centralized and decentralized servers.
I'm not defending it. Just acknowledging the reality. The next TMZ for private art gatherings is percolating in someone's garage at the moment.
This tool is free, and as far as I can tell it runs locally. If you're not selling anything, and there's no profit motive, then I don't think you can reasonably call it "snake oil".
At worst, it's a waste of time. But nobody's being deceived into purchasing it.
If this is a danger from "snake oil" of this type, it'd be from the other side, where artists are intentionally tricked into believing that tools like this mean that AI isn't or won't be a threat to their copyrights in order to get them to stop opposing it so strongly, when in fact the tool does nothing to prevent their copyrights from being violated.
I don't think that's the intention of Nightshade, but I wouldn't put past someone to try it.
So then of course, you also cannot sell your work, as those might put it online. And you cannot show your art to big crowds, as some will make pictures and put it online. So ... you can become a literal underground artists, where only some may see your work. I think only some will like that.
But I actually disagree, there are plenty of ways to be an artist now - but most should probably think about including AI as a tool, if they still want to make money. But with the exception of some superstars, most artists are famously low on money - and AI did not introduce this. (all the professional artists I know, those who went to art school - do not make their income with their art)
Everything old is new again. It's the same thing with any DRM that happens on the client side. As long as it's viewable by humans, someone will figure out a way to feed that into a machine.
Other people pointed out they appreciated this prose. It’s easy to forget what exactly people are asking for when they talk about regulating the training of machine learning models.
> leveraging state backed violence to deter the things they don’t want
I just want to say: I really appreciate the stark terms in which you've put this.
The thing that has come to be called "intellectual property" is actually just a threat of violence against people who arrange bytes in a way that challenges power structures.
I heard that flooding the net with AI generated art would do much much more harm to generative AI than this whatever is this. Yes, this must be some snake oil salesman, those take it seriously turn AIs own weapon against AI.
I'm thinking — is it possible to create something on a global level similar to what they did in Snapchat: some sort of image flickering that would be difficult to parse, but still acceptable for humans?
Sorry i do not use Snapchat and with googeling "Snapchat image flickering" i did not find a good result. Could you elaborate this a bit more or provide me with a link where this is described? Thank you very much. :)
My guess. Is that at some poi t of time You will not be able to use any generated image or video in commercial. Because of 100% copyright claim for using parts of copyrighted image. Like YouTube those days. When some random beeps matches with someone music...
A few months ago I made a proof-of-concept on how finetuning Stable Diffusion XL on known bad/incoherent images can actually allow it to output "better" images if those images are used as a negative prompt, i.e. specifying a high-dimensional area of the latent space that model generation should stay away from: https://news.ycombinator.com/item?id=37211519
There's a nonzero chance that encouraging the creation of a large dataset of known tampered data can ironically improve generative AI art models by allowing the model to recognize tampered data and allow the training process to work around it.
This seems like a pretty pointless "arms race" or "cat and mouse game". People who want to train generative image models and who don't care about what artists think about it at all can just do some basic post-processing on the images that is just enough to destroy the very carefully tuned changes this Nightshade algorithm makes. Something like resampling it to slightly lower resolution and then using another super-resolution model on it to upsample it again would probably be able to destroy these subtle tweaks without making a big difference to a human observer.
In the future, my guess is that courts will generally be on the side of artists because of societal pressures, and artists will be able to challenge any image they find and have it sent to yet another ML model that can quickly adjudicate whether the generated image is "too similar" to the artist's style (which would also need to be dissimilar enough from everyone else's style to give a reasonable legal claim in the first place).
Or maybe artists will just give up on trying to monetize the images themselves and focus only on creating physical artifacts, similar to how independent musicians make most of their money nowadays from touring and selling merchandise at shows (plus Patreon). Who knows? It's hard to predict the future when there are such huge fundamental changes that happen so quickly!
>Or maybe artists will just give up on trying to monetize the images themselves and focus only on creating physical artifacts, similar to how independent musicians make most of their money nowadays from touring and selling merchandise at shows (plus Patreon).
As is, art already isn't a sustainable career for most people who can't get a job in industry. The most common monetization is either commissions or hiding extra content behind a pay wall.
To be honest I can see more proverbial "Furry artists" sprouting up in a cynical timeline. I imagine like every other big tech that the 18+ side of this will be clamped down hard by the various powers that be. Which means NSFW stuff will be shielded a bit by the advancement and you either need to find underground training models or go back to an artist. .
It's not particularly that hard. The furry nsfw models are already the most well developed and available models you can get right now. And they are spitting out stuff that is almost indistinguishable from regular art.
> musicians make most of their money nowadays from touring and selling merchandise at shows
Be reminded that this is - and has always been - the mainstream model of the lineages of what have come to be called "traditional" and "Americana" and "Appalachian" music.
The Grateful Dead implemented this model with great finesse, sometimes going out of their way to eschew intellectual property claims over their work, in the belief that such claims only hindered their success (and of course, they eventually formalized this advocacy and named it "The Electronic Frontier Foundation" - it's no coincidence that EFF sprung from deadhead culture).
It is a funny appearance (weird viewpoint) that artists are furious loosing their monopily in stealing and cloning components from other artists, recomposing into a similar but new thing.
And that OpenArt on the analogy of OpenSource is a non-existing thing (I know, I know, different things, source code is not for the generic audience and can be hidden on will, unlike art, just having some generative thoughts artefact here ;) )
This feels like it'll actually help make AI models better versus worse once they train on these images. Artists are basically, for free, creating training data that conveys what types of noise does not change the intended meaning of the image to the artist themselves.
The number of people who are going to be able to produce high fidelity art with off the shelf tools in the near future is unbelievable.
It’s pretty exciting.
Being able to find a mix of styles you like and apply them to new subjects to make your own unique, personalized, artwork sounds like a wickedly cool power to give to billions of people.
In terms of art, population tends to put value not on the result, but origin and process. People will just look down on any art that’s AI generated in a couple of years when it becomes ubiquitous.
> population tends to put value not on the result, but origin and process
I think population tends to value "looks pretty", and it's other artists, connoisseurs, and art critics who value origin and process. Exit Through the Gift Shop sums this up nicely
I disagree. I definitely value modern digital art more than most historical art, because it just looks better. If AI art looks better (and in some cases it does) then I'll prefer that.
This is already the case. Art is a process, a form of human expression, not an end result.
I'm sure OpenAI's models can shit out an approximation of a new Terry Pratchett or Douglas Adams novel, but nobody with any level of literary appreciation would give a damn unless fraud was committed to trick readers into buying it. It's not the author's work, and there's no human message behind it.
> Being able to find a mix of styles you like and apply them to new subjects to make your own unique, personalized, artwork sounds like a wickedly cool power to give to billions of people.
And in the process, they will obviate the need for Nightshade and similar tools.
AI models ingesting AI generated content does the work of destroying the models all by itself. Have a look at "Model Collapse" in relation to generative AI.
I know this is an unpopular thing to say these days, but I still think the internet is amazing.
I have more access to information now than the most powerful people in the world did 40 years ago. I can learn about quantum field theory, about which pop star is allegedly fucking which other pop star, etc.
If I don't care about the law I can read any of 25 million books or 100 million scientific papers all available on Anna's Archive for free in seconds.
Not really. There is a reason why we find realistic painting to be more fascinating than a photo and why some still practice it. The effort put in by another artist does affect our enjoyment.
For me it doesn’t. I’m generating images, realistic, 2.5d, 2d and I like them as much. I don’t feel (or miss) what you described. Or what any other arts guy describes, for that matter. Arts people are different, because they were trained to feel something a normal person wouldn’t. And that’s okay, a normal person without training wouldn’t see how much beauty and effort there is in an algorithm or a legal contract as well.
The word "we" is doing a lot of heavy lifting here. A large majority of consumers can't even tell apart AI-generated from handmade, let alone care who or what made the thing.
I want progressive fees on copyright/IP/patent usage, and worldwide gov cooperation/legislation (and perhaps even worldwide ability to use works without obtaining initial permission, although let's not go into that outlandish stuff)
I want a scaling license fee to apply (e.g. % pegged to revenue. This still has an indirect problem with different industries having different profit margins, but still seems the fairest).
And I want the world (or EU, then others to follow suit) to slowly reduce copyright to 0 years* after artists death if owned by a person, and 20-30 years max if owned by a corporation.
And I want the penalties for not declaring usage** / not paying fees, to be incredibly high for corporations... 50% gross (harder) / net (easier) profit margin for the year? Something that isn't a slap on the wrist and can't be wriggled out of quite so easily, and is actually an incentive not to steal in the first place.)
[*]or whatever society deems appropriate.
[**]Until auto-detection (for better or worse) gets good enough.
IMO that would allow personal use, encourages new entrants to market, encourages innovation, incentivises better behaviour from OpenAI et al.
> And I want the world (or EU, then others to follow suit) to slowly reduce copyright to 0 years* after artists death if owned by a person, and 20-30 years max if owned by a corporation.
Why death at all?
It's icky to trigger soon after death, it's bad to have copyright vary so much based on author age, and it's bad for many works to still have huge copyright lengths.
It's perfectly fine to let copyright expire during the author's life. 20-30 years for everything.
Extremely naive to think that any of this could be enforced to any adequate level. Copyright is fundamentally broken and putting some plasters on it is not going to do much especially when these plasters are several decades too late.
I just tested it with Azure AI image classification and it worked - so this cat is yet to adapt to the mouse’s latest idea.
I still feel it is absolutely wrong to roam around the internet and scrape images (without consent) in order to power one’s cash cow AI. I hope more methods to protect artworks (including audio and other formats) become more accessible.
Artists copy from each other all the time. Arguably, culture exists because of copying (folk stories by necessity); copyright makes culture top-down and stagnant, and you can't avoid it because they have the money to shove it right in your face. Who wants trickle-down culture?
I might be missing something because I don't know much about the architecture of either Nightshade or AI art generators, but I wonder if you could try to have a GAN-like architecture (an extra model trying to trick the model) for the part of the generator that labels images to build resistance to Nightshade-like filters.
It doesn't even have to be a full GAN, you only need to train the discriminator side to filter out the data. Clean reference images + Nightshade would be the generator side.
What the article doesn't illustrate is that it destroys fine detail in the image, even in the thumbnails of the reference paper:
https://arxiv.org/pdf/2310.13828.pdf
Also... Maybe I am naive, but it seems rather trivial to work around with a quick prefilter? I don't know if tradition denoising would be enough, but worst case you could run img2img diffusion.
The poisoned images aren't intended to be viewed, rather scraped and pass a basic human screen. You wouldn't be able to denoise as you'd have to denoise the entire dataset, the entire point is that these are virtually undetectable from typical training set examples, but they can push prompt frequencies around at will with a small number of poisoned examples.
Long-term I think the real problem for artists will be corporations generating their own high quality targeted datasets from a cheap labor pool, completely outcompeting them by a landslide.
In the short-to-medium term, we're seeing huge improvements in the data efficiency of generative models. We haven't really started to see self-training in diffusion models, which could improve data efficiency by orders of magnitude. Current models are good at generalisation and are getting better at an incredible pace, so any efforts to limit the progress of AI by restricting access to training data is a speedbump rather than a roadblock.
Art is already democratized. It has been for decades. Everyone can pick it up at zero cost. Even you!
The poorest people have historically produced great art. Training a model, however? Expensive. Running it locally? Expensive. Paying the sub? Expensive.
Nothing is being democratized, the only thing this does is devaluing the blood and sweat people have put into their work so FAANG can sell it to lazy suckers.
This is fantastic. If companies want to create AI models, they should license the content they use for the training data. As long as there are not sufficient legal protections and the EU/Congress do not act, tools like these can serve as a stopgap and maybe help increase pressure on policymakers
It's going to be interesting to see how the lawsuits against OpenAI by content creators plays out. If the courts rule that AI generated content is a derivative work of all the content it was trained on it could really flip the entire gen AI movement on its head.
If it were a derivative work[1] (and sufficiently transformational) then it's allowed under current copyright law and might not be the slam dunk ruling you were hoping for.
My biggest fear is that the big players will drop a few billion dollars to silence the copyright holders with power go away, and new rules are put in place that will make open-source models that can't do the same essentially illegal.
If the courts do rule that way, I would expect a legislative race between different countries to amend the relevant laws. Visual generative AI is just too lucrative a thing.
Isn't this just teaching the models how to better understand pictures as humans do? As long as you feed them content that looks good to a human, wouldn't they improve in creating such content?
You would think the economists at UChicago would have told these researchers that their tool would achieve the opposite effect of what they intended, but here we are.
In this case, the mechanism for how it would work is effectively useless. It doesn't affect OpenAI or other companies building foundation models. It only works on people fine-tuning these foundation models, and only if the image is glazed to affect the same foundation model.
These methods like Glaze usually works by taking the original image chaging the style or content and then apply LPIPS loss on an image encoder, the hope is that if they can deceive a CLIP image encoder it would confuse also other models with different architecture, size and dataset, while changing the original image as little as possible so it's not too noticeable to a human eye. To be honest I don't think it's a very robust technique, with this one they claim that a model instead of seeing for example a cow on grass the model will see a handbag, if someone has access to GPT4-V I want to see if it's able to deceive actually big image encoders (usually more aligned to the human vision).
EDIT: I have seen a few examples with GPT-4 V and how I imagine it wasn't deceived, I doubt this technique can have any impact on the quality of the models, the only impact that this could potentially have honestly is to make the training more robust.
Each time there is an update to training algorithms and in response poisoning algorithms, artists will have to re-glaze, re-mist, and re-nightshade all their images?
Eventually I assume the poisoning artifacts introduced in the images will be very visible to humans as well.
Yeah, I've seen multiple artists complain about how glazing reduces image quality. It's very noticeable. That seems like an unavoidable problem given how AI is trained on images right now.
I'm glad to see tools like Nightshade starting to pop up to protect the real life creativity of artists. I like AI art, but I do feel conflicted about its potential long term effects towards a society that no longer values authentic creativity.
Is the existence of the AI tool not itself a product of authentic creativity? Does eliminating barriers to image generation not facilitate authentic creativity?
No, it facilitates commoditization. Art – real art – is fundamentally a human-to-human transaction. Once everyone can fire perfectly-rendered perfectly-unique pieces of 'art' at each other, it'll just become like the internet is today: filled with extremely low-value noise.
To protect an individual's image property rights from image generating AI's -- wouldn't it be simpler for the IETF (or other standards-producing group) to simply create an
AI image exclusion standard
, similar to "robots.txt" -- which would tell an AI data-gathering web crawler that a given image or set of images -- was off-limits for use as data?
Entities training models have no incentive to follow such metadata. If we accept the premise that "more input -> better models" then there's every reason to ignore non-legally-binding metadata requests.
Robots.txt survived because the use of it to gatekeep valuable goodies was never widespread. Most sites want to be indexed, most URLs excluded by the robots file are not of interest to the search engine anyway, and use of robots to prevent crawling actually interesting pages is marginal.
If there was ever genuine uptake in using robots to gatekeep the really good stuff search engines would've stopped respecting it pretty much immediately - it isn't legally binding after all.
>Entities training models have no incentive to follow such metadata. If we accept the premise that "more input -> better models" then there's every reason to ignore non-legally-binding metadata requests.
Name two entities that were asked to stop using a given individuals' images that failed to stop using them after the stop request was issued.
>Robots.txt survived because the use of it to gatekeep valuable goodies was never widespread. Most sites want to be indexed, most URLs excluded by the robots file are not of interest to the search engine anyway, and use of robots to prevent crawling actually interesting pages is marginal.
Robots.txt survived because it was a "digital signpost" a "digital sign" -- sort of like the way you might put a "Private Property -- No Trespassing" sign in your yard.
Most moral/ethical/lawful people -- will obey that sign.
Some might not.
But the some that might not -- probably constitute about a 0.000001% minority of the population, whereas the majority that do -- probably constitute about 99.99999% of the population.
"Robots.txt" is a sign -- much like a road sign is.
People can obey them -- or they can ignore them -- but they can ignore them only at their own peril!
It's a sign which provides a hint for what the right thing to do in a certain set of circumstances -- which is what the Law is; which is what the majority of Laws are.
People can obey them -- or they can choose to ignore them -- but only at their own peril!
Most will choose to obey them. Most will choose to "take the hint", proverbially speaking!
A few might not -- but that doesn't mean the majority won't!
>If there was ever genuine uptake in using robots to gatekeep the really good stuff search engines would've stopped respecting it pretty much immediately - it isn't legally binding after all.
Again, name two entities that were asked to stop using a given individuals' images that failed to stop using them after the stop request was issued.
And then what? The scrapers themselves already happily ignore copyright, they won't be inclined to obey a no-ai.txt. So someone would have to enforce the standard. Currently I see no organisation who would be willing to do this or even just technologically able - as even just detecting such scrapers is an extremely hard task.
Nevertheless, I hope that at some not-so-far point in the future there will be more legal guidance about this kind of stuff, i.e. it will be made clear that scraping violates copyright. This still won't solve the problem of detectability but it would at least increase the risk of scrapers, should they be caught.
>The scrapers themselves already happily ignore copyright, they won't be inclined to obey a no-ai.txt.
Name two entities that were asked to stop using a given individuals' images that failed to stop using them after the stop request was issued.
>Currently I see no organisation who would be willing to do this or even just technologically able - as even just detecting such scrapers is an extremely hard task.
// Part of Image Web Scraper For AI Image Generator ingestion psuedocode:
if fileExists("no-ai.txt") {
// Abort image scraping for this site -- move on to the next site
} else {
// Continue image scraping for this site
};
See? Nice and simple!
Also -- let me ask you this -- what happens to the intellectual property (or just plain property) rights of Images on the web after the author dies? Or say, 50 years (or whatever the legal copyright timeout is) after the author dies?
Legal grey area perhaps?
Also -- what about Images that exist in other legal jurisdictions -- i.e., other countries?
How do we know what set of laws are to apply to a given image?
?
Point is: If you're going to endorse and/or construct a legal framework (and have it be binding -- keep in mind you're going to have to traverse the legal jurisdictions of many countries, many countries!) -- you might as well consider such issues.
Also -- at least in the United States, we have Juries that can override any Law (Separation of Powers) -- that is, that which is considered "legally binding" -- may not be quite so "legally binding" if/when properly explained to a proper jury in light of extenuating (or just plain other) circumstances!
So kindly think of these issues prior to making all-encompasing proposals as to what you think should be "legally binding" or not.
I comprehend that you are just trying to solve a problem; I comprehend and empathize; but the problem might be a bit greater than you think, and there might be one if not serveral unexplored partial/better (since no one solution, legal or otherwise, will be all-encompassing) solutions -- because the problem is so large in scope -- but all of these issues must be considered in parallel -- or errors, present or future will occur...
Setting aside the efficacy of this tool, I would be very interested in the legal implications of putting designs in your art that could corrupt ML models.
For instance, if I set traps in my home which hurt an intruder we are both guilty of crimes (traps are illegal and are never considered self defense, B&E is illegal).
Would I be responsible for corrupting the AI operator's data if I intentionally include adversarial artifacts to corrupt models, or is that just DRM to legally protect my art from infringement?
edit:
I replied to someone else, but this is probably good context:
DRM is legally allowed to disable or even corrupt the software or media that it is protecting, if it detects misuse.
If an adversarial-AI tool attacks the model, it then becomes a question of whether the model, having now incorporated my protected art, is now "mine" to disable/corrupt, or whether it is in fact out of bounds of DRM.
So for instance, a court could say that the adversarial-AI methods could only actively prevent the training software from incorporating the protected media into a model, but could not corrupt the model itself.
None whatsoever. There is no right to good data for model training, nor does any contractual relationship exist between you and and a model builder who scrapes your website.
If you're assuming this is open-shut, you're wrong. I asked this specifically as someone who works in security. A court is going to have to decide where the line is between DRM and malware in adversarial-AI tools.
The way Nightshade works (assuming it does work) is by confusing the features of different tags with each other. To argue that this is illegal would be to argue that mistagging a piece of artwork on a gallery is illegal.
If you upload a picture of a dog to DeviantArt and you label it as a cat, and a model ingests that image and starts to think that cats look like dogs, would anybody claim that you are breaking a law? If you upload bad code to Github that has bugs, and an AI model consumes that code and then reproduces the bugs, would anyone argue that uploading badly written code to Github is a crime?
What if you uploaded some bad code to Github and then wrote a comment at the top of the code explaining what the error was, because you knew that the model would ignore that comment and would still look at the bad code. Then would you be committing a crime by putting that code on Github?
Even if it could be proven that your intention was for that code or that mistagged image to be unhelpful to training, it would still be a huge leap to say that either of those activities were criminal -- I would hope that the majority of HN would see that as a dangerous legal road to travel down.
No, it's much closer to (in fact, it is simply) asking if adversarial AI tools count as DRM or as malware. And a court is going to have to decide whether the model and or its output counts as separate software, which it is illegal for DRM to intentionally attack.
DRM can, for instance, disable its own parent tool (e.g. a video game) if it detects misuse, but it can't attack the host computer or other software on that computer.
So is the model or its output, having been trained on my art, a byproduct of my art, in which case I have a legal right to 'disable' it, or is it separate software that I don't have a right to corrupt?
I see it as no different than mapmakers inventing a nonexistent alley, to check who copies their maps verbatim ("trap street"). Even if this caused, for example, a car crash because of an autonomous driver, the onus I think would be on the one that made the car and used the stolen map for navigation, and not on the one that created the original map.
I find the AI training topic interesting, because it's really data/information that is involved. Forget about the fact that it's images or stories or Reddit posts, it's all data.
We are born and then exposed to the torrent of data from the world around us, mostly fed to us by other humans, this is what models are trying to tap.
Unfortunately our learning process is completely organic and takes decades and decades and decades; there's no way to put a model through this easily.
Perhaps we need to seed the web with AI agents who converse and learn as much like regular human beings as possible and assemble the dataset that way. Although having an agent browse and find an image to learn to draw from is still gonna make people reee even if that's exactly what a young and aspiring human artist would be doing.
Don't talk about humans being sacred; we already voted to let corporations be people, for the 1% to exist and "lobby", breaking our democracy so that they can get tax breaks and make corrupt under the table deals. None of us stopped that from happening...
Does it survive AI upscaling or img2img? If not - then it's useless. Nobody trains AI models without any preprocessing. This is basically a tool for 2022.
For this to work, wouldn't you have to have an enormous number of artists collaborating on "poisoning" their images the same way (cow to handbag) while somehow keeping it secret form ai trainers that they were doing this?
It seems to me that even if the technology works perfectly as intended, you're effectively just mislabeling a tiny fraction of the training data.
1. They don't need an enormous number of artists; the research paper showed significant results with even 50 poisoned image samples in the dataset, which is enough to be contained in even a single artist's online gallery.
2. They don't need to keep it a secret; the goal is to remove these images from the training data, in a way that would be much more efficient than simply adding a "please don't include my art in your ai scraper" message next to your pictures.
In so far as anger goes against AIs being trained on particular intellectual properties.
A made up scenario¹ is that a person who is training an AI, goes to the local
library and checks out 600 books on art.
The person then lets the AI read all of them.
After which they are returned to the library and another 600 books are borrowed
Then we can imagine the AI somehow visiting a lot of museums and galleries.
The AI will now have been trained on the style and looks of a lot
of art from different artists
All the material has been obtained in a legal manner.
Is this an acceptable use?
Or can an artist still assert that the AI was trained with their
IP without consent?
Clearly this is one of the ways a human would go about learning
about styles, techniques etc..
¹ Yes you probably cannot borrow 600 books at a time.
How does the AI read the books? I dont know. Simplicity would be
that the researcher takes a photo of each page.
This would be extremmly slow but for this hypothetical it is acceptable.
I think the key difference here is that the most prominent image generation AIs are commercial and for-profit. The scenarios you describe are comparing a commercial AI to a private person. You cannot get a library card for a company, and you cannot bring a photography crew to a gallery without permission.
I’m completely flabbergasted by the number of comments implying copyright concepts such as “fair use” or “derivative work” apply to trained ML models. Copyright is for _people_, as are the entailing rights, responsibilities and exemptions.
This has gone far beyond anthropomorphising and we need to like get it together, man!
Oh come on, you’re being insincere. Wether or not the model is learning from the work just like people is hotly debated as if it would make a difference. Fair use is even brought up. Fair use! Even if it applied, these training sets collate all of everything
I really don't understand the anxiety of artists towards AI - as if creatives haven't always borrowed and imitated. Every leading artist has had acolytes, and while it's true no artist ever had an acolyte as prodigiously productive as AI will be, I don't see anything different between a young artist looking to Picasso for cues and Stable Diffusion or DALL-E doing the same. Styles and methods haven't ever been subject to copyright - and art would die the moment that changed.
The only explanation I can find for this backlash is that artists are actually worried just like the rest of us that pretty soon AI will produce higher quality more inventive work faster and more imaginatively than they can - which is very natural, but not a reason to inhibit an AI's creative education.
This has been litigated over and over again, and there have been plenty of good points made and concerns raised over it by those who it actually affects. It seems a little bit disingenuous (especially in this forum) to say that that conclusion is the "only explanation" you can come up with. And just to avoid prompting you too much: trust me, we all know or can guess why you think AI art is a good thing regardless of any concerns one might bring up.
Could you please stop posting unsubstantive comments and flamebait? You've unfortunately been doing it repeatedly. It's not what this site is for, and destroys what it is for.
Imitation isn’t the problem so much as it is that ML generated images are composed of a mush of the images it was trained on. A human artist can abstract the concepts underpinning a style and mimic it by drawing all-new lineart, coloration, shading, composition, etc, while the ML model has to lean on blending training imagery together.
Furthermore there’s a sort of unavoidable “jitter” in human-produced art that varies between individuals that stems from vastly different ways of thinking, perception of the world, mental abstraction processes, life experiences, etc. This is why artists who start out imitating other artists almost always develop their imitations into a style all their own — the imitations were already appreciably different from the original due to the aforementioned biases and those distinctions only grow with time and experimentation.
There would be greatly reduced moral controversy surrounding ML models if they lacked that mincemeat/pink slime aspect.
I love it. This undermines the notion of ground truth. What separates correct information from incorrect information? Maybe nothing! I love how they acknowledge the never ending attack versus defense game. In stark contrast to "our AI will solve all your problems".
No, it's resistant to transformation. Rotation, cropping, scaling, the image remains poisonous. The only antidote known currently is active artist cooperation.
I think it's worthwhile for such discussion to happen in the open. If the tool can be defeated through simple means, it's better for everybody to know that, right?
Is it possible to reliably detect whether an image is poisoned? If not then it achieves the goal of punishing entities which indiscriminately harvest data.
Only protection is adding giant gaping vaginas to your art, nothing less will deter scraping. If the Email spam community showed us something in the last 40 years is that no amount of defensive tech measures will work except financial disincentives.
Any AI art/video/photography/music/etc generator company who generates revenue needs to add watermarks to let the public know its AI generator. This should be forced via legislation in all countries.
If they don't then whatever social network or other services where things can shared/viewed by large groups to millions & are posted publicly need to be labeled "We can not verify veracity of this content."
I want a real internet ..this AI stuff is just triple fold increasing fake crap on the Internet and in turn / time our trust in it!
For visual artists who don't want visible artifacting in the art they feature online, would it be possible to upload these alongside your un-poisoned art, but have them only hanging out in the background? So say having one proper copy and a hundred poisoned copies in the same server, but only showing the un-poisoned one?
Might this "flood the zone" approach also have -some- efficacy against human copycats?
I wonder how this tool works if it's actually model independent. My understanding so far was that in principle each possible model has some set of pathological inputs for which the classification will be different than what a user sees - but that this set is basically different for each model. So did they actually manage to build an "universal" poison? If yes, how?
I wonder if this is illegal in some countries. In France for example, there is the following law: "Obstructing or distorting the operation of an automated data processing system is punishable by five years' imprisonment and a fine of €150,000.".
If you ask me, this is 100% applicable in this case, so I wonder what a judge would rule.
Remember when the music industry tried to use technology to stop music pirating?
This will work about as well...
Oh, I forget, fighting music pirating was considered an evil thing to do on HN. "pirating is not stealing, is copyright infringement", right? Unlike training neural nets on internet content which of course is "stealing".
FWIW, you're the only use of the word "steal" in this comment thread.
Many people would in fact argue that training AI on people's art without permission is copyright infringement, since the thing it (according to detractors) does is infringe copyright by generating knockoffs of people's work.
You will see some people use the term "stealing" but they're usually referring to how these AIs are sold/operated by for-profit companies that want to make money off artists' work without compensating them. I think it's not unreasonable to call that "stealing" even if the legal definition doesn't necessarily fit 100%.
The music industry is also not really a very good comparison point for independent artists... there is no Big Art equivalent that has a stranglehold on the legislature and judiciary like the RIAA/MPAA do.
The difference is that “pirating” is mostly done by individuals for private use, whereas training is mostly done by megacorporations looking to make more money.
Musicians overwhelmingly do not even attempt to clear samples. This also isn't a great comparison since samples are taken directly out of the audio, not turned into a part of a pattern used to generate new sounds like what AI generators do with images
I wonder if we know enough about any of these systems to make such claims. This is all predicated on the fact that this tool will be in widespread use. If it is somehow widely used beyond the folks who have seen it at the top of HN, won't the big firms have countermeasures, ready to deploy?
It's an arms race the bigger players will win, and it undermines the quality of the images. But it feels natural that artists would want to do something since they don't feel like anyone else is protecting them right now.
The intention is good, from an AI-opponent's perspective. I don't think will work practically, though. The drawbacks for actual users of the image galleries, plus the level of complexity involved in poisoning the samples makes this unfeasible to implement at the scale required.
The opening website is so poor - "what is nightshade" - then a whole paragraph that tells nothing, then another paragraph.. then no examples. This whole description should be reworked to be shorter and more to the point.
The image generation models now are at the point where they can produce their own synthetic training images. So I'm not sure how big of an impact something like this would have.
would it have been that hard to include a sample photo and how it looks with the nightshade filter side by side in a 3 page document describing how it would look in great detail
Baffling to see anyone argue against this technology when it is a non-issue to any model by simply acquiring only training data you have permission to use.
The reason people are arguing against this technology is that no one is using them in the way you describe. They actually wouldn't even be economically viable in that case.
If it is not economically viable for you to be ethical, then you do not deserve economic success.
Anyone arguing against this technology following the line of reasoning you present is operating in adverse to the good of society. Especially if their only motive is economic viability.
I think people 100% have the right to use this on their images, but:
> simply acquiring only training data you have permission to use
Currently it's generally infeasible to obtain licenses at the required scale.
When attempting to develop a model that can describe photos for visually impaired users, I had even tried to reach out to obtain a license from Getty. They repeatedly told me that they don't license images for machine learning[0].
I think it's easy to say "well too bad, it doesn't deserve to exist" if you're just thinking about DALL-E 3, but there's a huge number of positive and far less-controversial applications of machine learning that benefit from web-scale pretraining and foundation models - spam filtering, tumour segmentation, voice transcription, language translation, defect detection, etc.
I don't believe it's a "doesn't deserve to exist" situation, because these things genuinely can be used for the public good.
However - and this is a big however - I don't believe it deserves the legal protection to be used for profit.
I am of the opinion that if you train your model on data that you do not hold the rights for, your usage should be handled similarly to most fair use laws. It's fine to use it for your personal projects, for research and education, etc. but it is not OK to use it for commercial endeavors.
I think the artists need to agree to stop making art altogether. That ought to get people’s attention. Then the AI people might (be socially pressured or legally forced to) put their tools away.
No, they'll just demand that artists produce more art so they can continue scraping, because if you work in tech you're allowed to be entitled, you're The Face Of The Future and all you're trying to do is Save The World, all these decels are just obstacles to be destroyed.
However much we might wish that it was not true, ideas are not rivalrous. If you share an idea with another person, they now have that idea too.
If you share words on paper, then someone with eyes and a brain might memorize them (or much more likely, just grasp and retain the ideas conveyed in the words).
If you let someone hear your music, then the ideas (phrasing, style, melody, etc) in that music are transferred.
If you let people see a visual work, then the stylistic and content elements of that work are potentially absorbed by the audience.
We have copyright to protect specific embodiments, but mostly if you try to share ideas with others without letting them use the ideas you shared, then you are in for a life of frustration and escalating arms race.
I completely sympathize with anyone who had a great idea and spent a lot of effort to realize it. If I invented/created something awesome I would be hurt and angry if someone “copied” it. But the hard cold reality is that you cannot “own” an idea.
> But the hard cold reality is that you cannot “own” an idea.
The above comment is true about the properties of information, as explained via the lens of economics. [1]
However, one ignores ownership as defined by various systems (including the rule of law and social conventions) at one's own peril. Such systems can also present a "hard cold reality" that can bankrupt or ostracize you.
[1] Don't let the apparent confidence and technicality of the language of economists fool you. Economics isn't the only game in town. There are other ways to model and frame the world.
[2] Dangling footnote warning. I think it is instructive to recognize that the field of economics has historically shown a kind of inferiority complex w.r.t. physics. Some economists ascribe to the level of rigor found in physics and that is well and good, but perhaps that effort should not be taken too seriously nor too far, since economics as a field operates at a different level. IMO, it would be wise for more in the field to eat a slice of humble pie.
[3] Ibid. It is well-known that economists can be "hired guns" used to "prove" a wide variety of things, many of which are subjective. My point: you can hire an economist to shore up one's political proposals. Is the same true of physicists? Hopefully not to the same degree. Perhaps there are some cases of hucksterism, but nothing like the history of economists-wagging-the-dog! At some point, the electron tunnels or it does not.
Many terms of art from economics are probably not widely-known here.
> In economics, a good is said to be rivalrous or a rival if its consumption by one consumer prevents simultaneous consumption by other consumers, or if consumption by one party reduces the ability of another party to consume it. - Wikipedia: Rivalry (economics)
Also: we should recognize that stating something as rivalrous or not is descriptive (what exists) not normative (what should be).
We're not trying to keep the AI from learning general ideas, we're trying to keep it from memorizing specific expressions[0]. There's a growing body of research to show that these models are doing a lot of memorizing, even if they're not regurgitating that data. For example, Google's little "ask GPT to repeat a word forever" trick, which will make GPT-4 spit out verbatim training data[1].
If there was a training process that let us pick a minimal sample of examples and turn it into a general purpose art generator or text generator, I think people would have been fine with that. But that's not what any of these models do. They were trained on shittons of creative expression, and there's statistical evidence that the models retain that expression, in a way that is fundamentally different from how humans remember, misremember, adapt, remix, and/or "play around with" other people's creativity.
[0] You called these "embodiments", but I believe you're trying to invoke the idea/expression divide, so I'll run with that.
[1] Or at least it did. OpenAI now filters out conversations that trip the bug.
I don't see the parallel between this offensive tool and DRM. I could, say buy a perpetual license to an image from the artist, so that I can print it and put it on my wall, while it can simultaneously be poisonous to an AI system. I can even steal it and print it, while it is still poisonous to an AI system.
The closest parallel I can think of is that humans can ingest chocolate but dogs should not.
What you've described is the literal, dictionary definition of Digital Rights Management - a technology to restrict the use of a digital asset beyond the contractually-agreed terms. Copying is only one of many uses that the copyright-holder may wish to prevent. The regional lockout on a DVD had nothing to do with copy-protection, but it was still DRM.
It's about the arm's race: DRM will always be cracked (with a sufficiently motivated customer.) AI poisoning will always be cracked (with a sufficiently motivated crawler.)
This doesn’t stop anyone from viewing or scraping the work, though, so in no way is it DRM. It just causes certain methods of computer interpretation of an image to interpret it in an odd way vs. human viewers. They can still learn from them.
No, I disagree. There is no principle of the universe or across human civilizations that says that you have a right to eat because you produced a creative work.
The way societies work is that the members of the society contribute and benefit in prescribed ways. Societies with lots of excess production may at times choose to allow creative works to be monetized. Societies without much surplus are extremely unlikely to do so, eg a society with not enough food for everyone to eat in the middle of a famine is extremely unlikely to feed people who only create art; those people will have to contribute in some other way.
I think it is a very modern western idea (less than a century old) that many artists can dedicate themselves solely to producing the art they want to produce. In all other times artists either had day jobs or worked on commission.
The tragedy of "your business model is not my problem" as a spreading idea is that while you're right since distribution is where the money is (not creation), intellectual property is de-facto weakened today and IP piracy is widely considered an acceptable thing.
So is sabotaging solutions that would make creative work of the same (or superior) quality more affordable. Your ability to produce expensive illustrations hinders my ability to produce cheap textbooks.
Not everybody equates automated scraping for training models and human experience. Just like any other “data wants to be free” type of discussion, the philosophical and ethical considerations are anything but cut-and-dried, and they’re far more consequential than the technical and economics-in-a-vacuum ones. The general public will quite possibly see things differently than the “oh well, artists— that’s the free market for ya, and you lost” crowd.
You don't copyright ideas, you copyright works. And these artists' productions are works, not abstract ideas, with copyrights, and they are being violated. This is simple law. Why do people have such a hard time with this? Are you the one training the models and you need to find a cognitive escape out of the illegality and wrong-doing of your activities?
It’s not obvious to me that using a copyrighted image to train a model is copyright infringement. It’s certainly not copyright infringement when used to train a human who may end up creating works that are influenced by (but not copies of) the original works.
Now, if the original copyrighted work can be extracted or reproduced from the model, that’s obviously copyright infringement.
>This is simple law. Why do people have such a hard time with this?
Because this isn’t simple law. It feels like simple infringement, but there’s no actual copying going on. You can’t open up the database and find a given duplicate of a work. Instead you have some abstraction of what it takes to get to a given work.
Also it’s important to point out that nothing in the law is sure. A good lawyer, a sympathetic judge, a bored/interested/contrarian juror, etc can render “settled law” unsettled in an instant. The law is not a set of board game rules.
“One may well ask: ‘How can you advocate breaking some laws and obeying others?’ The answer lies in the fact that there are two types of laws: just and unjust. I would be the first to advocate obeying just laws. One has not only a legal but a moral responsibility to obey just laws. Conversely, one has a moral responsibility to disobey unjust laws. I would agree with St. Augustine that ‘an unjust law is no law at all.’”
Let's talk about ownership in a broader sense. In practice, one cannot effectively own (retain possession of) something without some combination of physical capability or coercion (or threat of coercion). Meaning: maintaining ownership of anything (physical or otherwise) often depends on the rule of law.
Then let's use a more precise term that is also present in law: monopoly.
You can't monopolize an idea.
Copyright law is a prescription, not a description. Copyright law demands that everyone play along with the lie that is intellectual monopoly. The effectiveness of that demand depends on how well it can be enforced.
Playing pretend during the age of the printing press may have been easy enough to coordinate, but it's practically impossible here in the digital age.
If we were to increase enforcement to the point of effectiveness, then what society would be left to participate? Surely not a society I am keen to be a part of.
My issue with this line of argument is that it’s anthropomorphizing machines. It’s fine to compare how humans do a task with how a machine does a task, but in the end they are very different from each other, organic vs hardware and software logic.
First, you need to prove that generative AI works fundamentally the same way as humans at the task of learning. Next you have to prove that it recalls information in the same way as humans. I don’t think anyone would say these are things that we can prove, let alone believe they do. So what we get is comments like they are similar.
What this means, is these systems will fall into different categories of law around copyright and free-use. What’s clear is that there are people who believe that they are harmed by the use of their work in training these systems and it reproducing that work in some manner later on (the degree to which that single work or the corpus of their work influences that final product is an interesting question). If your terms of use/copyright/license says “you may not train on this data”, then should that be protected in law? If a system like nightshade can effectively influence a training model enough to make it clear that something protected was used in its training, is that enough proof that the legal protections were broken?
>First, you need to prove that generative AI works fundamentally the same way as humans at the task of learning. Next you have to prove that it recalls information in the same way as humans.
No, you don't need to prove any of those things. They're irrelevant. You'd need to prove that the AI is itself morally (or, depending on the nature of the dispute, legally) equivalent to a human and therefore deserving of (or entitled to) the same rights and protections as a human. Since it is pretty indisputably the case that software is not currently legally equivalent to a human, you're stuck with the moral argument that it ought to be, but I think we're very far from a point where that position is warranted or likely to see much support.
No, it's not. It's merely pointing out the similarity between the process of training artists (by ingesting publicly available works) and ML models (which ingest publicly available works).
> First, you need to prove that generative AI works fundamentally the same way as humans at the task of learning.
Given that there is no comprehensive model for how humans actually learn things, that would be an unfeasible requirement.
We are machines. We just haven't evenly accepted it yet.
Our biology is mechanical, and lay people don't possess an intuition about this. Unless you've studied molecular biology and biochemistry, it's not something that you can easily grasp.
Our inventions are mechanical, too, and they're reaching increasing levels of sophistication. At some point we'll meet in the middle.
The way these ML models and humans operate are indeed quite different.
Humans work by abstracting concepts in what they see, even when looking at the work of others. Even individuals with photographic memories mentally abstract things like lighting, body kinetics, musculature, color theory, etc and produce new work based on those abstractions rather than directly copying original work (unless the artist is intentionally plagiarizing). As a result, all new works produced by humans will have a certain degree of originality to them, regardless of influences due to differences in perception, mental abstraction processes, and life experiences among other factors. Furthermore, humans can produce art without any external instruction or input… give a 5 year old that’s never been exposed to art and hasn’t been shown how to make art a box of crayons and it’s a matter of time before they start drawing.
ML models are closer to highly advanced collage makers that take known images and blend them together in a way that’s convincing at first glance, which is why it’s not uncommon to see elements lifted directly from training data in the images they produce. They do not abstract the same way and by definition cannot produce anything that’s not a blend of training data. Give them no data and they cannot produce anything.
It’s absolutely erroneous to compare them to humans, and I believe it will continue to be so until ML models evolve into something closer to AGI which can e.g. produce stylized work with nothing but photographic input that it’s gathered in a robot body and artistic experimentation.
The first perceptron was explicitly designed to be a trainable visual pattern encoder. Zero assumptions about potential feelings of the ghost in the machine need to be made to conclude the program is probably doing what humans studying art say they assume is happening in their head when you show both of them a series of previous artists' works. This argument is such a tired misdirection.
> What this means, is these systems will fall into different categories of law around copyright and free-use.
No they won't.
A human who uses a computer as a tool (under all the previous qualifications of fair use) is still a human doing something in fair use.
Adding a computer to the workflow of a human doesn't make fair use disappear.
A human can use photoshop, in fair use. They can use a camera. They can use all sorts of machines.
The fact that photoshop is not the same as a human brain is simply a completely unrelated non sequitur. Same applies to AI.
And all the legal protections that are offered to someone who uses a regular computer, to use photoshop in fair use, are also extended to someone who uses AI in fair use.
What we actually need to prove is whether such technology is a net benefit to society all else is essentially hand waving. There is no natural right to poorly named intellectual property and even if there was such a matter would never be decided based on the outcome of a philosophical argument because we don't decide anything that way.
>My issue with this line of argument is that it’s anthropomorphizing machines. It’s fine to compare how humans do a task with how a machine does a task, but in the end they are very different from each other, organic vs hardware and software logic.
There's an entire branch of philosophy that calls these assumptions into question:
>Martin Heidegger viewed humanism as a metaphysical philosophy that ascribes to humanity a universal essence and privileges it above all other forms of existence. For Heidegger, humanism takes consciousness as the paradigm of philosophy, leading it to a subjectivism and idealism that must be avoided.
>Processes of technological and non-technological posthumanization both tend to result in a partial "de-anthropocentrization" of human society, as its circle of membership is expanded to include other types of entities and the position of human beings is decentered. A common theme of posthumanist study is the way in which processes of posthumanization challenge or blur simple binaries, such as those of "human versus non-human", "natural versus artificial", "alive versus non-alive", and "biological versus mechanical".
And? Even if neural networks learn the same way humans do, this is not an argument against taking measures against one's art being used as training data, since there are different implications if a human learns to paint the same way as another human vs. if an AI learns to paint the same way as a human. If the two were exactly indistinguishable in their effects no one would care about AIs, not even researchers.
And yet, some people don't even want their artwork studied in schools. Even if you argue that an AI is "human enough" the artists should still have the right to refuse their art being studies.
Is it strange to you that cars and pedestrians are both subject to different rules? They both utilise friction and gravity to travel along the ground. I'm curious if you see a difference between them, and if you could describe what it is.
Both cars and pedestrians can be videotaped in public, without asking for their explicit permission. That video can be manipulated by a computer to produce an artwork that is then put on public display. No compensation need be offered to anyone.
This is not one artist inspiring another. This is all artists providing their work for free to immensely capitalized corporations for the corporations sole profit.
People keep making metaphors as if the AI is an entity in this transaction: it’s not! The AI is only the mechanism by which corporations launder IP.
>This is all artists providing their work for free to immensely capitalized corporations for the corporations sole profit.
No, the artists would be within their rights to do that if they chose to. This is corporations taking all the work of all artists regardless of the terms under which it was provided.
Would it change your view if only open-source models were allowed to use the art in their training sets? What if a "capitalized corporation" starts using the open-source model?
This is such a nothing argument. Yes, new artists are inspired by other artists and sometimes make art similar to others, but a huge part of learning and doing art is to find a unique style.
But that’s not even the important part of the argument. A lot of artists work for commission, and are hired for their style. If an AI can be trained without explicit permission from their images, they lose work because a user can just prompt “in the style of”.
There’s no real great solution, outside of law, because the possibility of doing that is already here. But I’ve seen this argument so much and it’s just low effort
That is not how artists learn. This is a false equivalence used to justify the imitation and copying of artists’ work.
Artists’ work isn’t derivative in the same way that AI work is. Artists create work based on other sources of inspiration, some of them almost or completely to the disregard of other art.
Many artists don’t even go to art school. And those that do, do not spend most (all) of that time learning how to copy or imitate other artists.
I’m not expressing an opinion of whether GenAI is unethical or illegal - I think that’s a really difficult issue to wrestle with - just that this argument is a post-hoc rationalisation made in ignorance of how good artists work (not to say ignorance of the difference between illustration and art, conceptual art training vs say a foundation course etc).
If that's true, then it should be fine for that human to paint with the brush of his AI tool. Why should that human artist be restricted in the types of tools he uses to create his artwork?
Very true. I was watching a video yesterday learning how to make brush work digitally. While there were examples, they were just examples but the rest was specific techniques and demonstrations.
It is only natural to see a moral difference between people going to school and learn from your art because they are passionate about it, versus someone on the internet just scraping as many images as possible and automating the learning process.
his handle is KingOfCoders - self-aggrandizing, insufferable, impotent in its attempts to be meta.
He thinks he's an artist because he now has the ability to curate a dataset based off of one artist's work and prompt more art generated in that style. He did it, so clearly he is an artist now.
Most artists are happy to see more people getting into art and joining the community. More artists means the skills of this culture get passed down to the next generation.
Obviously a billion dollar corporation using their work to create an industrial tool designed to displace them is very different.
Artists learning to innovate a trade defend their trade from incursion by bloodthirsty, no-value-adding vampiric middle men attempting to cut them out of the loop.
This is a tired argument; whether or not the diffusion models are "learning", they are a tool of capital to fuck over human artists, and should be resisted for that reason alone.
As a human artist I don't feel the same as you, and I somehow doubt that you care all that much about what we think anyways. You already made up your mind about the tech, so don't feel the need to protect us from "a tool of capital [sic]" to fortify your argument.
imagine being a photographer that takes decades to perfect their craft. sure another student can study and mimic your style. but it's still different than some computer model "ingesting" vast amount of photos and vomiting something similar for $5.99 in aws cpu cost so that some prompt jockey can call themselves an AI artist and make money off of other peoples talent.
i get that this is cynical and does not encompass all ai art, but why not let computers develop their own style wihout ingesting human art? that's when it would actually be AI art
Exactly. Artists should drop the pretentious philosophical bumbling and accept what this is, a fight for their livelihood. Which is, in every sense, completely warranted and good.
Putting blame on the technology and trying to limit public access to software will not go anywhere. Your fight for regulation needs to be with publishers and producers, not with the teen trying to make a cool new wallpaper or the office-man trying to make an aesthetic powerpoint presentation.
> they are a tool of capital to fuck over human artists
So are the copyright and intellectual property laws that artists rely on. From my perspective, you are the capital and I am the one being fucked. So are you ready to abolish all that?
As a representative of a lot of things but hardly any capital who uses diffusion models to get something I would otherwise not pay a human artist for anyway, I testify that, the models are not exclusively what you describe them to be.
I do not support indiscriminate banning of anything and everything that can potentially be used to fuck someone over.
Human beings and LLMs are essentially equivalent, and their processes of "learning" are essentially equivalent, yet human artists are not affected by tools like Nightshade. Odd.
Sigh. Once again: I always love it when techbros say that AI learning and human learning are exactly the same, because reading one thing at a time at a biological pace and remembering takeaway ideas rather than verbatim passages is obviously exactly the same thing as processing millions of inputs at once and still being able to regurgitate sources so perfectly that verbatim copyrighted content can be spit out of an LLM that doesn't 'contain' its training material.
I'm just glad that so many more people have caught on to the bullshit than this time last year, or even six months ago.
I really don't even get the endgame. Art gets "democratized", so anyone who doesn't want their style copied stops putting stuff on the internet, and eventually all human art is trained, so the only new contributions are genAI. Maybe we could get a few centuries worth of stuff of "robot unicorn in the style of Artist X with a flair of Y" permutations, but even ignoring the centipede, that just sounds... boring. worthless.
Since techbros are stupid: "Note that people could always do these kinds of repurposing, and it was never a problem from a copyright perspective. We have a problem now because those things are being done (1) in an automated way (2) at a billionfold greater scale (3) by companies that have vastly more power in the market than artists, writers, publishers, etc. Incidentally, these three reasons are also why AI apologists are wrong when claiming that training image generators on art is just like artists taking inspiration from prior works."
A human artist does not need to look at and memorize 100000 pictures in any span of time, period. Current AI does.
We needed huge amounts of human labor to fund and build Versailles. I'm sure many died as a result. Now we have machines that save many of those lives and labor.
I'm not sold on your argument. I'm not an artist but I don't see how an artist using Nightshade is breaking the law. From an anti-AI point of view, you illegally took my artwork and used it. How is it my fault you didn't understand what you were stealing?
Currently it's not legally defined what is stealing and what is fair use with respect to this, which is of course why k feel it is such a strong issue, however, poisoning data and intentionally masking it to hide said poisoning is rather blatantly illegal.
Rather clearly I think most people support individual IP protection and that's not really a contested issue, however what is fair use and where things fall in that gray area is where things do get dicey.
Duty of Care, additionally from an alternative perspective there are a number of laws covering the transmission of data that is intended to cause harm to another system.
(I understand that this is not a popular point, but I really want to emphasize that I am talking about what is _currently legal_ right now, not at all about the ethicality of large companies using people's data. The latter is a much harder topic.
This is mainly about what the law considers to be legal or not legal and is trying to avoid the more emotional side of the topic.)
This is excellent. We need more tools like this, for text content as well. For software we need GPL 4 with ML restrictions (make your model open source or not at all). Potentially even DRM for text.
Paper is here: https://arxiv.org/abs/2310.13828
This seems to introduce levels of artifacts that many artists would find unacceptable: https://twitter.com/sini4ka111/status/1748378223291912567
The rumblings I'm hearing are that this a) barely works with last-gen training processes b) does not work at all with more modern training processes (GPT-4V, LLaVA, even BLIP2 labelling [1]) and c) would not be especially challenging to mitigate against even should it become more effective and popular. The Authors' previous work, Glaze, also does not seem to be very effective despite dramatic proclamations to the contrary, so I think this might be a case of overhyping an academically interesting but real-world-impractical result.
[1]: Courtesy of /u/b3sn0w on Reddit: https://imgur.com/cI7RLAq https://imgur.com/eqe3Dyn https://imgur.com/1BMASL4
The screenshots you sent in [1] are inference, not training. You need to get a Nightshaded image into the training set of an image generator in order for this to have any effect. When you give an image to GPT-4V, Stable Diffusion img2img, or anything else, you're not training the AI - the model is completely frozen and does not change at all[0].
I don't know if anyone else is still scraping new images into the generators. I've heard somewhere that OpenAI stopped scraping around 2021 because they're worried about training on the output of their own models[1]. Adobe Firefly claims to have been trained on Adobe Stock images, but we don't know if Adobe has any particular cutoffs of their own[2].
If you want an image that screws up inference - i.e. one that GPT-4V or Stable Diffusion will choke on - you want an adversarial image. I don't know if you can adversarially train on a model you don't have weights for, though I've heard you can generalize adversarial training against multiple independent models to really screw shit up[3].
[0] All learning capability of text generators come from the fact that they have a context window; but that only provides a short term memory of 2048 tokens. They have no other memory capability.
[1] The scenario of what happens when you do this is fancifully called Habsburg AI. The model learns from it's own biases, reinforcing them into stronger biases, while forgetting everything else.
[2] It'd be particularly ironic if the only thing Nightshade harms is the one AI generator that tried to be even slightly ethical.
[3] At the extremes, these adversarial images fool humans. Though, the study that did this intentionally only showed the images for a small period of time, the idea being that short exposures are akin to a feed-forward neural network with no recurrent computation pathways. If you look at them longer, it's obvious that it's a picture of one thing edited to look like another.
Hey you know what might not be AI generated post-2021? Almost everything run through Nightshade. So given it's defeated, which is pretty likely, artists have effectively tagged their own work for inclusion.
7 replies →
Correct me if I'm wrong but I understand image generators as relying on auto-labeled images to understand what means what, and the point of this attack to make the auto-labelers mislabel the image, but as the top-level comment said it's seemingly not tricking newer auto-labelers.
1 reply →
Even if no new images are being scraped to train the foundation text-to-image models, you can be certain that there is a small horde of folk still scraping to create datasets for training fine-tuned models, LoRAs, Textual Inversions, and all the new hotness training methods still being created each day.
If it doesn't work during inference I really doubt it will have any intended effect during training, there is simply too much signal and the added adversarial noise works on the frozen and small proxy model they used (CLIP image encoder I think) but it doesn't work on a larger model and trained on a different dataset, if there is any effect during training it will probably just be the model learning that it can't take shortcuts (the artifacts working on the proxy model showcase gaps in its visual knowledge).
Generative models like text-to-image have an encoder part (it could be explicit or not) that extract the semantic from the noised image, if the auto-labelers can correctly label the samples then the encoded trained on both actual and adversarial images will learn to not take the same shortcuts that the proxy model has taken making the model more robust, I cannot see an argument where this should be a negative thing for the model.
The context windows of LLMs are now significantly larger than 2048 tokens, and there are clever ways to autopopulate context window to remind it of things.
[3] sounds really interesting - do you have a link?
1 reply →
Yeah. At worst a simple img2img diffusion step would mitigate this, but just eyeballing the examples, traditional denoisers would probably do the job?
Denoising is probably a good preprocessing step anyway.
It’s a common preprocessing step and I believe that’s how glaze (this lab’s previous work) was defeated.
I can’t really see any difference in those images on the Twitter example when viewing it on mobile
The animation when you change images makes it harder to see the difference, I opened the three images each in its own tab and the differences are more apparent when you change between each other instantly.
8 replies →
At full size it's super obvious - I made a side-by-side:
https://i.imgur.com/I6EQ05g.png
7 replies →
Something similar to jpeg artifacts on any surface with a normally smooth color gradient, in some cases rather significant.
I didn't see it immediately either, but there's a ton of added noise. The most noticeable bit for me was near the standing person's bent elbow, but there's a lot more that becomes obvious when flipping back and forth between browser tabs instead of swiping on Twitter.
look at the green drapes to the right, or any large uniform colored space. It looks similar to bad JPEG artifacts.
I don't have great vision, but me neither. They're indistinguishable to me (likewise on mobile).
1 reply →
It's really noticeable on desktop, like compressing an 800kb jpeg to 50kb. Maybe on mobile you won't notice, but on desktop the image looks blown out.
It took me a minute too but on the fast you can see some blocky artifacting by the elbow and a few spots elsewhere like curtain upper left.
The gradient on the bat has blocks in it instead of being smooth.
Maybe it's more about "protecting" images that artists want to publicly share to advertise work, but it's not appropriate for final digital media, etc.
In short, anti-AI watermark.
1 reply →
Seems obvious that the people stealing would be adjusting their process to negate these kinds of countermeasures all the time. I don't see this as an arms race the artists are going to win. Not like the LLM folks can consider actually paying their way...the business plan pretty much has "...by stealing everything we can get our hands on..." in the executive summary.
Sir /u/b3nsn0w is courteous, `/nod`.
The artifacts are a non-issue. It's intended images with nightshade are intended to be silently scrapped and avoid human filtering.
The artifacts are extremely an issue for artists who don't want their images damaged for the possibility of them not being trained by AI.
It's a bad tradeoff.
5 replies →
do you mean scrapped or scraped?
1 reply →
> The artifacts are a non-issue.
According to which authority?
Huge market for snake oil here. There is no way that such tools will ever win, given the requirements the art remain viewable to human perception, so even if you made something that worked (which this sounds like it doesn’t) from first principles it will be worked around immediately.
The only real way for artists or anyone really to try to hold back models from training on human outputs is through the law, ie, leveraging state backed violence to deter the things they don’t want. This too won’t be a perfect solution, if anything it will just put more incentives for people to develop decentralized training networks that “launder” the copyright violations that would allow for prosecutions.
All in all it’s a losing battle at a minimum and a stupid battle at worst. We know these models can be created easily and so they will, eventually, since you can’t prevent a computer from observing images you want humans to be able to observe freely.
The level of claims accompanied by enthusiastic reception from a technically illiterate audience make it sound, smell, and sound like snake oil without much deep investigation.
There is another alternative to the law. Provide your art for private viewing only, and ensure your in person audience does not bring recording devices with them. That may sound absurd, but it's a common practice during activities like having sex.
That doesn't sound like a viable business model. There seems to be a non-trivial bootstrap problem involved -- how do you become well-known enough to attract audiences to private venues in sufficient volume to make a living? -- and would in no way diminish demand for AI-generated artwork which would still continue to draw attention away from you.
The thing is people want the benefits of having their stuff public but not bear the costs. Scraping has been mostly a solved problem especially when it comes to broad crawling. Put it under a login, there, no more AI "stealing" your work.
6 replies →
This would just create a new market for art paparazzis who would find any and all means to inflitrate such private viewings with futuristic miniature cameras and other sensors and selling it for a premium. Less than 24 hours later the files end up on hundreds or thousands of centralized and decentralized servers.
I'm not defending it. Just acknowledging the reality. The next TMZ for private art gatherings is percolating in someone's garage at the moment.
6 replies →
True I can imagine that kind of thing becoming popular.
>There is no way that such tools will ever win, given the requirements the art remain viewable to human perception
On the other hand, the adversarial environment might push models towards a representation more aligned with human perception, which is neat.
The ol' Analog Gap. https://en.m.wikipedia.org/wiki/Analog_hole
> Huge market for snake oil here.
This tool is free, and as far as I can tell it runs locally. If you're not selling anything, and there's no profit motive, then I don't think you can reasonably call it "snake oil".
At worst, it's a waste of time. But nobody's being deceived into purchasing it.
If this is a danger from "snake oil" of this type, it'd be from the other side, where artists are intentionally tricked into believing that tools like this mean that AI isn't or won't be a threat to their copyrights in order to get them to stop opposing it so strongly, when in fact the tool does nothing to prevent their copyrights from being violated.
I don't think that's the intention of Nightshade, but I wouldn't put past someone to try it.
There's an academic paper being published.
Snake oil for the sake of getting published is a very real problem that does exist.
Religion is also deceptive and snake-oil even if it does not involve profit driven motivations.
2 replies →
This is the hard reality. There is no putting this genie back in the bottle.
The only way to be an artist now is to have a unique style of your own, and to never make it online.
"and to never make it online."
So then of course, you also cannot sell your work, as those might put it online. And you cannot show your art to big crowds, as some will make pictures and put it online. So ... you can become a literal underground artists, where only some may see your work. I think only some will like that.
But I actually disagree, there are plenty of ways to be an artist now - but most should probably think about including AI as a tool, if they still want to make money. But with the exception of some superstars, most artists are famously low on money - and AI did not introduce this. (all the professional artists I know, those who went to art school - do not make their income with their art)
3 replies →
Everything old is new again. It's the same thing with any DRM that happens on the client side. As long as it's viewable by humans, someone will figure out a way to feed that into a machine.
"A law, ie, leveraging state backed violence to deter the things they don’t want."
We all know what a law is you don't need to clarify. It makes your prose less readable.
Other people pointed out they appreciated this prose. It’s easy to forget what exactly people are asking for when they talk about regulating the training of machine learning models.
> leveraging state backed violence to deter the things they don’t want
I just want to say: I really appreciate the stark terms in which you've put this.
The thing that has come to be called "intellectual property" is actually just a threat of violence against people who arrange bytes in a way that challenges power structures.
I heard that flooding the net with AI generated art would do much much more harm to generative AI than this whatever is this. Yes, this must be some snake oil salesman, those take it seriously turn AIs own weapon against AI.
I'm thinking — is it possible to create something on a global level similar to what they did in Snapchat: some sort of image flickering that would be difficult to parse, but still acceptable for humans?
Sorry i do not use Snapchat and with googeling "Snapchat image flickering" i did not find a good result. Could you elaborate this a bit more or provide me with a link where this is described? Thank you very much. :)
If humans can process it, you can train a model to do the same.
You don’t need it to visible. You only need it to be scrapped to poison the models. I think that’s the idea.
My guess. Is that at some poi t of time You will not be able to use any generated image or video in commercial. Because of 100% copyright claim for using parts of copyrighted image. Like YouTube those days. When some random beeps matches with someone music...
It should be like that. I agree
A few months ago I made a proof-of-concept on how finetuning Stable Diffusion XL on known bad/incoherent images can actually allow it to output "better" images if those images are used as a negative prompt, i.e. specifying a high-dimensional area of the latent space that model generation should stay away from: https://news.ycombinator.com/item?id=37211519
There's a nonzero chance that encouraging the creation of a large dataset of known tampered data can ironically improve generative AI art models by allowing the model to recognize tampered data and allow the training process to work around it.
Great lora post, thanks for sharing this again! Not sure how I missed as I'm especially interested in sd content.
This seems like a pretty pointless "arms race" or "cat and mouse game". People who want to train generative image models and who don't care about what artists think about it at all can just do some basic post-processing on the images that is just enough to destroy the very carefully tuned changes this Nightshade algorithm makes. Something like resampling it to slightly lower resolution and then using another super-resolution model on it to upsample it again would probably be able to destroy these subtle tweaks without making a big difference to a human observer.
In the future, my guess is that courts will generally be on the side of artists because of societal pressures, and artists will be able to challenge any image they find and have it sent to yet another ML model that can quickly adjudicate whether the generated image is "too similar" to the artist's style (which would also need to be dissimilar enough from everyone else's style to give a reasonable legal claim in the first place).
Or maybe artists will just give up on trying to monetize the images themselves and focus only on creating physical artifacts, similar to how independent musicians make most of their money nowadays from touring and selling merchandise at shows (plus Patreon). Who knows? It's hard to predict the future when there are such huge fundamental changes that happen so quickly!
>Or maybe artists will just give up on trying to monetize the images themselves and focus only on creating physical artifacts, similar to how independent musicians make most of their money nowadays from touring and selling merchandise at shows (plus Patreon).
As is, art already isn't a sustainable career for most people who can't get a job in industry. The most common monetization is either commissions or hiding extra content behind a pay wall.
To be honest I can see more proverbial "Furry artists" sprouting up in a cynical timeline. I imagine like every other big tech that the 18+ side of this will be clamped down hard by the various powers that be. Which means NSFW stuff will be shielded a bit by the advancement and you either need to find underground training models or go back to an artist. .
>need to find underground training models
It's not particularly that hard. The furry nsfw models are already the most well developed and available models you can get right now. And they are spitting out stuff that is almost indistinguishable from regular art.
> This seems like a pretty pointless "arms race" or "cat and mouse game".
If there is any "point" of this, it's that's going to push the AI models to become better at capturing how humans see things.
> musicians make most of their money nowadays from touring and selling merchandise at shows
Be reminded that this is - and has always been - the mainstream model of the lineages of what have come to be called "traditional" and "Americana" and "Appalachian" music.
The Grateful Dead implemented this model with great finesse, sometimes going out of their way to eschew intellectual property claims over their work, in the belief that such claims only hindered their success (and of course, they eventually formalized this advocacy and named it "The Electronic Frontier Foundation" - it's no coincidence that EFF sprung from deadhead culture).
It is a funny appearance (weird viewpoint) that artists are furious loosing their monopily in stealing and cloning components from other artists, recomposing into a similar but new thing.
And that OpenArt on the analogy of OpenSource is a non-existing thing (I know, I know, different things, source code is not for the generic audience and can be hidden on will, unlike art, just having some generative thoughts artefact here ;) )
the point is you could circumvent one nightshade, but as long as the cat and mouse game continues there can be more
This feels like it'll actually help make AI models better versus worse once they train on these images. Artists are basically, for free, creating training data that conveys what types of noise does not change the intended meaning of the image to the artist themselves.
The number of people who are going to be able to produce high fidelity art with off the shelf tools in the near future is unbelievable.
It’s pretty exciting.
Being able to find a mix of styles you like and apply them to new subjects to make your own unique, personalized, artwork sounds like a wickedly cool power to give to billions of people.
In terms of art, population tends to put value not on the result, but origin and process. People will just look down on any art that’s AI generated in a couple of years when it becomes ubiquitous.
> population tends to put value not on the result, but origin and process
I think population tends to value "looks pretty", and it's other artists, connoisseurs, and art critics who value origin and process. Exit Through the Gift Shop sums this up nicely
I disagree. I definitely value modern digital art more than most historical art, because it just looks better. If AI art looks better (and in some cases it does) then I'll prefer that.
3 replies →
This is already the case. Art is a process, a form of human expression, not an end result.
I'm sure OpenAI's models can shit out an approximation of a new Terry Pratchett or Douglas Adams novel, but nobody with any level of literary appreciation would give a damn unless fraud was committed to trick readers into buying it. It's not the author's work, and there's no human message behind it.
6 replies →
https://en.wikipedia.org/wiki/Labor_theory_of_value
According to Marx, value is only created with human labour. This is not just a Marxist theory, it is an observation.
There may be lots of over-priced junk that makes you want to question this idea. But let's not nit-pick on that.
In two years time people will not see any value in AI art, quite correctly because there is not much human labour in creating it.
7 replies →
Nope, but I already look down on artists who refuse to integrate generative AI into their processes.
16 replies →
> Being able to find a mix of styles you like and apply them to new subjects to make your own unique, personalized, artwork sounds like a wickedly cool power to give to billions of people.
And in the process, they will obviate the need for Nightshade and similar tools.
AI models ingesting AI generated content does the work of destroying the models all by itself. Have a look at "Model Collapse" in relation to generative AI.
It'll be about as wickedly tool as the ability to get on the internet, e.g. commoditized, transactional, and boring.
I know this is an unpopular thing to say these days, but I still think the internet is amazing.
I have more access to information now than the most powerful people in the world did 40 years ago. I can learn about quantum field theory, about which pop star is allegedly fucking which other pop star, etc.
If I don't care about the law I can read any of 25 million books or 100 million scientific papers all available on Anna's Archive for free in seconds.
2 replies →
And we only had to alienate millions of people from their labor to do it.
Absolutely agree we should allow people to accumulate equity through effective allocation of their labor.
And I also agree that we shouldn’t build systems that alienate people from that accumulated equity.
Yeah, sadly those millions of people don’t matter in the grand scheme of things and were never going to profit off their work long term
5 replies →
Is this utilitarianism?
Worth it.
Not really. There is a reason why we find realistic painting to be more fascinating than a photo and why some still practice it. The effort put in by another artist does affect our enjoyment.
For me it doesn’t. I’m generating images, realistic, 2.5d, 2d and I like them as much. I don’t feel (or miss) what you described. Or what any other arts guy describes, for that matter. Arts people are different, because they were trained to feel something a normal person wouldn’t. And that’s okay, a normal person without training wouldn’t see how much beauty and effort there is in an algorithm or a legal contract as well.
The word "we" is doing a lot of heavy lifting here. A large majority of consumers can't even tell apart AI-generated from handmade, let alone care who or what made the thing.
1 reply →
I want progressive fees on copyright/IP/patent usage, and worldwide gov cooperation/legislation (and perhaps even worldwide ability to use works without obtaining initial permission, although let's not go into that outlandish stuff)
I want a scaling license fee to apply (e.g. % pegged to revenue. This still has an indirect problem with different industries having different profit margins, but still seems the fairest).
And I want the world (or EU, then others to follow suit) to slowly reduce copyright to 0 years* after artists death if owned by a person, and 20-30 years max if owned by a corporation.
And I want the penalties for not declaring usage** / not paying fees, to be incredibly high for corporations... 50% gross (harder) / net (easier) profit margin for the year? Something that isn't a slap on the wrist and can't be wriggled out of quite so easily, and is actually an incentive not to steal in the first place.)
[*]or whatever society deems appropriate.
[**]Until auto-detection (for better or worse) gets good enough.
IMO that would allow personal use, encourages new entrants to market, encourages innovation, incentivises better behaviour from OpenAI et al.
> And I want the world (or EU, then others to follow suit) to slowly reduce copyright to 0 years* after artists death if owned by a person, and 20-30 years max if owned by a corporation.
Why death at all?
It's icky to trigger soon after death, it's bad to have copyright vary so much based on author age, and it's bad for many works to still have huge copyright lengths.
It's perfectly fine to let copyright expire during the author's life. 20-30 years for everything.
Extremely naive to think that any of this could be enforced to any adequate level. Copyright is fundamentally broken and putting some plasters on it is not going to do much especially when these plasters are several decades too late.
With this "solution" it looks like the world of art enters the cat-and-mouse game the ad blockers were playing for the last decade or two.
I just tested it with Azure AI image classification and it worked - so this cat is yet to adapt to the mouse’s latest idea.
I still feel it is absolutely wrong to roam around the internet and scrape images (without consent) in order to power one’s cash cow AI. I hope more methods to protect artworks (including audio and other formats) become more accessible.
Artists copy from each other all the time. Arguably, culture exists because of copying (folk stories by necessity); copyright makes culture top-down and stagnant, and you can't avoid it because they have the money to shove it right in your face. Who wants trickle-down culture?
2 replies →
I might be missing something because I don't know much about the architecture of either Nightshade or AI art generators, but I wonder if you could try to have a GAN-like architecture (an extra model trying to trick the model) for the part of the generator that labels images to build resistance to Nightshade-like filters.
It doesn't even have to be a full GAN, you only need to train the discriminator side to filter out the data. Clean reference images + Nightshade would be the generator side.
What the article doesn't illustrate is that it destroys fine detail in the image, even in the thumbnails of the reference paper: https://arxiv.org/pdf/2310.13828.pdf
Also... Maybe I am naive, but it seems rather trivial to work around with a quick prefilter? I don't know if tradition denoising would be enough, but worst case you could run img2img diffusion.
reply
The poisoned images aren't intended to be viewed, rather scraped and pass a basic human screen. You wouldn't be able to denoise as you'd have to denoise the entire dataset, the entire point is that these are virtually undetectable from typical training set examples, but they can push prompt frequencies around at will with a small number of poisoned examples.
> You wouldn't be able to denoise as you'd have to denoise the entire dataset
Doing that requires much less compute than training a large generative image model.
4 replies →
Long-term I think the real problem for artists will be corporations generating their own high quality targeted datasets from a cheap labor pool, completely outcompeting them by a landslide.
In the short-to-medium term, we're seeing huge improvements in the data efficiency of generative models. We haven't really started to see self-training in diffusion models, which could improve data efficiency by orders of magnitude. Current models are good at generalisation and are getting better at an incredible pace, so any efforts to limit the progress of AI by restricting access to training data is a speedbump rather than a roadblock.
It will democratize art.
Art is already democratized. It has been for decades. Everyone can pick it up at zero cost. Even you!
The poorest people have historically produced great art. Training a model, however? Expensive. Running it locally? Expensive. Paying the sub? Expensive.
Nothing is being democratized, the only thing this does is devaluing the blood and sweat people have put into their work so FAANG can sell it to lazy suckers.
1 reply →
then it won't be art anymore, it'll just be mountains of shit
sorta like what the laptop did for writing
1 reply →
This is fantastic. If companies want to create AI models, they should license the content they use for the training data. As long as there are not sufficient legal protections and the EU/Congress do not act, tools like these can serve as a stopgap and maybe help increase pressure on policymakers
It's going to be interesting to see how the lawsuits against OpenAI by content creators plays out. If the courts rule that AI generated content is a derivative work of all the content it was trained on it could really flip the entire gen AI movement on its head.
If it were a derivative work[1] (and sufficiently transformational) then it's allowed under current copyright law and might not be the slam dunk ruling you were hoping for.
[1] https://en.wikipedia.org/wiki/Derivative_work
4 replies →
My biggest fear is that the big players will drop a few billion dollars to silence the copyright holders with power go away, and new rules are put in place that will make open-source models that can't do the same essentially illegal.
If the courts do rule that way, I would expect a legislative race between different countries to amend the relevant laws. Visual generative AI is just too lucrative a thing.
…then I'll keep enjoying my Stable Diffusion and pirated models.
> they should license the content they use for the training data
You mean like OpenAI and Adobe ?
Only the free and open source models didn't licensed any content for the training data.
Adobe is training off of images stored in their cloud systems, per their Terms of Service.
OpenAI has provided no such documentation or legal guarantees, and it is still quite possible they scraped all sorts of copyright materials.
14 replies →
There is a small difference between any and all. OpenAI certainly didn't licence all of the image they use for training.
source for OpenAI paying anyone a dime? don't you think that would set a precedent that everyone else deserves their cut?
Isn't this just teaching the models how to better understand pictures as humans do? As long as you feed them content that looks good to a human, wouldn't they improve in creating such content?
You would think the economists at UChicago would have told these researchers that their tool would achieve the opposite effect of what they intended, but here we are.
In this case, the mechanism for how it would work is effectively useless. It doesn't affect OpenAI or other companies building foundation models. It only works on people fine-tuning these foundation models, and only if the image is glazed to affect the same foundation model.
These methods like Glaze usually works by taking the original image chaging the style or content and then apply LPIPS loss on an image encoder, the hope is that if they can deceive a CLIP image encoder it would confuse also other models with different architecture, size and dataset, while changing the original image as little as possible so it's not too noticeable to a human eye. To be honest I don't think it's a very robust technique, with this one they claim that a model instead of seeing for example a cow on grass the model will see a handbag, if someone has access to GPT4-V I want to see if it's able to deceive actually big image encoders (usually more aligned to the human vision).
EDIT: I have seen a few examples with GPT-4 V and how I imagine it wasn't deceived, I doubt this technique can have any impact on the quality of the models, the only impact that this could potentially have honestly is to make the training more robust.
Each time there is an update to training algorithms and in response poisoning algorithms, artists will have to re-glaze, re-mist, and re-nightshade all their images?
Eventually I assume the poisoning artifacts introduced in the images will be very visible to humans as well.
>Like Glaze, Nightshade is computed as a multi-objective optimization that minimizes visible changes to the original image.
It's still noticeably visible.
Yeah, I've seen multiple artists complain about how glazing reduces image quality. It's very noticeable. That seems like an unavoidable problem given how AI is trained on images right now.
I'm glad to see tools like Nightshade starting to pop up to protect the real life creativity of artists. I like AI art, but I do feel conflicted about its potential long term effects towards a society that no longer values authentic creativity.
Is the existence of the AI tool not itself a product of authentic creativity? Does eliminating barriers to image generation not facilitate authentic creativity?
No, it facilitates commoditization. Art – real art – is fundamentally a human-to-human transaction. Once everyone can fire perfectly-rendered perfectly-unique pieces of 'art' at each other, it'll just become like the internet is today: filled with extremely low-value noise.
Enjoy the short term novelty while you can.
4 replies →
To protect an individual's image property rights from image generating AI's -- wouldn't it be simpler for the IETF (or other standards-producing group) to simply create an
AI image exclusion standard
, similar to "robots.txt" -- which would tell an AI data-gathering web crawler that a given image or set of images -- was off-limits for use as data?
https://en.wikipedia.org/wiki/Robots.txt
https://www.ietf.org/
Entities training models have no incentive to follow such metadata. If we accept the premise that "more input -> better models" then there's every reason to ignore non-legally-binding metadata requests.
Robots.txt survived because the use of it to gatekeep valuable goodies was never widespread. Most sites want to be indexed, most URLs excluded by the robots file are not of interest to the search engine anyway, and use of robots to prevent crawling actually interesting pages is marginal.
If there was ever genuine uptake in using robots to gatekeep the really good stuff search engines would've stopped respecting it pretty much immediately - it isn't legally binding after all.
>Entities training models have no incentive to follow such metadata. If we accept the premise that "more input -> better models" then there's every reason to ignore non-legally-binding metadata requests.
Name two entities that were asked to stop using a given individuals' images that failed to stop using them after the stop request was issued.
>Robots.txt survived because the use of it to gatekeep valuable goodies was never widespread. Most sites want to be indexed, most URLs excluded by the robots file are not of interest to the search engine anyway, and use of robots to prevent crawling actually interesting pages is marginal.
Robots.txt survived because it was a "digital signpost" a "digital sign" -- sort of like the way you might put a "Private Property -- No Trespassing" sign in your yard.
Most moral/ethical/lawful people -- will obey that sign.
Some might not.
But the some that might not -- probably constitute about a 0.000001% minority of the population, whereas the majority that do -- probably constitute about 99.99999% of the population.
"Robots.txt" is a sign -- much like a road sign is.
People can obey them -- or they can ignore them -- but they can ignore them only at their own peril!
It's a sign which provides a hint for what the right thing to do in a certain set of circumstances -- which is what the Law is; which is what the majority of Laws are.
People can obey them -- or they can choose to ignore them -- but only at their own peril!
Most will choose to obey them. Most will choose to "take the hint", proverbially speaking!
A few might not -- but that doesn't mean the majority won't!
>If there was ever genuine uptake in using robots to gatekeep the really good stuff search engines would've stopped respecting it pretty much immediately - it isn't legally binding after all.
Again, name two entities that were asked to stop using a given individuals' images that failed to stop using them after the stop request was issued.
And then what? The scrapers themselves already happily ignore copyright, they won't be inclined to obey a no-ai.txt. So someone would have to enforce the standard. Currently I see no organisation who would be willing to do this or even just technologically able - as even just detecting such scrapers is an extremely hard task.
Nevertheless, I hope that at some not-so-far point in the future there will be more legal guidance about this kind of stuff, i.e. it will be made clear that scraping violates copyright. This still won't solve the problem of detectability but it would at least increase the risk of scrapers, should they be caught.
>The scrapers themselves already happily ignore copyright, they won't be inclined to obey a no-ai.txt.
Name two entities that were asked to stop using a given individuals' images that failed to stop using them after the stop request was issued.
>Currently I see no organisation who would be willing to do this or even just technologically able - as even just detecting such scrapers is an extremely hard task.
// Part of Image Web Scraper For AI Image Generator ingestion psuedocode:
if fileExists("no-ai.txt") {
} else {
};
See? Nice and simple!
Also -- let me ask you this -- what happens to the intellectual property (or just plain property) rights of Images on the web after the author dies? Or say, 50 years (or whatever the legal copyright timeout is) after the author dies?
Legal grey area perhaps?
Also -- what about Images that exist in other legal jurisdictions -- i.e., other countries?
How do we know what set of laws are to apply to a given image?
?
Point is: If you're going to endorse and/or construct a legal framework (and have it be binding -- keep in mind you're going to have to traverse the legal jurisdictions of many countries, many countries!) -- you might as well consider such issues.
Also -- at least in the United States, we have Juries that can override any Law (Separation of Powers) -- that is, that which is considered "legally binding" -- may not be quite so "legally binding" if/when properly explained to a proper jury in light of extenuating (or just plain other) circumstances!
So kindly think of these issues prior to making all-encompasing proposals as to what you think should be "legally binding" or not.
I comprehend that you are just trying to solve a problem; I comprehend and empathize; but the problem might be a bit greater than you think, and there might be one if not serveral unexplored partial/better (since no one solution, legal or otherwise, will be all-encompassing) solutions -- because the problem is so large in scope -- but all of these issues must be considered in parallel -- or errors, present or future will occur...
2 replies →
Setting aside the efficacy of this tool, I would be very interested in the legal implications of putting designs in your art that could corrupt ML models.
For instance, if I set traps in my home which hurt an intruder we are both guilty of crimes (traps are illegal and are never considered self defense, B&E is illegal).
Would I be responsible for corrupting the AI operator's data if I intentionally include adversarial artifacts to corrupt models, or is that just DRM to legally protect my art from infringement?
edit:
I replied to someone else, but this is probably good context:
DRM is legally allowed to disable or even corrupt the software or media that it is protecting, if it detects misuse.
If an adversarial-AI tool attacks the model, it then becomes a question of whether the model, having now incorporated my protected art, is now "mine" to disable/corrupt, or whether it is in fact out of bounds of DRM.
So for instance, a court could say that the adversarial-AI methods could only actively prevent the training software from incorporating the protected media into a model, but could not corrupt the model itself.
None whatsoever. There is no right to good data for model training, nor does any contractual relationship exist between you and and a model builder who scrapes your website.
If you're assuming this is open-shut, you're wrong. I asked this specifically as someone who works in security. A court is going to have to decide where the line is between DRM and malware in adversarial-AI tools.
4 replies →
The way Nightshade works (assuming it does work) is by confusing the features of different tags with each other. To argue that this is illegal would be to argue that mistagging a piece of artwork on a gallery is illegal.
If you upload a picture of a dog to DeviantArt and you label it as a cat, and a model ingests that image and starts to think that cats look like dogs, would anybody claim that you are breaking a law? If you upload bad code to Github that has bugs, and an AI model consumes that code and then reproduces the bugs, would anyone argue that uploading badly written code to Github is a crime?
What if you uploaded some bad code to Github and then wrote a comment at the top of the code explaining what the error was, because you knew that the model would ignore that comment and would still look at the bad code. Then would you be committing a crime by putting that code on Github?
Even if it could be proven that your intention was for that code or that mistagged image to be unhelpful to training, it would still be a huge leap to say that either of those activities were criminal -- I would hope that the majority of HN would see that as a dangerous legal road to travel down.
That’s like asking if lying on a forum is illegal
No, it's much closer to (in fact, it is simply) asking if adversarial AI tools count as DRM or as malware. And a court is going to have to decide whether the model and or its output counts as separate software, which it is illegal for DRM to intentionally attack.
DRM can, for instance, disable its own parent tool (e.g. a video game) if it detects misuse, but it can't attack the host computer or other software on that computer.
So is the model or its output, having been trained on my art, a byproduct of my art, in which case I have a legal right to 'disable' it, or is it separate software that I don't have a right to corrupt?
1 reply →
I see it as no different than mapmakers inventing a nonexistent alley, to check who copies their maps verbatim ("trap street"). Even if this caused, for example, a car crash because of an autonomous driver, the onus I think would be on the one that made the car and used the stolen map for navigation, and not on the one that created the original map.
https://en.wikipedia.org/wiki/Trap_street
Japan is considering it, I think? https://news.ycombinator.com/item?id=38615280
How would that situation be remotely related?
I find the AI training topic interesting, because it's really data/information that is involved. Forget about the fact that it's images or stories or Reddit posts, it's all data.
We are born and then exposed to the torrent of data from the world around us, mostly fed to us by other humans, this is what models are trying to tap.
Unfortunately our learning process is completely organic and takes decades and decades and decades; there's no way to put a model through this easily.
Perhaps we need to seed the web with AI agents who converse and learn as much like regular human beings as possible and assemble the dataset that way. Although having an agent browse and find an image to learn to draw from is still gonna make people reee even if that's exactly what a young and aspiring human artist would be doing.
Don't talk about humans being sacred; we already voted to let corporations be people, for the 1% to exist and "lobby", breaking our democracy so that they can get tax breaks and make corrupt under the table deals. None of us stopped that from happening...
Does it survive AI upscaling or img2img? If not - then it's useless. Nobody trains AI models without any preprocessing. This is basically a tool for 2022.
For this to work, wouldn't you have to have an enormous number of artists collaborating on "poisoning" their images the same way (cow to handbag) while somehow keeping it secret form ai trainers that they were doing this? It seems to me that even if the technology works perfectly as intended, you're effectively just mislabeling a tiny fraction of the training data.
1. They don't need an enormous number of artists; the research paper showed significant results with even 50 poisoned image samples in the dataset, which is enough to be contained in even a single artist's online gallery.
2. They don't need to keep it a secret; the goal is to remove these images from the training data, in a way that would be much more efficient than simply adding a "please don't include my art in your ai scraper" message next to your pictures.
In so far as anger goes against AIs being trained on particular intellectual properties.
A made up scenario¹ is that a person who is training an AI, goes to the local library and checks out 600 books on art. The person then lets the AI read all of them. After which they are returned to the library and another 600 books are borrowed
Then we can imagine the AI somehow visiting a lot of museums and galleries.
The AI will now have been trained on the style and looks of a lot of art from different artists
All the material has been obtained in a legal manner.
Is this an acceptable use?
Or can an artist still assert that the AI was trained with their IP without consent?
Clearly this is one of the ways a human would go about learning about styles, techniques etc..
¹ Yes you probably cannot borrow 600 books at a time. How does the AI read the books? I dont know. Simplicity would be that the researcher takes a photo of each page. This would be extremmly slow but for this hypothetical it is acceptable.
I think the key difference here is that the most prominent image generation AIs are commercial and for-profit. The scenarios you describe are comparing a commercial AI to a private person. You cannot get a library card for a company, and you cannot bring a photography crew to a gallery without permission.
I’m completely flabbergasted by the number of comments implying copyright concepts such as “fair use” or “derivative work” apply to trained ML models. Copyright is for _people_, as are the entailing rights, responsibilities and exemptions. This has gone far beyond anthropomorphising and we need to like get it together, man!
You act like computers and ML models aren't just tools used by people.
What did I write to give you that impression?
7 replies →
No one is saying a model is the legal entity. The legal entities are still people and corporations.
Oh come on, you’re being insincere. Wether or not the model is learning from the work just like people is hotly debated as if it would make a difference. Fair use is even brought up. Fair use! Even if it applied, these training sets collate all of everything
I feel like I’m taking crazy pills TBQH
I really don't understand the anxiety of artists towards AI - as if creatives haven't always borrowed and imitated. Every leading artist has had acolytes, and while it's true no artist ever had an acolyte as prodigiously productive as AI will be, I don't see anything different between a young artist looking to Picasso for cues and Stable Diffusion or DALL-E doing the same. Styles and methods haven't ever been subject to copyright - and art would die the moment that changed.
The only explanation I can find for this backlash is that artists are actually worried just like the rest of us that pretty soon AI will produce higher quality more inventive work faster and more imaginatively than they can - which is very natural, but not a reason to inhibit an AI's creative education.
This has been litigated over and over again, and there have been plenty of good points made and concerns raised over it by those who it actually affects. It seems a little bit disingenuous (especially in this forum) to say that that conclusion is the "only explanation" you can come up with. And just to avoid prompting you too much: trust me, we all know or can guess why you think AI art is a good thing regardless of any concerns one might bring up.
[flagged]
Could you please stop posting unsubstantive comments and flamebait? You've unfortunately been doing it repeatedly. It's not what this site is for, and destroys what it is for.
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
4 replies →
Imitation isn’t the problem so much as it is that ML generated images are composed of a mush of the images it was trained on. A human artist can abstract the concepts underpinning a style and mimic it by drawing all-new lineart, coloration, shading, composition, etc, while the ML model has to lean on blending training imagery together.
Furthermore there’s a sort of unavoidable “jitter” in human-produced art that varies between individuals that stems from vastly different ways of thinking, perception of the world, mental abstraction processes, life experiences, etc. This is why artists who start out imitating other artists almost always develop their imitations into a style all their own — the imitations were already appreciably different from the original due to the aforementioned biases and those distinctions only grow with time and experimentation.
There would be greatly reduced moral controversy surrounding ML models if they lacked that mincemeat/pink slime aspect.
I love it. This undermines the notion of ground truth. What separates correct information from incorrect information? Maybe nothing! I love how they acknowledge the never ending attack versus defense game. In stark contrast to "our AI will solve all your problems".
Won't a simple downsample->upsample be the antidote?
No, it's resistant to transformation. Rotation, cropping, scaling, the image remains poisonous. The only antidote known currently is active artist cooperation.
Or Img2Img.
How do you train your upsampler? (Also: why are you seeking to provide an “antidote”?)
> why are you seeking to provide an “antidote”
I think it's worthwhile for such discussion to happen in the open. If the tool can be defeated through simple means, it's better for everybody to know that, right?
2 replies →
I apologize. I was trying to respond to inflammatory language ("poison") with similarly hyperbolic terms, and I should know better than to do that.
Let me rephrase: Would AI-powered upscaling/downscaling (not a simple deterministic mathematical scaling) not defeat this at a conceptual level?
>why are you seeking to provide an “antidote”
To train a model on the data.
14 replies →
Why would you train one?
Doing the work to increase OpenAIs moat
Obviously AIs can just train on images that aren't poisoned.
Is it possible to reliably detect whether an image is poisoned? If not then it achieves the goal of punishing entities which indiscriminately harvest data.
3 replies →
Only protection is adding giant gaping vaginas to your art, nothing less will deter scraping. If the Email spam community showed us something in the last 40 years is that no amount of defensive tech measures will work except financial disincentives.
Any AI art/video/photography/music/etc generator company who generates revenue needs to add watermarks to let the public know its AI generator. This should be forced via legislation in all countries.
If they don't then whatever social network or other services where things can shared/viewed by large groups to millions & are posted publicly need to be labeled "We can not verify veracity of this content."
I want a real internet ..this AI stuff is just triple fold increasing fake crap on the Internet and in turn / time our trust in it!
For visual artists who don't want visible artifacting in the art they feature online, would it be possible to upload these alongside your un-poisoned art, but have them only hanging out in the background? So say having one proper copy and a hundred poisoned copies in the same server, but only showing the un-poisoned one?
Might this "flood the zone" approach also have -some- efficacy against human copycats?
Do not fight the AI, it's a lost cause, embrace it.
I wonder how this tool works if it's actually model independent. My understanding so far was that in principle each possible model has some set of pathological inputs for which the classification will be different than what a user sees - but that this set is basically different for each model. So did they actually manage to build an "universal" poison? If yes, how?
It misclassifies objects in clip, which is used for label generation.
I wonder if this is illegal in some countries. In France for example, there is the following law: "Obstructing or distorting the operation of an automated data processing system is punishable by five years' imprisonment and a fine of €150,000.".
If you ask me, this is 100% applicable in this case, so I wonder what a judge would rule.
Remember when the music industry tried to use technology to stop music pirating?
This will work about as well...
Oh, I forget, fighting music pirating was considered an evil thing to do on HN. "pirating is not stealing, is copyright infringement", right? Unlike training neural nets on internet content which of course is "stealing".
FWIW, you're the only use of the word "steal" in this comment thread.
Many people would in fact argue that training AI on people's art without permission is copyright infringement, since the thing it (according to detractors) does is infringe copyright by generating knockoffs of people's work.
You will see some people use the term "stealing" but they're usually referring to how these AIs are sold/operated by for-profit companies that want to make money off artists' work without compensating them. I think it's not unreasonable to call that "stealing" even if the legal definition doesn't necessarily fit 100%.
The music industry is also not really a very good comparison point for independent artists... there is no Big Art equivalent that has a stranglehold on the legislature and judiciary like the RIAA/MPAA do.
The difference is that “pirating” is mostly done by individuals for private use, whereas training is mostly done by megacorporations looking to make more money.
A more apt comparison is sampling.
AI is sampling other's works.
Musicians can and do sample. They also obtain clearance for commercial works, pay royalties if required, AND credit the samples if required.
AI "art" does none of that.
Musicians overwhelmingly do not even attempt to clear samples. This also isn't a great comparison since samples are taken directly out of the audio, not turned into a part of a pattern used to generate new sounds like what AI generators do with images
1 reply →
I wonder if we know enough about any of these systems to make such claims. This is all predicated on the fact that this tool will be in widespread use. If it is somehow widely used beyond the folks who have seen it at the top of HN, won't the big firms have countermeasures, ready to deploy?
How long will this work?
It's an arms race the bigger players will win, and it undermines the quality of the images. But it feels natural that artists would want to do something since they don't feel like anyone else is protecting them right now.
The intention is good, from an AI-opponent's perspective. I don't think will work practically, though. The drawbacks for actual users of the image galleries, plus the level of complexity involved in poisoning the samples makes this unfeasible to implement at the scale required.
Wonder if the AI companies are already so far ahead that they can use their AI to detect and avoid any poisoning?
Too little, too late. There's already very large high quality datasets to train AI art generators.
The naive idea of wanting to protect artists is actually protecting the monopoly of big companies.
Some projects against this behavior:
https://github.com/syncblob/Obey-AI-Luddites
The opening website is so poor - "what is nightshade" - then a whole paragraph that tells nothing, then another paragraph.. then no examples. This whole description should be reworked to be shorter and more to the point.
Well, at least for sdxl it's not working neither in LoRa nor dreambooth finetunes.
Cute. The effectiveness of any technique like this will be short-lived.
What we really need is clarification of the extent that copyright protection extends to similar works. Most likely from an AI analysis of case law.
If you decrease quality of art, you give AI all the advantage in the market.
> More specifically, we assume the attacker:
> • can inject a small number of poison data (image/text pairs) to the model’s training dataset
I think thoes are bad assumption, labelling is more and more done by some labelling AI.
Usually clip, which is actually how this works — the examples are modified to be misclassified in clip, but look passable to a human.
The image generation models now are at the point where they can produce their own synthetic training images. So I'm not sure how big of an impact something like this would have.
would it have been that hard to include a sample photo and how it looks with the nightshade filter side by side in a 3 page document describing how it would look in great detail
Put a TOC on all your content that says “by using my content for AI you agree to pay X per image” and then send them a bill once you see it in an AI.
Once you see what exactly? “AI” isn’t some image filter from the early 2000s.
Baffling to see anyone argue against this technology when it is a non-issue to any model by simply acquiring only training data you have permission to use.
The reason people are arguing against this technology is that no one is using them in the way you describe. They actually wouldn't even be economically viable in that case.
If it is not economically viable for you to be ethical, then you do not deserve economic success.
Anyone arguing against this technology following the line of reasoning you present is operating in adverse to the good of society. Especially if their only motive is economic viability.
2 replies →
I don't know if asking permission of every copyright holder of every image on the Internet is as simple as you're implying.
I think people 100% have the right to use this on their images, but:
> simply acquiring only training data you have permission to use
Currently it's generally infeasible to obtain licenses at the required scale.
When attempting to develop a model that can describe photos for visually impaired users, I had even tried to reach out to obtain a license from Getty. They repeatedly told me that they don't license images for machine learning[0].
I think it's easy to say "well too bad, it doesn't deserve to exist" if you're just thinking about DALL-E 3, but there's a huge number of positive and far less-controversial applications of machine learning that benefit from web-scale pretraining and foundation models - spam filtering, tumour segmentation, voice transcription, language translation, defect detection, etc.
[0]: https://i.imgur.com/iER0BE2.png
I don't believe it's a "doesn't deserve to exist" situation, because these things genuinely can be used for the public good.
However - and this is a big however - I don't believe it deserves the legal protection to be used for profit.
I am of the opinion that if you train your model on data that you do not hold the rights for, your usage should be handled similarly to most fair use laws. It's fine to use it for your personal projects, for research and education, etc. but it is not OK to use it for commercial endeavors.
3 replies →
What are LLMs that was trained with public domain content only?
I would believe there is enough content out there to get reasonably good results.
Trying to convince an AI it sees something and a human they don’t is probably a losing battle.
This timeline is getting quite similar to the second season of Pantheon.
How is there not a single example on that website?
I think the artists need to agree to stop making art altogether. That ought to get people’s attention. Then the AI people might (be socially pressured or legally forced to) put their tools away.
No, they'll just demand that artists produce more art so they can continue scraping, because if you work in tech you're allowed to be entitled, you're The Face Of The Future and all you're trying to do is Save The World, all these decels are just obstacles to be destroyed.
Sounds like free adversarial data augmentation.
Wouldn't this be applicable to text too?
Another way would be, for every 1 piece of art you make, post 10 AI generated arts, so that the SNR is really bad.
Artist are now fully dependent on Software Engineers for protecting the future of their career lol
Why are there no examples?
My hope is these type of "poisoning tools" become ubiquitous for all content types on the web, forcing AI companies to, you know, license things.
This is the DRM problem again.
However much we might wish that it was not true, ideas are not rivalrous. If you share an idea with another person, they now have that idea too.
If you share words on paper, then someone with eyes and a brain might memorize them (or much more likely, just grasp and retain the ideas conveyed in the words).
If you let someone hear your music, then the ideas (phrasing, style, melody, etc) in that music are transferred.
If you let people see a visual work, then the stylistic and content elements of that work are potentially absorbed by the audience.
We have copyright to protect specific embodiments, but mostly if you try to share ideas with others without letting them use the ideas you shared, then you are in for a life of frustration and escalating arms race.
I completely sympathize with anyone who had a great idea and spent a lot of effort to realize it. If I invented/created something awesome I would be hurt and angry if someone “copied” it. But the hard cold reality is that you cannot “own” an idea.
> But the hard cold reality is that you cannot “own” an idea.
The above comment is true about the properties of information, as explained via the lens of economics. [1]
However, one ignores ownership as defined by various systems (including the rule of law and social conventions) at one's own peril. Such systems can also present a "hard cold reality" that can bankrupt or ostracize you.
[1] Don't let the apparent confidence and technicality of the language of economists fool you. Economics isn't the only game in town. There are other ways to model and frame the world.
[2] Dangling footnote warning. I think it is instructive to recognize that the field of economics has historically shown a kind of inferiority complex w.r.t. physics. Some economists ascribe to the level of rigor found in physics and that is well and good, but perhaps that effort should not be taken too seriously nor too far, since economics as a field operates at a different level. IMO, it would be wise for more in the field to eat a slice of humble pie.
[3] Ibid. It is well-known that economists can be "hired guns" used to "prove" a wide variety of things, many of which are subjective. My point: you can hire an economist to shore up one's political proposals. Is the same true of physicists? Hopefully not to the same degree. Perhaps there are some cases of hucksterism, but nothing like the history of economists-wagging-the-dog! At some point, the electron tunnels or it does not.
There are other games in town.
But whatever game gives the most predictive power is going to win.
2 replies →
Many terms of art from economics are probably not widely-known here.
> In economics, a good is said to be rivalrous or a rival if its consumption by one consumer prevents simultaneous consumption by other consumers, or if consumption by one party reduces the ability of another party to consume it. - Wikipedia: Rivalry (economics)
Also: we should recognize that stating something as rivalrous or not is descriptive (what exists) not normative (what should be).
I think ideas being rivalrous is intrinsic, and therefore descriptive and normative.
3 replies →
We're not trying to keep the AI from learning general ideas, we're trying to keep it from memorizing specific expressions[0]. There's a growing body of research to show that these models are doing a lot of memorizing, even if they're not regurgitating that data. For example, Google's little "ask GPT to repeat a word forever" trick, which will make GPT-4 spit out verbatim training data[1].
If there was a training process that let us pick a minimal sample of examples and turn it into a general purpose art generator or text generator, I think people would have been fine with that. But that's not what any of these models do. They were trained on shittons of creative expression, and there's statistical evidence that the models retain that expression, in a way that is fundamentally different from how humans remember, misremember, adapt, remix, and/or "play around with" other people's creativity.
[0] You called these "embodiments", but I believe you're trying to invoke the idea/expression divide, so I'll run with that.
[1] Or at least it did. OpenAI now filters out conversations that trip the bug.
I don't see the parallel between this offensive tool and DRM. I could, say buy a perpetual license to an image from the artist, so that I can print it and put it on my wall, while it can simultaneously be poisonous to an AI system. I can even steal it and print it, while it is still poisonous to an AI system.
The closest parallel I can think of is that humans can ingest chocolate but dogs should not.
What you've described is the literal, dictionary definition of Digital Rights Management - a technology to restrict the use of a digital asset beyond the contractually-agreed terms. Copying is only one of many uses that the copyright-holder may wish to prevent. The regional lockout on a DVD had nothing to do with copy-protection, but it was still DRM.
It's about the arm's race: DRM will always be cracked (with a sufficiently motivated customer.) AI poisoning will always be cracked (with a sufficiently motivated crawler.)
A huge amount of DRM effort has been spent in the watermarking area, which is similar, but not exactly the same.
This doesn’t stop anyone from viewing or scraping the work, though, so in no way is it DRM. It just causes certain methods of computer interpretation of an image to interpret it in an odd way vs. human viewers. They can still learn from them.
It absolutely is DRM, just a different form than media encryption. It's a purely-digital mechanism of enforcing rights.
16 replies →
Being able to fairly monetise your creative work and put food on the table is a bit rivalrous though, don’t you think?
No, I disagree. There is no principle of the universe or across human civilizations that says that you have a right to eat because you produced a creative work.
The way societies work is that the members of the society contribute and benefit in prescribed ways. Societies with lots of excess production may at times choose to allow creative works to be monetized. Societies without much surplus are extremely unlikely to do so, eg a society with not enough food for everyone to eat in the middle of a famine is extremely unlikely to feed people who only create art; those people will have to contribute in some other way.
I think it is a very modern western idea (less than a century old) that many artists can dedicate themselves solely to producing the art they want to produce. In all other times artists either had day jobs or worked on commission.
6 replies →
No, rivalrous has a specific meaning https://en.wikipedia.org/wiki/Rivalry_(economics)
The tragedy of "your business model is not my problem" as a spreading idea is that while you're right since distribution is where the money is (not creation), intellectual property is de-facto weakened today and IP piracy is widely considered an acceptable thing.
So is sabotaging solutions that would make creative work of the same (or superior) quality more affordable. Your ability to produce expensive illustrations hinders my ability to produce cheap textbooks.
Not everybody equates automated scraping for training models and human experience. Just like any other “data wants to be free” type of discussion, the philosophical and ethical considerations are anything but cut-and-dried, and they’re far more consequential than the technical and economics-in-a-vacuum ones. The general public will quite possibly see things differently than the “oh well, artists— that’s the free market for ya, and you lost” crowd.
You don't copyright ideas, you copyright works. And these artists' productions are works, not abstract ideas, with copyrights, and they are being violated. This is simple law. Why do people have such a hard time with this? Are you the one training the models and you need to find a cognitive escape out of the illegality and wrong-doing of your activities?
It’s not obvious to me that using a copyrighted image to train a model is copyright infringement. It’s certainly not copyright infringement when used to train a human who may end up creating works that are influenced by (but not copies of) the original works.
Now, if the original copyrighted work can be extracted or reproduced from the model, that’s obviously copyright infringement.
OpenAI etc should ensure they don’t do that.
6 replies →
>This is simple law. Why do people have such a hard time with this?
Because this isn’t simple law. It feels like simple infringement, but there’s no actual copying going on. You can’t open up the database and find a given duplicate of a work. Instead you have some abstraction of what it takes to get to a given work.
Also it’s important to point out that nothing in the law is sure. A good lawyer, a sympathetic judge, a bored/interested/contrarian juror, etc can render “settled law” unsettled in an instant. The law is not a set of board game rules.
6 replies →
> This is simple law.
“One may well ask: ‘How can you advocate breaking some laws and obeying others?’ The answer lies in the fact that there are two types of laws: just and unjust. I would be the first to advocate obeying just laws. One has not only a legal but a moral responsibility to obey just laws. Conversely, one has a moral responsibility to disobey unjust laws. I would agree with St. Augustine that ‘an unjust law is no law at all.’”
1 reply →
Illegality and wrongdoing are completely distinct categories.
I'm not convinced that most copyright infringements are immoral regardless of their legal status.
If you post your images for the world to see, and someone uses that image, you are not harmed.
The idea that the world owes you something after you deliberately shared it with others seems bizarre.
6 replies →
If it were true, then we wouldn't have that great difference in opinions on this topic.
2 replies →
That may be the law, although we are probably years of legal proceedings away from finding out.
It obviously is not "simple law".
Nothing is being reproduced. Just the ideas being reused.
> ... you cannot “own” an idea.
Let's talk about ownership in a broader sense. In practice, one cannot effectively own (retain possession of) something without some combination of physical capability or coercion (or threat of coercion). Meaning: maintaining ownership of anything (physical or otherwise) often depends on the rule of law.
Then let's use a more precise term that is also present in law: monopoly.
You can't monopolize an idea.
Copyright law is a prescription, not a description. Copyright law demands that everyone play along with the lie that is intellectual monopoly. The effectiveness of that demand depends on how well it can be enforced.
Playing pretend during the age of the printing press may have been easy enough to coordinate, but it's practically impossible here in the digital age.
If we were to increase enforcement to the point of effectiveness, then what society would be left to participate? Surely not a society I am keen to be a part of.
15 replies →
Kick ass.
I now declare that I own Fortnite.
Where’s my money, Epic?
[flagged]
I don't think we have to compare the process when the end result (another's work being influenced by what they have previously seen) is the same.
1 reply →
[dead]
[dead]
[dead]
[flagged]
My issue with this line of argument is that it’s anthropomorphizing machines. It’s fine to compare how humans do a task with how a machine does a task, but in the end they are very different from each other, organic vs hardware and software logic.
First, you need to prove that generative AI works fundamentally the same way as humans at the task of learning. Next you have to prove that it recalls information in the same way as humans. I don’t think anyone would say these are things that we can prove, let alone believe they do. So what we get is comments like they are similar.
What this means, is these systems will fall into different categories of law around copyright and free-use. What’s clear is that there are people who believe that they are harmed by the use of their work in training these systems and it reproducing that work in some manner later on (the degree to which that single work or the corpus of their work influences that final product is an interesting question). If your terms of use/copyright/license says “you may not train on this data”, then should that be protected in law? If a system like nightshade can effectively influence a training model enough to make it clear that something protected was used in its training, is that enough proof that the legal protections were broken?
>First, you need to prove that generative AI works fundamentally the same way as humans at the task of learning. Next you have to prove that it recalls information in the same way as humans.
No, you don't need to prove any of those things. They're irrelevant. You'd need to prove that the AI is itself morally (or, depending on the nature of the dispute, legally) equivalent to a human and therefore deserving of (or entitled to) the same rights and protections as a human. Since it is pretty indisputably the case that software is not currently legally equivalent to a human, you're stuck with the moral argument that it ought to be, but I think we're very far from a point where that position is warranted or likely to see much support.
5 replies →
> that it’s anthropomorphizing machines.
No, it's not. It's merely pointing out the similarity between the process of training artists (by ingesting publicly available works) and ML models (which ingest publicly available works).
> First, you need to prove that generative AI works fundamentally the same way as humans at the task of learning.
Given that there is no comprehensive model for how humans actually learn things, that would be an unfeasible requirement.
7 replies →
We are machines. We just haven't evenly accepted it yet.
Our biology is mechanical, and lay people don't possess an intuition about this. Unless you've studied molecular biology and biochemistry, it's not something that you can easily grasp.
Our inventions are mechanical, too, and they're reaching increasing levels of sophistication. At some point we'll meet in the middle.
2 replies →
The way these ML models and humans operate are indeed quite different.
Humans work by abstracting concepts in what they see, even when looking at the work of others. Even individuals with photographic memories mentally abstract things like lighting, body kinetics, musculature, color theory, etc and produce new work based on those abstractions rather than directly copying original work (unless the artist is intentionally plagiarizing). As a result, all new works produced by humans will have a certain degree of originality to them, regardless of influences due to differences in perception, mental abstraction processes, and life experiences among other factors. Furthermore, humans can produce art without any external instruction or input… give a 5 year old that’s never been exposed to art and hasn’t been shown how to make art a box of crayons and it’s a matter of time before they start drawing.
ML models are closer to highly advanced collage makers that take known images and blend them together in a way that’s convincing at first glance, which is why it’s not uncommon to see elements lifted directly from training data in the images they produce. They do not abstract the same way and by definition cannot produce anything that’s not a blend of training data. Give them no data and they cannot produce anything.
It’s absolutely erroneous to compare them to humans, and I believe it will continue to be so until ML models evolve into something closer to AGI which can e.g. produce stylized work with nothing but photographic input that it’s gathered in a robot body and artistic experimentation.
5 replies →
The first perceptron was explicitly designed to be a trainable visual pattern encoder. Zero assumptions about potential feelings of the ghost in the machine need to be made to conclude the program is probably doing what humans studying art say they assume is happening in their head when you show both of them a series of previous artists' works. This argument is such a tired misdirection.
> What this means, is these systems will fall into different categories of law around copyright and free-use.
No they won't.
A human who uses a computer as a tool (under all the previous qualifications of fair use) is still a human doing something in fair use.
Adding a computer to the workflow of a human doesn't make fair use disappear.
A human can use photoshop, in fair use. They can use a camera. They can use all sorts of machines.
The fact that photoshop is not the same as a human brain is simply a completely unrelated non sequitur. Same applies to AI.
And all the legal protections that are offered to someone who uses a regular computer, to use photoshop in fair use, are also extended to someone who uses AI in fair use.
2 replies →
Why do you have to prove that? There is no replication (except in very rare cases), how someone draws a line should not be copyrightable.
16 replies →
What we actually need to prove is whether such technology is a net benefit to society all else is essentially hand waving. There is no natural right to poorly named intellectual property and even if there was such a matter would never be decided based on the outcome of a philosophical argument because we don't decide anything that way.
5 replies →
>My issue with this line of argument is that it’s anthropomorphizing machines. It’s fine to compare how humans do a task with how a machine does a task, but in the end they are very different from each other, organic vs hardware and software logic.
There's an entire branch of philosophy that calls these assumptions into question:
https://en.wikipedia.org/wiki/Posthumanism
https://en.wikipedia.org/wiki/Antihumanism
>Martin Heidegger viewed humanism as a metaphysical philosophy that ascribes to humanity a universal essence and privileges it above all other forms of existence. For Heidegger, humanism takes consciousness as the paradigm of philosophy, leading it to a subjectivism and idealism that must be avoided.
>Processes of technological and non-technological posthumanization both tend to result in a partial "de-anthropocentrization" of human society, as its circle of membership is expanded to include other types of entities and the position of human beings is decentered. A common theme of posthumanist study is the way in which processes of posthumanization challenge or blur simple binaries, such as those of "human versus non-human", "natural versus artificial", "alive versus non-alive", and "biological versus mechanical".
And? Even if neural networks learn the same way humans do, this is not an argument against taking measures against one's art being used as training data, since there are different implications if a human learns to paint the same way as another human vs. if an AI learns to paint the same way as a human. If the two were exactly indistinguishable in their effects no one would care about AIs, not even researchers.
But the 'different implications' only exist in the heads of said artists?
EDIT: removed a part.
13 replies →
And yet, some people don't even want their artwork studied in schools. Even if you argue that an AI is "human enough" the artists should still have the right to refuse their art being studies.
44 replies →
Is it strange to you that cars and pedestrians are both subject to different rules? They both utilise friction and gravity to travel along the ground. I'm curious if you see a difference between them, and if you could describe what it is.
Both cars and pedestrians can be videotaped in public, without asking for their explicit permission. That video can be manipulated by a computer to produce an artwork that is then put on public display. No compensation need be offered to anyone.
6 replies →
This is not one artist inspiring another. This is all artists providing their work for free to immensely capitalized corporations for the corporations sole profit.
People keep making metaphors as if the AI is an entity in this transaction: it’s not! The AI is only the mechanism by which corporations launder IP.
>This is all artists providing their work for free to immensely capitalized corporations for the corporations sole profit.
No, the artists would be within their rights to do that if they chose to. This is corporations taking all the work of all artists regardless of the terms under which it was provided.
Would it change your view if only open-source models were allowed to use the art in their training sets? What if a "capitalized corporation" starts using the open-source model?
This is such a nothing argument. Yes, new artists are inspired by other artists and sometimes make art similar to others, but a huge part of learning and doing art is to find a unique style.
But that’s not even the important part of the argument. A lot of artists work for commission, and are hired for their style. If an AI can be trained without explicit permission from their images, they lose work because a user can just prompt “in the style of”.
There’s no real great solution, outside of law, because the possibility of doing that is already here. But I’ve seen this argument so much and it’s just low effort
That is not how artists learn. This is a false equivalence used to justify the imitation and copying of artists’ work. Artists’ work isn’t derivative in the same way that AI work is. Artists create work based on other sources of inspiration, some of them almost or completely to the disregard of other art.
Many artists don’t even go to art school. And those that do, do not spend most (all) of that time learning how to copy or imitate other artists.
I’m not expressing an opinion of whether GenAI is unethical or illegal - I think that’s a really difficult issue to wrestle with - just that this argument is a post-hoc rationalisation made in ignorance of how good artists work (not to say ignorance of the difference between illustration and art, conceptual art training vs say a foundation course etc).
It's not the learning per se what's concerning here but the ease of production (e.g. generate thousands of images in a day)
AI is just a tool in someone's hand, there's a human who intends something
If that's true, then it should be fine for that human to paint with the brush of his AI tool. Why should that human artist be restricted in the types of tools he uses to create his artwork?
3 replies →
You might as well compare a Xerox copier to a human.
Art schools don't teach people how to paint in different artistic styles. They teach materials and technique.
Very true. I was watching a video yesterday learning how to make brush work digitally. While there were examples, they were just examples but the rest was specific techniques and demonstrations.
It is only natural to see a moral difference between people going to school and learn from your art because they are passionate about it, versus someone on the internet just scraping as many images as possible and automating the learning process.
Not being snarky, but if you believe that, then you clearly are not an artist.
his handle is KingOfCoders - self-aggrandizing, insufferable, impotent in its attempts to be meta.
He thinks he's an artist because he now has the ability to curate a dataset based off of one artist's work and prompt more art generated in that style. He did it, so clearly he is an artist now.
(salty salty~)
Human learning =/= machine learning
Most artists are happy to see more people getting into art and joining the community. More artists means the skills of this culture get passed down to the next generation.
Obviously a billion dollar corporation using their work to create an industrial tool designed to displace them is very different.
The memetic weapons humans unleashed on other humans at art school to deter copying are brutal. Just wait until critique.
"Sorry, this is not art, is AI generated trash."
3 replies →
This seems more like looking at other artists and being totally incapacitated by some little touch in the painting you're looking at.
The Nam-shub of Hockney?
Artists learning to innovate a trade defend their trade from incursion by bloodthirsty, no-value-adding vampiric middle men attempting to cut them out of the loop.
This is a tired argument; whether or not the diffusion models are "learning", they are a tool of capital to fuck over human artists, and should be resisted for that reason alone.
As a human artist I don't feel the same as you, and I somehow doubt that you care all that much about what we think anyways. You already made up your mind about the tech, so don't feel the need to protect us from "a tool of capital [sic]" to fortify your argument.
15 replies →
great comment!
imagine being a photographer that takes decades to perfect their craft. sure another student can study and mimic your style. but it's still different than some computer model "ingesting" vast amount of photos and vomiting something similar for $5.99 in aws cpu cost so that some prompt jockey can call themselves an AI artist and make money off of other peoples talent.
i get that this is cynical and does not encompass all ai art, but why not let computers develop their own style wihout ingesting human art? that's when it would actually be AI art
6 replies →
Exactly. Artists should drop the pretentious philosophical bumbling and accept what this is, a fight for their livelihood. Which is, in every sense, completely warranted and good.
Putting blame on the technology and trying to limit public access to software will not go anywhere. Your fight for regulation needs to be with publishers and producers, not with the teen trying to make a cool new wallpaper or the office-man trying to make an aesthetic powerpoint presentation.
> they are a tool of capital to fuck over human artists
So are the copyright and intellectual property laws that artists rely on. From my perspective, you are the capital and I am the one being fucked. So are you ready to abolish all that?
8 replies →
As a representative of a lot of things but hardly any capital who uses diffusion models to get something I would otherwise not pay a human artist for anyway, I testify that, the models are not exclusively what you describe them to be.
I do not support indiscriminate banning of anything and everything that can potentially be used to fuck someone over.
15 replies →
"Cameras are a tool of captial to fuck over human portrait artists"
It's funny that these people use the langauge of communism, but apparently see artwork as purley an economic activity.
5 replies →
Human beings and LLMs are essentially equivalent, and their processes of "learning" are essentially equivalent, yet human artists are not affected by tools like Nightshade. Odd.
As another posted out, modern models like BLIP or GPT4V aren't affected by this either.
Humans don't fall for optical illusions? News to me.
1 reply →
Lol, the scale is other-worldly different...
Sigh. Once again: I always love it when techbros say that AI learning and human learning are exactly the same, because reading one thing at a time at a biological pace and remembering takeaway ideas rather than verbatim passages is obviously exactly the same thing as processing millions of inputs at once and still being able to regurgitate sources so perfectly that verbatim copyrighted content can be spit out of an LLM that doesn't 'contain' its training material.
I'm just glad that so many more people have caught on to the bullshit than this time last year, or even six months ago.
I really don't even get the endgame. Art gets "democratized", so anyone who doesn't want their style copied stops putting stuff on the internet, and eventually all human art is trained, so the only new contributions are genAI. Maybe we could get a few centuries worth of stuff of "robot unicorn in the style of Artist X with a flair of Y" permutations, but even ignoring the centipede, that just sounds... boring. worthless.
Since techbros are stupid: "Note that people could always do these kinds of repurposing, and it was never a problem from a copyright perspective. We have a problem now because those things are being done (1) in an automated way (2) at a billionfold greater scale (3) by companies that have vastly more power in the market than artists, writers, publishers, etc. Incidentally, these three reasons are also why AI apologists are wrong when claiming that training image generators on art is just like artists taking inspiration from prior works."
https://www.aisnakeoil.com/p/generative-ais-end-run-around-c...
A human artist cannot look at and memorize 100000 pictures in a day, and cannot paint 100000 pictures in a day.
I am SO tired of this non-argument
A human artist does not need to look at and memorize 100000 pictures in any span of time, period. Current AI does.
We needed huge amounts of human labor to fund and build Versailles. I'm sure many died as a result. Now we have machines that save many of those lives and labor.
What's your non-argument?
2 replies →
These AIs are not people. They do not learn.
Define learn.
3 replies →
[flagged]
I'm not sold on your argument. I'm not an artist but I don't see how an artist using Nightshade is breaking the law. From an anti-AI point of view, you illegally took my artwork and used it. How is it my fault you didn't understand what you were stealing?
Currently it's not legally defined what is stealing and what is fair use with respect to this, which is of course why k feel it is such a strong issue, however, poisoning data and intentionally masking it to hide said poisoning is rather blatantly illegal.
Rather clearly I think most people support individual IP protection and that's not really a contested issue, however what is fair use and where things fall in that gray area is where things do get dicey.
4 replies →
Why is it my responsibility to ensure my data doesn’t inadvertently harm your model that I have zero knowledge of?
Duty of Care, additionally from an alternative perspective there are a number of laws covering the transmission of data that is intended to cause harm to another system.
(I understand that this is not a popular point, but I really want to emphasize that I am talking about what is _currently legal_ right now, not at all about the ethicality of large companies using people's data. The latter is a much harder topic.
This is mainly about what the law considers to be legal or not legal and is trying to avoid the more emotional side of the topic.)
Delighted to see it. Fuck AI art.
This is excellent. We need more tools like this, for text content as well. For software we need GPL 4 with ML restrictions (make your model open source or not at all). Potentially even DRM for text.