Comment by rich_sasha
10 hours ago
I wonder what legal gymnastics are needed for "I can scrape anything off the web ignoring copyright and build a product from this, but you can’t even display what’s on my webpage elsewhere”.
Perhaps that is it in fact. The act of protecting it from scraping means you object. 99% of the blogged contents etc. Big AI helped themselves to was just… there. Public. Not free from copyright but still not paywalled.
Precedent is pretty clear: competitive uses bad, transformative uses good. Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content.
> Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content.
It seems unreasonable to stop there though; the agentic bots are designed and marketed as able to compete with the initially-scraped sources.
I'm not convinced that a competitive use at one remove should be treated as not competitive.
I think that's more true in image generation than in text? At least, all the money is in LLMs that write code, not LLMs that write O'Reilley-style coding books.
2 replies →
I think LLMs providers pretty directly compete with content they scrape like Wikipedia and SO...
Remember the golden rule of the golden rules:
Who has the gold makes the rules.
7 replies →
No kidding. I dunno what kind of traffic loss Wikipedia has had but SO is really dead these days
If they weren't competing with AI then why is AI killing it?
15 replies →
> Precedent is pretty clear:
What cases are you citing when you say this?
Bartz v Anthropic. Though the plaintiffs did get something, it was because of the piracy to the original works (competing against the legal market for the books), not the use of them to train the LLM.
5 replies →
Perhaps https://en.wikipedia.org/wiki/Warner_Bros._Entertainment_Inc....
but xcancel is not putting ads or making a commercial product
How clear is it when I google a recipe and get an AI-generated recipe that's clearly derived from the top three results and then placed above those results? That sounds like it's both transformative (in that the recipe created by the AI may not match any one of the scraped recipes perfectly) and also competitive (in that the AI takes page views away from the pages it got the recipes from)
> How clear is it when I google a recipe and get an AI-generated recipe
Recipes can’t be copyrighted
Here’s one discussion about this https://www.nycbar.org/reports/secret-ingredients-how-to-pro...
1 reply →
Transformation.
Taking something someone else made and showing it as-is, bypassing their own restrictions: No no.
Taking something someone else made, modify it or use parts of it in some bigger thing or completely change it: Fine, if you have money and/or run a company
twitter didnt make it though. they have a license to it
So in theory if you took twitter content and then transformed it so it "summarizes" all tweets with an AI rather than posting the exact text, would that be allowed?
Because that's stupid. These laws are stupid.
That's exactly the way UK courts are heading, see Getty vs Stability AI. The court ruled that there's no infringment because the model doesn't store exact copies, just derived weights, and therefore when it generates new images those aren't copies of protected works.
7 replies →
The point is that you can't steal someone else's content 1:1. But you can use it for a different use (say, display the tweet in an article, then comment on it).
except a bunch of paywalled stuff did end up in training corpora