Comment by zqwt3k
8 hours ago
Thanks Elon, now we know that scraping is illegal! Very good to clarify that for future proceedings against the AI thieves.
8 hours ago
Thanks Elon, now we know that scraping is illegal! Very good to clarify that for future proceedings against the AI thieves.
Everything is both legal and illegal until a lawsuit happens. Then it collapses depending a little bit on the facts and mostly on who has the better lawyers.
I suspect Nitter's first round with lawyers pointed out that scraping is legal, but now they have been threatened with something else than scraping - Elon claims something else the way Nitter runs is illegal, such as the use of fake accounts to circumvent an access control device (DMCA 1201).
At the end of the day, as an individual or a team or a company, regardless of the statue and case law, you have to perform the calculus on your monetary and legal resources versus your counterparty.
Obviously ungrounded and frivolous cases tend to be easier to defend asymmetrically, but if I was X's legal team, there's no shortage of semi-plauisble claims I could throw at the wall and see what sticks.
As an example of this imbalance in action, BrightData is a 'gray area' company that basically does this exact kind of scraping. They have somehow won against Meta Platforms suing them, and even got X's lawsuit against them for scraping -- identical (?) activity to XCancel -- dismissed.
According to Wikipedia:
> In May 2024, a federal judge dismissed the suit, ruling that Bright Data did not violate X's terms of service or copyright by scraping publicly accessible data.[24] The judge emphasized that such scraping practices are generally legal and that restricting them could lead to information monopolies
But does XCancel have the resources of a company like Bright Data, that's funded and used by companies like Deloitte and Moodys?
If the name Bright Data is ringing a bell to anyone, it’s probably because they are a (the?) primary offender running the LG TV “residential proxy” (botnet)
2 replies →
Notice it says "terms of service or copyright". If X's lawyers have any intelligence, they'll have a reason why XCancel is not identical to Bright Data. Perhaps this time, instead of claiming it's a copyright violation, they'll claim it's wire fraud because multiple accounts are used.
I wonder if anti-SLAPP laws could be used to shield nitter: https://en.wikipedia.org/wiki/Strategic_lawsuit_against_publ...
Scraping is legal, but IIRC scraping authenticated contents is not.
Isn't Nitter abusing account sign-in for this?
Anyone can make an account, though, and instantly access that content, so Elon doesn't really have any leg to stand on by claiming they're private. It's not the same as Cambridge Analytica scraping stuff you had to have certain privileges to see by tricking the system into granting those privileges.
5 replies →
To fix this:
1) Nitter offers a deal: download our browser extension, sign up for Twitter, and we'll give you some kind of perk (Amazon credits, whatever).
2) Browser extension surfs Twitter on the user's behalf, scraping and sending copies to Nitter.
3) We find out how serious the legal system is about prosecuting scraping.
3 replies →
Yes, that's what they went on to point out.
Aurora Store does the same thing to facilitate downloading Play Store apps.
I like this framing, like Schrödinger's cat.
Thin skins over at there at X
I wonder what legal gymnastics are needed for "I can scrape anything off the web ignoring copyright and build a product from this, but you can’t even display what’s on my webpage elsewhere”.
Perhaps that is it in fact. The act of protecting it from scraping means you object. 99% of the blogged contents etc. Big AI helped themselves to was just… there. Public. Not free from copyright but still not paywalled.
Precedent is pretty clear: competitive uses bad, transformative uses good. Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content.
> Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content.
It seems unreasonable to stop there though; the agentic bots are designed and marketed as able to compete with the initially-scraped sources.
I'm not convinced that a competitive use at one remove should be treated as not competitive.
3 replies →
I think LLMs providers pretty directly compete with content they scrape like Wikipedia and SO...
24 replies →
> Precedent is pretty clear:
What cases are you citing when you say this?
7 replies →
but xcancel is not putting ads or making a commercial product
How clear is it when I google a recipe and get an AI-generated recipe that's clearly derived from the top three results and then placed above those results? That sounds like it's both transformative (in that the recipe created by the AI may not match any one of the scraped recipes perfectly) and also competitive (in that the AI takes page views away from the pages it got the recipes from)
2 replies →
Transformation.
Taking something someone else made and showing it as-is, bypassing their own restrictions: No no.
Taking something someone else made, modify it or use parts of it in some bigger thing or completely change it: Fine, if you have money and/or run a company
twitter didnt make it though. they have a license to it
So in theory if you took twitter content and then transformed it so it "summarizes" all tweets with an AI rather than posting the exact text, would that be allowed?
Because that's stupid. These laws are stupid.
7 replies →
except a bunch of paywalled stuff did end up in training corpora
Pretty sure it's legal when you do it for your own use (same as browsing a website) but it's illegal to redistribute web scraped results.
It's not that cut and dry or else search engines wouldn't be legal. It depends on how much is used, for what context, etc. This very well may wind up being fair use.
Search engines "modify" it ie. show snippets + direct to the actual site.
In general fair use pretty much always requires it to be transformative and/or point to the source. Simply scraping it to prevent people from going to X isn't free use in any definition I've heard.
1 reply →
Thats "Billionaire Use" -- its like a "Fair Use" exemption but for billionaires.
Aaron Swartz died so that "AI" billionaires can live.
[flagged]
1 reply →
Now I want to know, what happens if you redistribute an "AI summary" of the copyright material?
scraping content is mostly legal, redistributing content is not.
if you started doing the same to, say, instagram content both meta and individual creators would sue you as well.
sites like archive.ph are in a similar bucket btw, and yet nobody's complaining (except websites seeing people evading their paywall). but at the end of the day it's not really fair to apply laws differentially on the basis of whose political ideas we like more.
Meta has no exclusive rights to the content on Instagram, and X has no exclusive rights to the content on X. They have a non-exclusive license to republish it, etc.
It’s the rehosting of the content not the scraping
Edit: not a moral stance
I love how the content belongs to them when someone else reposts it but it belongs to the user if the content is illegal. Such a double standard with these social media and AI companies. Why do we put up with it?
> Why do we put up with it?
You let it happen. Once people stop letting it happen, it'll stop. But social media is apparently the new "opium of the masses" so here we are and no one wants to do anything.
2 replies →
Did the users of X consent to XCancel copying their posts to their servers?
27 replies →
That is what all LLMs could do in 2023, verbatim, before they trained it out of them in order to keep up the pretense that there is no plagiarism. Now they all obfuscate the original or refuse to cite.
There's a button on X that allows me to repost someone else's content.
there's a setting to disable that
That’s not rehosting
Edit: iPhone autocorrected my OP which meant to say rehosting not reposting
Grok can do that if you give it an HN thread or other websites.