Comment by jacquesm
11 days ago
Because he planned to make the data public, Meta just wants to use it to enrich its shareholders. See also: Google, OpenAI, Anthropic and every other big player in this space besides.
11 days ago
Because he planned to make the data public, Meta just wants to use it to enrich its shareholders. See also: Google, OpenAI, Anthropic and every other big player in this space besides.
There’s also the fact that Swartz was physically trespassing and attaching unauthorized machines into networking closets to run scraping on a university network to exfiltrate the scrapes to the public, versus just scraping public facing web from the public web to train a model. That’s a little different and while the feds were heavy-handed against Swartz these computer crime laws were well known and it was less heavy handed than the hacker crackdowns of the 90s if you want to look at precedents.
Meta has opened sourced almost all of their models trained on this data.
But not the training data.
Not sure what your point is - the articles they trained on are copyrighted and not theirs to open source. The only they can legally open source is the weights, which they have done repeatedly.
My point is both Aaron was trying to make the journals public and Meta made models that encapsulate data from the journals public.
With added filters (censorship)