It’s in the same neighborhood but isn’t really apples to apples. Distilling LLMs is to take a synthesized result that comes from huge amounts of innovation and computation, while the other is scraping what already exists as is. It is fair to say you stole our multi-billion dollar intellectual output in that scenario.
> It’s in the same neighborhood but isn’t really apples to apples. Distilling LLMs is to take a synthesized result that comes from huge amounts of innovation and computation, while the other is scraping what already exists as is.
Hang on, why is scraping the public pool of knowledge not taking "a synthesized result that comes from huge amounts of innovation and computation"?
You think that that all those github repos that LLMs trained on, were not the result of innovation and computation?
How many years of human innovation and cycles of computation during compilation were involved in bringing something like GCC or LLVM to their current status?
Those LLMs trained on every single research paper available online - were those papers not the synthesised result of billions of dollars of research, effort and (importantly, for you anyway) computation?
LLMs trained on the collected works of every author in existence. Were all those works just "as is"?
> It is fair to say you stole our multi-billion dollar intellectual output in that scenario.
Thousands of years of human innovation taken without any permission.
Everyone should steal everything not nailed from other AI companies. Then steal everything nailed and take the nails too. At least this way a tiniest bit might return back to society.
a synthesized result that comes from huge amounts of innovation and computation
The published algorithms like the transformer architecture are not patentable. You spent a lot of money on compute and China used the uncopyright-able output to steer its own training models? Too bad. I feel especially unsympathetic to OpenAI, who went from being a presenting itself as a benevolent nonprofit to a very-much-for-private private entity over night.
Don’t forget that the Chinese models are also built on top of huge amounts of “stolen data” as well, beyond the distilled. So it’s basically all of the above. However, there’s no mechanism for the NYT or an author or anyone in the US to sue the Chinese companies that took their work.
Legitimate Salvage !
It’s in the same neighborhood but isn’t really apples to apples. Distilling LLMs is to take a synthesized result that comes from huge amounts of innovation and computation, while the other is scraping what already exists as is. It is fair to say you stole our multi-billion dollar intellectual output in that scenario.
> It’s in the same neighborhood but isn’t really apples to apples. Distilling LLMs is to take a synthesized result that comes from huge amounts of innovation and computation, while the other is scraping what already exists as is.
Hang on, why is scraping the public pool of knowledge not taking "a synthesized result that comes from huge amounts of innovation and computation"?
You think that that all those github repos that LLMs trained on, were not the result of innovation and computation?
How many years of human innovation and cycles of computation during compilation were involved in bringing something like GCC or LLVM to their current status?
Those LLMs trained on every single research paper available online - were those papers not the synthesised result of billions of dollars of research, effort and (importantly, for you anyway) computation?
LLMs trained on the collected works of every author in existence. Were all those works just "as is"?
> It is fair to say you stole our multi-billion dollar intellectual output in that scenario.
No, we didn't. We simply took the model as-is.
The Chinese models are the result of just as many papers, GitHub repos, etc… AND the synthesized results of those.
5 replies →
> comes from huge amounts of innovation
Thousands of years of human innovation taken without any permission.
Everyone should steal everything not nailed from other AI companies. Then steal everything nailed and take the nails too. At least this way a tiniest bit might return back to society.
a synthesized result that comes from huge amounts of innovation and computation
The published algorithms like the transformer architecture are not patentable. You spent a lot of money on compute and China used the uncopyright-able output to steer its own training models? Too bad. I feel especially unsympathetic to OpenAI, who went from being a presenting itself as a benevolent nonprofit to a very-much-for-private private entity over night.
You could say that both of them stole, but different stuff.
Don’t forget that the Chinese models are also built on top of huge amounts of “stolen data” as well, beyond the distilled. So it’s basically all of the above. However, there’s no mechanism for the NYT or an author or anyone in the US to sue the Chinese companies that took their work.
1 reply →
You can't steal intellectual property, only infringe on the copyright holder.
2 replies →
Sorry, no leg to stand on and relatively speaking no sympathy for the model makers or the data-center builders…
I assume you don’t use any LLMs then
1 reply →
bet