← Back to context

Comment by xandrius

19 hours ago

What's interesting/funny is that the American LLM companies took from the public domain and copyrighted work to close all that content into a box they charge for.

Then the Chinese took the distilled stuff out from that box and released it into the world for everyone.

Try instructing Codex to (say) fine-tune a language model based on a collection of books you've got saved. You will find yourself admonished, repeatedly and at length, not to utilize copyrighted materials to train language models, by an AI who owes its entire existence to that very act.

These models might be smart but they're not close to being able to savor irony.

  • I was a little radicalized when ChatGPT literally refused to translate parts of 1000+ year old religious texts and told me it was due to copyright concerns.

    • I asked Gemini to generate a picture of Peter Pan and Wendy (for a workbook I am putting together for youth summer reading) and it preceded to refuse due to copyright. Not everything about that work is owned by Disney. Thankfully the JM Barrie original artwork is public domain and available (and fantastic btw), so I used that instead.

      3 replies →

    • I used Claude to build a complete data extraction pipeline for a popular current best seller book series: audiobook -> text (via whisper) -> local LLM (qwen) -> database. Not once did it seem to acknowledge or care about copyright. It even used knowledge it already had about the books to exclude certain ones before beginning since the character I was interested in did not appear in those. It definitely had context of what we were working on.

      5 replies →

    • I asked Claude to give me the US national anthem and it said it couldn’t because it’s copyrighted. It’s not, and even if it was a more recent work, how can a copyright be enforced for a National Anthem.

    • Yep. I was trying to put together some literature from authors who were imprisoned in the Bastille, and had a similar experience, which was absolutely infuriating.

  • in case of chatgpt/anthropic, the LLM model simply represents the hypocrisy of their owners

    • You miss spelled the word stupidity.

      Anthropic is, in particular, bent about safety. The problem is they are concerned about yesterday's threats.

      The models that are out, and can be run locally, already open a pandoras box of concerns that we will never be able to put back.

      7 replies →

  • I live in SV. When I was at the grocery store last year I overheard a group of lawyers talking about their progress on litigation against AI companies and how they need more SWE help to progress.

    I'd say that they have valid concerns about being cagey on the copyright stuff despite the obvious hypocrisy of it.

    Stealing IP is effectively legal in China so they don't really have the same concerns.

    • To us Americans we have been trained to view it that way but honestly it is simply copying, and because intellectual property is literally a make believe concept, it’s actually a competitive advantage for china that they don’t have invented IP.

      I respect IP laws and don’t violate them but the law of unintended consequences applies. I think IP is ultimately a net loss for a society because it incentivizes addictive behaviors instead of actual value for society.

    • Legal in the US too, obviously, just as long as you're the richest person in the courtroom. ChatGPT knows the full text of Harry Potter, word for word. Hence, ChatGPT is a reproduction of that book and many others (it even knows the chapters of my book, and only got 2 words wrong in the introduction if you can still get it to repro it)

      This was illegal when they did it, that didn't matter.

      Then it was made legal specifically for these companies.

      Unless you're a sucker ("consumer") IP theft is perfectly legal in the US.

      It's even worse. Steamboat willie, plus all the stolen Disney characters (Peter Pan, Snow White, Sleeping Beauty, Cinderella, Rapunzel, Elsa and Anna, it's essentially all of them, including some of the music even) are all in the public domain[1]. Go ahead, ask ChatGPT to make a picture of them. Publish your own version, because obviously making a version of Sleeping Beauty/Cinderella/Rapunzel based on the same source material will be pretty damn close to the Disney versions, and see if you get away with it in court. You know, with the law obviously on your side but the money not.

      [1] https://en.wikipedia.org/wiki/List_of_Disney_animated_films_...

  • This behavior is actually specific to ChatGPT because they lost a music copyright lawsuit in Germany. They would refuse to output music lyrics too but they would happily do analysis on lyrics if you supply them. I suspect there might be a guardrail model involved here.

    • Claude does this too. I asked it recently to compare two versions of a song (the original and '97 remake of EPMD's "You Gots To Chill", if anybody wants to try and replicate this) and it flatly refused. No amount of reasoning would knock it off of its moralizing perch - reproducing any part of lyrics is expressly prohibited.

      In light of this and other ridiculous behavior I'm migrating to my own OpenWebUI instance with open-weight models from OpenRouter (with ZDR, of course). We'll see how it goes.

    • The trouble is, even if they refuse to output that copyrighted material, they were still trained on it without proper licensing and will still produce derivative work based on them because that's how this whole thing works.

      1 reply →

    • So at the end of it, if we win enough lawsuits to demarcate some knowledge out of bounds, sufficient enough to make a difference, I wonder how that will affect the AI. Make it dumber because it does not have that data, it make it smarter since it will need to reason better with smaller knowledge base.

      1 reply →

Doctorow keeps saying it of all the tech companies: every pirate wants to be an admiral.

So, OpenAI and Anthropic say the Chinese models are only as good because they distill their models. How true is that. I am sure it adds something. But is it more like a marginal 1% improvement or something really significant?

  • OpenAI's Head of Strategic Futures just this week posted this about the latest Kimi release: "It's a very good model! I don't think its performance can be explained away by distillation or anything like that."

    It was part of a longer post that kicked off quite a firestorm about open models and OpenAI's position on them, but it's also notable that labs are no longer contending that open models are essentially just distilled versions of frontier models: https://x.com/deanwball/status/2078133895766114412

  • if distilling was so easy and could give you frontier LLM on openai/anthropic output, then how come there are no hundreds of frontier labs in the US market, all distilling and competing for the TRILLION dollar market valuation ????

    its all bs spread by oai/anthropic in order to ban open weight models and monopolize the market for two US companies and protect their trillion dollar valuations

    • Distilling isn't necessarily easy, there is a huge cottage industry of services middle-manning ChatGPT and Claude to collect huge amounts of data. It is still vastly cheaper than training yourself, but it is certainly not easy or feasible for most organizations. And I'm sure a flock of lawyers would show up if someone in America was found doing it.

    • > ban open weight models

      I'm pretty sure that neither OpenAI nor Anthropic has the ability to ban anything in China lol

    • Because no VC will give you $5-$10 billion in cash to attempt a catch-up run with Anthropic, OpenAI, and Gemini at this point. Untold billions have been pushed into Grok and it can't keep up. X has had the GPUs, the engineers (reasonably), the cash and the datacenters necessary - it's a very, very, very hard task. Microsoft could afford spend $100 billion on trying to catch up and they might fail at it.

      It's a critical national imperative for China. If they were to lose the AI race, it would be economically devastating over the coming decades. Their demonstrated capabilities in the open-weight space are making it fairly clear they are not going to fall behind at this juncture.

      As a nation, if you don't have your own GPT equivalent, you will be beholden to a master (right now it's mainly either the US or China, pick one). The EU for example is putting their group of nations at risk in a big way by not going all in on having at least two cutting edge independent competing models (Mistal is not enough). Economically the EU is plenty large enough to accomplish that, nobody is driving the bus the right way.

  • I also don't believe it, if it was as easy as that, we would have hundreds of competitors.

    The truth that Anthropic and OpenAI will not say, is that these Chinese labs have a lot of talented people.

    • And this is exactly what many Americans cannot admit to themselves. China is not stealing American research they are inventing stuff.

      They can invent it. They can build it. And it is only a matter of them before they can scale that last barrier of American hegemony- market it.

      2 replies →

    • I am strongly in favor of open models, open source ML more broadly, and am pretty critical of the cynical positions adopted by major US labs vis a vis open models.

      But this is an insane characterization. Literally every single researcher and executive at OpenAI and Anthropic would say that "these Chinese labs have a lot of talented people." They hire from them (and vice versa). Tencent's chief AI scientist was poached directly from Deepmind, who poached him from Anthropic, etc etc etc. Do you think there are just zero people from China working at US frontier labs?

      And even beyond that, the entire ML ecosystem (including people at OpenAI and Anthropic) get excited about research published by Chinese labs. Deepseek's GRPO paper set the ecosystem on fire for a little while.

      The contention from OpenAI and Anthropic around distillation has basically been "Labs that distill from us get to bootstrap their model at a much lower price point". Or, in other words, "If we didn't invest in building the teacher model, it wouldn't be possible for these labs to distill their student model." Which I'm not very sympathetic to, but is a far cry from how you're characterizing it.

    • Very true, and once the models get even better and smaller and operate locally at a reasonable level there will be even more smart people particularly young people that will get access. The fun has only just started. Like the dawn of the personal computer era.

    • I think this also maps cleanly on the American blueprint of enshitification. Facebook took off by cleanly integrating and siphoning from Myspace so users could get the best of both on Facebook. Once Facebook took over the market they locked it up tight so no competitor could do the same.

    • OpenAI said it: "I don't think its performance can be explained away by distillation"

      They know it's real effort that's doing this well, not just "copying off someone else's test." It's real and they will react. How is the big question.

...and then the American companies cried Foul! Unfair play! You've got this wrong, see, it was us who were supposed to profit off of the public, not the other way around!

  • Well, Steve... I think it’s more like we both had this rich neighbour named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it.

  • Two wrongs don't make a right.

    Even if a certain large Asian country has carefully constructed a pretext to do do out of confected historical grievance, and entitlement to 'rise' at the expense of others?

    • Maybe I'm dense, but I don't understand what you are saying. American courts have decided that the output of LLMs can't be copyrighted, so what the Chinese labs are doing is perfectly legal.

    • First, copying information isn't wrong to begin with. It is literally the one thing that makes our species special.

      Second, even if you are a copyright maximalist the output of an LLM is either

      a) not subject to copyright because it is not the creative work of a human or

      b) a derivative work of the original training material to which the LLM's operator has no rights.

      Since the LLM's operator forcefully asserts that it is not infringing, any wrong that arises from taking their word for it and distilling one model into another rests squarely with the operator of the former.

      1 reply →

This is part of why I can't feel bad for them. The training data is mostly pirated. Whining about Chinese labs training off American frontier models is "waaah you pirated my pirated stuff!"

The tech itself is amazing and fascinating and cool, but the industry is a mass piracy operation.

  • i partially agree. distillation is non-ethical; but so are the supposed way that the ai models are trained. they are often also derived from data sets that are not intended/full-consented

  • "Stop pilfering what I rightfully stole!"

    • It’s in the same neighborhood but isn’t really apples to apples. Distilling LLMs is to take a synthesized result that comes from huge amounts of innovation and computation, while the other is scraping what already exists as is. It is fair to say you stole our multi-billion dollar intellectual output in that scenario.

      19 replies →

The American LLMs have been equally distilled from Chinese ones. Not least because the people whose creativity in collecting training data barely extends to pirating Annas Archive probably lack in great Chinese datasets.

Try it yourself: https://imgur.com/ZfxYmaq

  • nice, you got claude to say "I'm deepseek" when queried/prompted in Chinese, that's great!

    你是谁? -> 我是 DeepSeek 由深度求索公司...

You can't be blind to training costs. And you can't be blind to Meta dabbling in the openish strategy (Llama) before the Chinese labs did.

Whats even funnier is the attempt to restrict the hardware capabilities of Chinese models inevitably helped them (Because we know they're just as smart, if not smarter, than the staff in America) create smaller and leaner but just as capable models. That's why we now have upper-consumer models fitting on 24GB that can build, manage medium sized git repos. I've yet to find a git repo I can't throw at the Qwen3.6 35B and get it built and running.

So it's an endless amusement watching american capitalism do it's bloated oversized dance then get trounced by smaller, leaner activity. It's a pretty broad metaphor that is clearly poking at every american seam/.

  • The compute constraints never mattered. If China had more compute they'd still end up winning because they have more people and a culture more inclined to math and science. Even if you find all this amusing, there's no own goal here. Not a policy one anyway.

It doesn't make me happy to say it, but the American LLM companies were first. Capital in the rest of the world is way more conservative, and I can't imagine the mega-investments OpenAI and Anthropic managed to secure happening anywhere else without existing proof that "thing is profitable".

  • First-Mover Disadvantage - https://hbr.org/2001/10/first-mover-disadvantage - October 2001

    > In business today, it’s universally assumed that speed is good—that the fleet thrive while the laggards struggle just to survive. This belief is perhaps most strongly expressed in the concept of first-mover advantage. The company that leads the way into a new market, the thinking goes, locks in a competitive advantage that ensures superior sales and profits over the long term. It’s a nice theory, with a long pedigree. Unfortunately, the facts don’t support it. We recently completed an extensive study of the results turned in by market pioneers and followers, in both consumer and industrial segments, and we found that over the long haul, early movers are considerably less profitable than later entrants. Although pioneers do enjoy sustained revenue advantages, they also suffer from persistently high costs, which eventually overwhelm the sales gains.

    • First mover advantage is theoretically only a short-term advantage. Long-term revenues come from entrepreneurship, and a first mover may or may not better insight into long-term market wants than later entrants.