Comment by himata4113

1 month ago

Does this matter? Distillation is not illegal by every definition of the word.

There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them.

Another example is that it appears that the upper limit of what you can do is ultimately dependent on people working on the model, otherwise grok would be a LOT more competitive pre-cursor acquisition.

And lastly, kimi architecture is vastly different than that of fable as it uses mechanisms developed by... kimi themselves. US AI labs are inspired by opensource advancements just as much as open source labs are inspired by traces from models such as fable.

Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.

edit: (moved this to bottom) The only argument they have here is that they use GB300 GPU's which for some reason should not be available to chinese citizens.

Perhaps even more importantly, the current frontier LLM models are self-admittedly the product of enormous quantities of copyright infringement and even less savory inputs, so calling them out for distilling the fruit of that tainted tree reads as highly hypocritical at best.

  • fair use, both in the original training and in distillation, or rather, anthropic has no copyright at all over the output tokens

  • No - distillation is not data inputs.

    Raw materials vs. Value add.

    They are different things, like ore and metal.

    Distillation is a new thing we need to understand, it's probably closer to IP than not.

    • The "data inputs" were also, very much, somebody's "value added" IP.

      We're talking about things like text people wrote, not some kind of raw data floating out in the ether.

      1 reply →

    • Are you suggesting data input is further from IP than distillation?

      That would stun me, but it's a little hard to read.

    • LLM outputs are not copyrightable or rather the user who generated them owns it.

      That entirely settles it and there isn’t much else to say about.

      If Anthropic feels that other countries are violating their EULA well they are free to stop doing business with them.

      1 reply →

    • Do you think writing books (and Wikipedia articles, and stack overflow articles, and github repos, and, and, and, and ...) is not a value add?? What terrible claim.

  • While I agree on a moral level, I think there is a distinction to be made. Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way. I think this is much less true for distillation (which is kind of the whole point).

    ed: to clarify, I totally agree that a huge chunk of the value in LLMs is coming from the source material. My point was just that training an LLM takes more resources and expertise than distilling from an existing LLM so I don't think the equivalence between training and distilling is entirely justified.

    • I like this comment because its argument only makes sense if you assume that the entire world's output of books and art did not require a huge amount of resources and expertise to make, nor did it add any value.

      It's the most CS-major take ever!

      47 replies →

    • > raining a SOTA model takes a huge amount of resources and expertise

      Writing books, building Wikipedia, and answering questions on online forums takes a lot of resources and expertise that scraping didn't. So at the very least, we're already one rung down the "maybe you should've asked" ladder.

    • I suspect that, in aggregate, all of the informational output of humanity prior to 2020 has taken more resources to produce than the last few years of LLM research.

    • Why is it less true for distillation? Everyone technically has access to Fable but Moonshot came up with the model. How can you objectively claim one is adding value while the other is not?

      If that is the whole point you need to clarify why this is the case on an objective level.

      I would say building a comparable model using any means necessary (just like what Anthropic and OAI did) at a lower cost is actually more valuable to soceity and Monshoot is arguably generating more value with less.

    • Probably not as much effort as writing books and creating art the models were trained on.

    • The value of LLM's come from replacing what generated its training data.

      If the distilled model is cheaper, then it's just LLM's getting LLM'ed.

    • Still, AFAIK Kimi's architecture (just like that of other LLMs from Chinese labs) is different from those of OpenAI and Anthropic's model in a nontrivial way. So the expertise is still there, and I guess resource use too (although Chinese labs tend to optimize this, thanks to the restrictions they have on GPU use).

      EDIT: just wanted to add that resource optimization is usually where the contribution of Chinese labs is, so you shouldn't reaad the above parenthesis as a negative comment.

    • As an author, that's a genuinely disheartening thing to read.

      It took me a year to write a book. It took OpenAI and Anthropic a fraction of a second to ingest it. Do you understand now why I give zero shits if it takes Anthropic a billion to train a model, and Moonshot 10k in API cost to distill it?

    • > training an LLM takes more resources and expertise than distilling from an existing LLM

      This is not automatically true. Training and distillation use the same underlying infra and method and there is no intrinsic differences in between.

    • They add value on top of other people’s work, often against licensing, and then commercialize this product, ie profiting from making a product out of other people’s IP.

    • >Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way.

      producing the entire body of human knowledge that Silicon Valley companies absorbed like the Borg did not just take more resources but also a fair amount of blood and sweat, certainly more than the LLM so on that front that comparison also seems entirely justified.

    • I'm sure it takes a lot of time and resources to plan and pull off an epic heist but it is unusual to see people like Thomas Crown being accused of creating value, as they're usually accused of committing theft.

    • Yeah, there's a difference. One party spends a bunch of resources doing something illegal and extremely immoral. The other party spends little money doing something legal and morally neutral.

    • You can argue that reverse engineering anything is as hard if not harder than engineering something. I can’t imagine distillation is any different.

      3 replies →

  • I disagree that LLM models are the product of enormous quantities of copyright infringement.

    The recent announcement that AI-assisted research produced a counterexample to the Jacobian conjecture--a long-standing open problem in algebraic geometry--shows the original value AI can create. The result was not copied from a textbook; it emerged from AI learning from existing material, much as a human does, and then applying that knowledge in a new way. If that's a violation of copyright, then a human doing the exact same thing would be a copyright violation too. But it isn't.

It matters because everyone imagines the inevitable "closing of the gap" between closed and open source, but the rate at which open source catches up with closed source seems to depend on being able to train on and distill the outputs of open source models. As long as performance of open source models is at least partially dependent on frontier-model outputs, then that gap will remain in place by definition.

>Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.

If the distillation is irrelevant to why it is competitive, why do they do it then? Obviously is helps improve their benchmarks/performance to some degree, otherwise they wouldn't need to do it.

  • Never claimed that it is irrelevant. And kimi k3 is on the same level and sometimes outperforms fable 5 - that cannot be explained by distillation. The reason why gap is not closed is simply the fact that fable was trained months ago so in theory the frontier labs are still 1 (small) step ahead.

    Although I will reiterate the fact that distillation is not the primary reason why these models are performing so competitively.

    • Closed-source models have to deal with the current frontier being heavily regulated. Fable, at its old level, was "too good" to be released and they had to add an additional safety layer to sanitize the outputs. Lowering the quality of the models so they are safer and more steerable has been something all the closed-source models have been doing for a while, a requirement that many open source models don't need to deal with.

      If Kimi k3 really were above Fable 5 then there invariably the USG would have to consider their restrictions on model capabilities excessive, or one would have to admin closed source models are held to more restrictive safety standards than open source models.

      >Although I will reiterate the fact that distillation is not the primary reason why these models are performing so competitively.

      How would you know this? How could you ascertain exactly how much performance is attributable to their unique engineering/research? If they really were so competitive they could surely make a model that isn't dependent on distilling Fable or other frontier models.

      2 replies →

I agree distillation isn't illegal; I also think Moonshot/Kimi is very impressive. But the more interesting question is whether labs like Moonshot can be a real competitor to OpenAI/Anthropic. If you can only play catchup (however quickly you do that), then you're never going to be at the frontier - I think that's why distillation matters.

  • People that think the Chinese are only able to copy western tech are in for a wakeup call.

    Actually, that has already happened in many domains, it's just that most western people (USA especially) won't admit it.

    • my assertion isn't that china isn't able to surpass western AI. I think it may well happen. I've been to china many times and am well aware of how ahead they are in many technological/societal areas.

      at the same time, I don't buy the idea that distillation is unimportant in assessing what Chinese labs are capable of. If it wasn't, why did Kimi's release timing coincide so well with Fable's launch?

      and if Anthropic hadn't released Fable, would we have Kimi today? If the answer is no, then I think that's still a very important point to consider.

      2 replies →

  • So many of the breakthroughs and architecture that make LLMs powerful in general today came from China, especially ones related to sparsity and MoE that have made inference and training substantially cheaper.

    Let's not forget how much people talked about "prompt engineering" before Deepseek mainstreamed the idea of thinking mode which is now universal

  • I think it depends on where you think we are on the S curve of intelligence growth. (Yes, I think it's an S curve, not an unbounded exponential). If you think we're near the peak than playing catch up (especially if you can play catch up quickly) is very rational.

    I know this isn't exactly a scientific test, but I had a local Qwen 3.6 27B model implement a fairly sizable feature today. There were a couple of bugs, mostly around me not giving sufficient specifications, but they were ironed out quickly when I pointed it out. I was able to ask the model to create instructions so next time it doesn't fall into the same pitfalls, and it did a great job. 27B local model! (And it was super fast too).

    I ran Fable 5 as a code review and it didn't really have any significant corrections.

    I guess my point here is that, for most work the frontier models are probably overkill anyway, and improving on overkill in a way that raises prices significantly is probably not a winning strategy.

    The only place I can think of where the super high powered models are "required" is if you want to do a ridiculous token burn like GasTown where you just have it run un-monitored on very long tasks. To me though, that's an experiment, not a real workflow. And the way these labs are like "oh we made this (broken) thing in a week using just agents!" always also follows with "and it cost $100,000+ in tokens!". Like, ok, I get it if you're doing research but that's the salary of an entire person.. that can actually learn and improve.

  • My entire point was that this was not achieved purely from distillation and claiming that is slander against open research.

  • > But the more interesting question is whether labs like Moonshot can be a real competitor to OpenAI/Anthropic.

    The answer depends on whether you think the AI researchers at Chinese labs are (or can be) as smart, motivated, and as good at math as those working at US labs - a not-insignificant proportion of whom are Chinese nationals.

> The only argument they have here is that they use GB300 GPU's which for some reason should not be available to chinese citizens

Note that Chinese companies are free to rent from GB300 clouds internationally. There are large datacenter hubs in Singapore and Malaysia serving chinese and other customers.

Though there is also reported [1] significant smuggling of Nvidia chips into China as well.

[1] https://epoch.ai/publications/chip-smuggling

Legal, illegal…

The word I would use is inevitable. It reminds me of the (PC) clones wars…

It is incredibly important to whether the US can maintain its AI lead. If foreign competition is closing the gap only by distillation, then the frontier labs can focus on preventing distillation and maintain their lead that way.

US dominance is also important for approaches to safety, especially political approaches. If the frontier models are all US-based, safety might be tackled via internal US policy. If other countries can independently train competitive models, international cooperation is required.

Edit: It is also important for the business model. Companies won't be able to justify tremendous training costs if competitors can replicate their product much more cheaply via distillation.

  • Do Americans even believe that US policy is likely to steer development in a way that’s safe and beneficial for humanity? The rest of the world certainly doesn’t. The US currently seems to primarily use their superpower status to be the world’s number one shit disturber and geopolitical antagonist.

    I don’t think China’s necessarily any better, but I’d rather have the most powerful models be open rather than under the exclusive control of the US executive.

    • China uses its power to make favorable deals and get people hooked on what it's slinging so it has captive customers. The US uses its power to bully and break rules that apply to everyone else for its own benefit. Kind of a big difference.

      2 replies →

    • I didn't mean to imply that the US is more likely than elsewhere to responsibly steer AI via policy. But I do think it is easier if it can be done internally as opposed to via international dealmaking.

    • >Do Americans even believe that US policy is likely to steer development in a way that’s safe and beneficial for humanity?

      No, no we do not.

    • > Do Americans even believe that US policy is likely to steer development in a way that’s safe and beneficial for humanity?

      No, at least not outside of this forum.

      We all mostly think these models and US policy are going to drive the exact opposite of that. Wealth will continue to get extracted and funneled to the top, and the rest of us are going to be left with the scraps and left to die while what little social safety nets we had continue to get eroded away alongside losing our jobs.

  • > If foreign competition is closing the gap only by distillation, then the frontier labs can focus on preventing distillation and maintain their lead that way.

    We already know it's false because you would have hundreds of competitors if it was that easy.

    The reason why these Chinese labs are releasing good models is simpler, they have access to a tremendous pool of talented people.

Of course it matters. Regardless of whether distillation is legal, there is a difference between training a model with and without distillation. For one thing, the distilled model wouldn't exist without the model it distilled.

Also, companies that use distillation may be competitive but seem unlikely to surpass the companies that are training these models from scratch.

  • >but seem unlikely to surpass the companies that are training these models from scratch

    Then why is it a problem?

    Another serious question.

    Trying to get my head around what the root of the objection is here. There must be some fear, but if that fear is not a fear of being surpassed in the market, then what is the fear?

    • A fear of competition causing a failure to recoup the trillions invested in AI via sky-high margins, and starting a (short) chain-reaction that causes the bubble to pop (or fizzle). A lot of people have a lot riding on the AI bubble not popping.

    • I didn't say it was a problem, I said it "mattered"

      If I wanted to argue that it's a problem, I'd just say that companies investing billions in training frontier models should reap the rewards. And distillation is essentially theft.

> Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.

They claim it because Anthropic are planning to push for protectionism. They just doubled their political spending to $40 million for the midterms to "push for AI regulation" Gee, I wonder what it is they are lobbying for. Certainly won't be OFAC sanctions right? ICTS import controls?

US GOV, under lobbying pressure from Anthropic and OpenAI are going to go full protectionism and restrict Chinese models, I'd almost be willing to bet money on it. They can't really enforce for individuals, but they can definitely tell US based hpyerscalers they can't host them, make it illegal to host the weights, and government procurement restrictions.

> The only argument they have here is that they use GB300 GPU's

I don't follow events closely, but the US has constantly flipflopped on what sort of GPUs the Chinese are allowed to have, not in small part because much of the AI boom's valuation is based on demand for US-made hardware, for which the Chinese have inexhaustible and well-financed demand.

So even this feels a bit hypocritical to me, but my understanding is that Chinese native AI hardware is getting good enough that labs dont feel a huge disadvantage by being forced to buy at home, even if they'd have preferred to buy US chips.

Which is a situation that was manufactured by the constant thread of having their access to advanced GPUs revoked.

It matters because it means that lab could not train that model without distilling another frontier model, and their progress would slow once they get properly cut-off. If I funded that lab, I would want to know that.

> Distillation is not illegal by every definition of the word.

I am waiting for a precedent on this one. In general, training on copyrighted material is legal, there is a lot of precedent there. But every now and then there is a case where the owner of the training material wins.

I don't remember the details but I believe one of these instances was when one company trained its AI on the knowledge base of another company and turned it into a competing product. Fair use was denied because of that direct competition. Distilling a LLM to make a competing LLM looks kind of like this, or maybe not, I don't know.

It would make sense for distillation to be legal in every way, LLMs are built on a broad interpretation of fair use, but sometimes, law is weird.

  • > I am waiting for a precedent on this one. In general, training on copyrighted material is legal, there is a lot of precedent there. But every now and then there is a case where the owner of the training material wins.

    You're making a fundamental assumption: that model outputs are subject to copyright. In the US that's only the case if a human is part of the creative process:

    https://www.copyright.gov/newsnet/2025/1060.html

    > It concludes that the outputs of generative AI can be protected by copyright only where a human author has determined sufficient expressive elements. This can include situations where a human-authored work is perceptible in an AI output, or a human makes creative arrangements or modifications of the output, but not the mere provision of prompts.

It also doesn't matter for a simpler, and much grander reason.

All LLMs are trained on the corpus of humanity's knowledge, the legacy of everyone who's ever lived and our civilization as a whole.

Anything that prevents or circumvents the accumulation or gatekeeping of this knowledge and puts it in the hands of more people (that are not AI company shareholders) is a good thing. Whether that is done by open sourcing the model weights, the training set, or by making the output better and cheaper, it is all fair game and is, as another poster mentioned, inevitable in the long run.

It certainly matters as familiar sounding words to their stock holders to please not drop the valuation.

Because what they want them to think is "the AI factory has unique proprietary technology that cannot be replicated"

What they don't want them to think is "it's relatively easy once you know the basics to bootstrap to near SOTA and so the commercial case for selling inference has an extremely short profitability horizon with little if any brand loyalty or lock in".

I wouldn't be surprised at all if US labs are also distilling Chinese models, except we'd never know since they can simply self-host them

  • We do know, because we used to have some Claude models identifying as Deepseek when prompted in Chinese

It does matter in that these LLM companies need to be run into the ground, and every embarrassing clod working for them run out of town.

It's showing that 'distillation' is a viable way to reclaim all of what they stole and hoard, and with enough luck their debts will come due in time for them to feel it.

> Does this matter? Distillation is not illegal by every definition of the word.

Correct, but it at least helps answer the question of "how do they make such good models for a fraction of the price???" The answer is someone else spends the untold billions and Chinese labs do a little tweaking.

It doesn't matter. It is most likely a pretext for upcoming actions mostly likely executed via yet another retarded executive order. The guy that posted this looks like he's drowned himself in the MAGA Koolaid.

Uh, what?

> Distillation is not illegal by every definition of the word

Note that Anthropic (and USG) alleges [0] not only that Kimi was distilled, but that they actively circumvented measures intended to stop distillation. There are multiple ways that's illegal, including:

- Civil breach of contract. Anthropic's TOS explicitly say you can't do what Kimi is alleged to have done.

- Economic espionage: 18 U.S.C. §1831 criminalizes obtaining a trade secret through theft, fraud, or deception while intending that it will benefit a foreign entity.

- Trade-secret misappropriation: if Anthropic could argue industrial-scale querying reconstructed proprietary aspects of Fable (like by showing it produces similar outputs, as others have done) then it's illegal under 18 U.S.C. §1832.

- California computer-access statute §502 bars knowingly accessing a computer system and, without permission, taking, copying, or using its data.

- Computer Fraud and Abuse Act protects against the case where restrictions against an activity are circumvented (like Kimi is alleged to have done).

> There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them.

A lack of prosecution does not make something legal. There is also the scale/commercialization thing, which isn't an issue with random tiny HF datasets/models. Remember: Kimi also sells K3 inference.

> kimi architecture is vastly different than that of fable

How do you know that? Do you work for Anthropic? Also, this has nothing to do with architecture, we are talking about data.

> US AI labs are inspired by opensource advancements just as much as open source labs are inspired by traces from models such as fable.

Cool. The difference is that one of those things is legal (because they chose to open-source) and one of those things is illegal theft of trade secrets (because it was stolen).

> Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.

1) this has nothing to do with other labs, just Moonshot (and Z.ai, MiniMax, DS)

2) slandering or not it happens to be completely true, so, there's that

[0] https://www.anthropic.com/news/detecting-and-preventing-dist...

  • All of the models stole the entirety of written knowledge on the internet to train. They are being sued for the few cases where we have some proof of what they did because of some whistleblowers, all the rest will just go unpunished. They breached Github TOS, robot.txt's, copyright, patents every form of IP protection under the sun from a billion sources. It's just ridiculous for the thieves to cry about someone else stealing from them.

  • I do agree that two wrongs don't make a right, the terms of service generally gives cooperation the power to sever the contract, but it does not make things illegal in the literal sense. The illegality usually comes from widescale fraud which includes accessing services you are banned from accessing.

    When I said "Does this matter?" I specially meant that distillation in itself, the data you get from distillation is first and foremost not owned by anthropic nor is it copyrightable. If a user willingly gives up their anthropic reasoning data/traces that is 100% legal no matter what the "terms of service" say as it's not enforceable and would fall apart in court.

    And what I explicitely pointed out that focusing so much on distillation is an attack on open research and claiming that the majority of advancements are thanks to US labs which is simply not true (at least not anymore this was somewhat true during deepseek R1 era), but that in itself was inspired by open research.

    > How do you know that? Do you work for Anthropic? Also, this has nothing to do with architecture, we are talking about data.

    Because anthropic would be the first ones to make that information public and the architecture is unique to kimi... They made it, they wrote papers on it, it's their research.

    P.S. none of the quoted laws apply here since no trade information is stolen, the one about circumventing distillation protection might hold up in court although unlikely.

    • > The illegality usually comes from widescale fraud which includes accessing services you are banned from accessing.

      Agree, and this is exactly what Anthropic is alleging.

      > data you get from distillation is first and foremost not owned by anthropic nor is it copyrightable. If a user willingly gives up their anthropic reasoning data/traces that is 100% legal no matter what the "terms of service" say as it's not enforceable and would fall apart in court.

      It's important to note this is NOT what happened. Anthropic was able to trace data directly back to employees at the company: "We attributed the campaign through request metadata, which matched the public profiles of senior Moonshot staff."

      > none of the quoted laws apply here since no trade information is stolen

      There is a lot of work showing Kimi models produce similar outputs to Anthropic models, which constitutes trade information. This is not dissimilar to past and ongoing IP suits against Anthropic and OpenAI by showing the models would recreate images of Mickey Mouse/NYT articles etc.

      For the record, I'm a researcher myself and I'm well aware how competent the researchers are at the open-source labs/how much they've contributed. But that's not at issue here, my disagreement with you is specific to your arguments about legality; you're conflating what you think should be legal with what actually is legal.

      2 replies →

  • >Civil breach of contract. Anthropic's TOS explicitly say you can't do what Kimi is alleged to have done.

    This is true, but Kimi also has a variety of defenses. Kimi can't raise unclean hands if Anthropic systematically violated others' terms of use, but it can raise copyright misuse (which is similar in some respects to unclean hands) as well as lack of standing to enforce restrictions in the contract due to the third party beneficiary principle (i.e., Kimi would argue that Anthropic cannot sue Kimi for derived IP that rightfully belongs to third parties whose terms of use were violated by Anthropic, and the proper party to sue Kimi, if any, would be those third parties). That latter argument usually fails in small-scale cases (ProCD) but has been successful in larger ones where the alternative would be anticompetitive.

  • >"A lack of prosecution does not make something legal"

    Plainly who gives a flying fuck. The US can claim whatever rules they want and so can China or any other country. On international level all those rules are artificial constructs unless they can be enforced. China can just say for example that they do not recognize copyrights /patents / whatever so it is "legal" for them.

    • This is illegal in China too, there's just an enforcement asymmetry. I understand what you're saying is de facto true, I'm just taking issue with people saying either

      1) its not illegal (it is)

      2) it shouldn't be illegal because Anthropic stole training data (thats not how the law works)

      1 reply →

  • > Civil breach of contract. Anthropic's TOS explicitly say you can't do what Kimi is alleged to have done.

    Ah yes, I remember when Anthropic crawlers abided by the TOS of the websites they slurped up.

    All your other points are downstream from this, which makes them pretty tenuous. Labs don't think that ToS or other explicit wishes of content providers apply to them, but they expect everyone else to abide by theirs.

    • To be clear, I think theft is also bad when Anthropic does it.

      US and CA law really don't care that Anthropic violated IP law elsewhere.

      1 reply →

It matters because the closed-source frontier labs spend lots of money on human data (RLHF / RLAIF with human oversight). Moonshot is accused of circumventing these costs. Frontier labs add research costs into their inference pricing. If the market doesn't permit them to sustain sufficient pricing to have a positive cash flow, then their business prospects become weaker and they risk insolvency. Furthermore, other leveraged companies are at risk.

The reason why the United States government is weighing in is because it's in the national interest of the US to have supremacy in "AI".

Legality or lack thereof is one of many data points about whether a thing is noteworthy.

Moonshot performing distillation is rational from their point of view. Reducing costs is in the interest of businesses. It's also rational for frontier labs and the US government to add obstacles to this process.

As consumers this is probably a positive development.

  • My parents put in countless hours and tens of thousands of dollars into raising me to the point where I could write an answer on StackOverflow

    And OpenAI scraped and distilled that answer and gave me nothing

    • What does your story have to do with Moonshot AI? Do you think they didn't also use the same corpus? Bizarre

    • And now people such as myself have access to open weight models with that information. I wasn't lucky enough to have parents put me through school, and LLMs have absolutely helped me further educate myself and play "catch up" on opportunities others have been given. So, the net effect has been (and is continuing to be) a democratization of information.

      6 replies →

  • Isn’t that the same argument they are making for replacing human labour?

    Circumventing costs.

    • There are many frames that one can place upon this issue. They do not contradict the other. There are moral framings (stole the internet so go eff yourselves, is one), but so is national security, and so is the doomer recursive self improvement risk, and then there is the framing purely on what this implies for future AI training.

      I mainly focus on the last.

      It will be hard for a frontier lab to justify spending the compute and data curation needed to advance AI further if that expenditure can be assimilated into your competitor's products within months/weeks. So reality will present labs with three choices:

      A. Cease spending massive amounts of money and compute improving those models.

      B. make those improved models more difficult to distill from, either through some regulatory regime, or some technical solution, which seems unlikely to me.

      C. making the best models available only to select partners and government.

      In all these potential outcomes, China, which lacks compute that U.S. labs enjoy, will likely stop seeing massive improvements in their AI models. Improvements to be sure, but right now they are enjoying gains from distillation AND their own model innovations, and these potential outcomes would largely stop one of those sources.

  • Oh, so mass theft is okay as long as American companies are doing it

  • Chatgpt routinely cites and uses papers I don't have access to because they're behind a paywall. I don't think OpenAI is paying for all that copyright. That's in my opinion way more serious.

    • Yes the fact that the scientific literature - created largely on the back of the tax payer - isn't open to all free of charge by force of law is a travesty. A cartel should not get to charge for access to the bulk of human knowledge. That is indeed a far more important issue than whether or not Moonshot violated the Anthropic ToS, possibly committing mass fraud in the course of doing so.

      I mean honestly if they did that why should I care? I'm happy to see copyright violated in a manner that leads to the creation of new technology. IP law exists strictly for the benefit of society and by all appearances AI is an incredibly powerful tool.

      Also while I'm at it libgen is a gift to humanity. Information wants to be free. Spreading and preserving knowledge is generally one of the most wholesome activities anyone can undertake as far as I'm concerned.

    • Your statement is orthogonal to my comment. Why reiterate the schadenfreude / fairness comment already stated several dozen times in this thread?