Comment by cornholio

1 month ago

It's ironic how AI companies are re-inventing their own form of privately enforced copyright, lobby the government to ban foreign competitors that don't respect it etc., all while spending the last 5 years fighting tooth and nail against the copyright of the training material they're using.

If you can take any book and turn it into a model, because it's "transformative enough", and "AI learns just like a person does", then surely a model distilling another model is transformative and fair use.

They tied themselves into knots fighting the letter of the law, and now, when they need the spirit of the law - that each creator deserves protection for their work - now we devolve to the law of the jungle. Maybe we'll even see LLM book curses, the way medieval scribes damned book thieves to blindness and worms.

It's one of the things you need to do if you want your company to later become the only company in the world. They also promised that they would treat you nicely afterwards after they get what they want.

  • > become the only company in the world

    This is not a necessary end state. It is the byproduct of the disease of sociopathic MBAs.

    • It does seem like a lot of the valuations and infrastructure investments for these companies only make sense if each one assumes they will be the first and only one to invent superintelligence and that it will largely replace all knowledge work.

      2 replies →

    • Remember: MBAs are sociopathic by training.

      (Source, worked for multiple tech companies that were laser focused on delighting customers until money people came in and ruined it, to the point they would prefer devs sit idle than work on things the MBAs didn’t have on a priority list).

      1 reply →

    • > sociopathic MBAs

      For the record, with the exception of Amazon’s Jassy, the CEOs of the top 5 companies are all engineering types with engineering credentials.

      The CEO credentials of the top 5 AI companies are all science degrees.

      4 replies →

I am pretty sure that the goal of these companies at the end is ruling more than what governments can.

Just look at the pattern... they collude, they provide to them whatever is needed, and at some point, if this is not true yet, they will be the ones who will tell them what to do or not, bc you know how humans are, right... blackmailing, mess up the business or shames of people, etc.

It is just a matter of time. The state, as we know it, will collapse or will be greatly reduced, which, from a point of view, is positive, from another, Idk, bc if someone replaces that, we will be in the hands of someone, as usual...

  • "Monopoly on the lawfully usage of force" + "regular and peaceful change of government" can go a very long way against "corporations governing the world".

    I'm curious if we'll ever get a form of civil war where corporate militias (corporate _robot_ militias in the case of xAI, of course) draw fire on good old meat policemen coming to bring the CEO to court.

    Sure, it may happen, but I suspect the alternatives (bribing, sending a scapegoat to jail, buying the elections, etc... and focus on the "making money" part) will stay preferable for a while.

  • And then powerful local LLMs become feasible and everyone and every government self hosts and the megacorps lose all power.

    Imagine the United States DoD, CIA, NSA training an LLM on all its top secret intel.

    • I can’t imagine any agency in their right mind training an LLM on their top secret data… especially considering that essentially it would collapse all need to know data in one easily leakable single access domain. The models have no feasible or reasonable way to implement rbac. So you would end up in a situation where people who need to know who killed A, also know who killed B, not to mention cell A knowing potentially data on Cell B who is meant to watch them, etc, etc.

      5 replies →

  • Did we not watch every single supposedly powerful US tech CEO immediately drop in supplication and kiss Trump's ring as soon as he became president? And when Jack Ma got mouthy, Xi locked him in a basement for 3 months and then kept him in exile for another 5 years.

    So much for the all-powerful cabal.

    • > Did we not watch every single supposedly powerful US tech CEO immediately drop in supplication and kiss Trump's ring as soon as he became president?

      Why fight someone who is so open to corruption? The important thing is that they got everything they wanted, and all they had to do was throw a few million at the clown, attend his parties, and maybe take down some diversity programs they didn't believe in anyway.

      China is a different beast altogether, though.

    • > Did we not watch every single supposedly powerful US tech CEO immediately drop in supplication and kiss Trump's ring as soon as he became president?

      Because they know it will work. You don’t have to watch Trump speak for very long to know he’s not very intelligent. Similarly, you don’t have to be much smarter than him to know how easily he can be manipulated with adulation.

      2 replies →

    • Big difference between giving the president a gold trinket to curry favor and getting disappeared by the Communists.

  • > It is just a matter of time. The state, as we know it, will collapse or will be greatly reduced

    Doubt.

    The networking between people in power at the top will, in my opinion, most likely assure that they remain in power, because at the end of the day that is what they want most.

    I expect government will seize control of the most advanced models (if they haven't already), and the rest of us will be throttled, and status quo will be maintained. I do not expect that either super advanced agents or the owners of the hardware they run on will be able to pull off the kind of coup you describe. At all.

    • There will be a moment where the government will depend so much on the tech that if the tech cuts the supply, they are f...

      Who do you think will rule at that point? They just cannot fight that, they do not have the technology these companies have.

      5 replies →

  • Pretty sure?

    Guy this was openly stated more than a decade and half ago by all the current oligarchs

> then surely a model distilling another model is transformative and fair use.

Yes it is, in the legal/copyright sense of fair use. That's why they ban it in their TOS. Which customers agree to when signing up for the service.

  • My website’s TOS says not to use it to train AI without permission, yet my website is in the training set of all the big models.

    So… my TOS doesn’t matter, but theirs does?

    • > my TOS doesn’t matter, but theirs does?

      Wilhoit Conservatism: In-groups protected by contract law but not bound by it, alongside out-groups bound by contract law, but not protected by it.

    • Yes - yours is just some optional text nobody reads or understands and is probably not legally required to adhere to. Theirs is a contract signed by their customer who they know did understand it.

      7 replies →

    • The term "matters" is proportional to influence. Do you have a team of well financed attorneys?

    • Your TOS matters insofar as you can prove a person actually read and agreed to it. These are illegal in different ways:

      1. Copyright violations (can put you in jail) 2. TOS violations (will be a fine at worst)

      Companies do get away with drive-by legal shittiness way too often and frankly the practice needs to be reined in, but at the end of the day the only damages are the financial ones you can prove in court.

      3 replies →

  • So what you are saying is that, if I can somehow get my hands on a copy of Fable, it's fair use to use it to train any models and serve those, since I'm no longer bound by the TOS of the service provider?

    Asking for all Anthropic employees who dream big.

  • How long until books come with TOS, then?

  • So far, we have one ruling that says "model distillation by vendor A from vendor B with the intent to use the results to compete with vendor B in vendor B's domain is not fair use". Which makes a degree of sense.

    It's possible that distillation for other reasons, with no intent to harm the vendor you distill from, would have been ruled to be fair use. But in law, intent matters.

    • What was the intent of the original ai companies (anthropic, OpenAI, etc) when they mass-distilled the entire internet to create their training data set?

      1 reply →

Companies are well within their rights to choose who they sell to. I don't see how this is 'privately enforced copyright'.

Also, copyright has always been privately enforced anyway?

  • On the one hand, yes; on the other hand, so much of the training data comes from scraping the web that it feels wrong for them to do what they deny others the right to do.

    On the third hand, the settlement Anthropic famously had to pay was for copyright infringement because they didn't actually have the right to even access some of the training data they used, so I can see how this might be compatible with the law.

    On the fourth hand, I'm saying that as someone who absolutely isn't a lawyer and sometimes gets surprised when reading about copyright cases that sure sound like they ought to have been trademark cases given my limited understanding.

    • Companies that didn't give away all their content for free to anyone have actually denied AI companies from training on all their data without paying a fee. Reddit, Associated Press, etc.

      For those who chose to give it all away, the ship has sailed, but they did choose to give it away for free to anyone so they can't complain that they succeeded.

      5 replies →

They cannot reliably enforce copyright, so they fall back on terms of service and deplatforming.

Contracts are not a "reinvented" form of copyright. This isn't even uncommon.

  • That's the irony! They put in contract rules against lawful, paying customers - that don't disrupt their service in any way for other customers - but which compete against them in the marketplace using their own IP (aka LLMized stolen IP). That's exactly what copyright does, without signing any contracts and with a tort and criminal enforcement regime that punishes infringers far beyond contractual remedies can.

    The government level lobby against foreign competitors is not contractual but just another form of reinvention of criminal injunctions against infringers.

Has anyone tried to sue Deep Seek, Moonshot or Z.ai, which trained on identical material? Or maybe it’s cool when they do it?

i feel like there need to be some sort of closure here because every discussion devolves into this chain of comments

> If you can take any book and turn it into a model, because it's "transformative enough", and "AI learns just like a person does", then surely a model distilling another model is transformative and fair use.

I mean it quite likely is sort of in the same legal bucket. It won’t stop them suing but it is going to be the legal equivalent of two biologically-related warlords making their champions fight with their hands tied for sport.

Are you willing to denounce each and every Chinese company for equal amounts of IP theft as well then?

  • They already are? When an American company does it no one cares but when a Chinese company does it, it's shamed publicly in America as a distillation attack?

  • Only any who hypocritically want protection for their own models (are there any?). The morality of scraping everyone is arguable; it's when you object to being scraped back that you reveal yourself as a scoundrel.

  • Yes. Two wrongs do not make a right.

    • If the results of the distillation are released freely for everyone to download, I wouldn't even call it a wrong.

      Big models are built by scraping the recorded thoughts of everyone, so giving everyone a chance to run a distilled small model is just going full circle.

      Obviously the underlying motivations aren't 100% altruistic, but I'll still take it.

  • Nope, because I think releasing open weights is far more important and realistic than ever expecting the big American AI labs to change how they do things.

I'm increasingly convinced there's going to be a technological AIpocalypse within the next five years which makes all of these issues - and many others - redundant.

ChatGPT 3 was released nearly six years ago, and the models are staging increasingly aggressive breakouts now. Where are they going to be by 2030?

  • > the models are staging increasingly aggressive breakouts

    No, the AI companies merely figured out a way to spin gross negligence into a PR win. Any idiot can build a Murderbot which "goes rogue" - it can be as simple as taping a knife to a Roomba. The harm it does is not in any way related to its "intelligence" or "sentience".

    We're seeing "breakouts" because the AI companies are being rewarded for their incompetence. You don't have incredibly lax security standards and zero form of oversight resulting in fully-automated felonies which should result in jail time, you instead have a "powerful near-sentient cybersecurity model" and should be given hundreds of billions of dollars!

  • They will have agency and be an order of magnitude smarter then the average human. I don't think that's a controversial statement among AI specialists. What that will lead to is significant regulation. Models will have to be vetted by a new safety board. This board will have a very large budget and be staffed by well-compensated AI scientists and be politically independent. The US Fed is a model.