← Back to context

Comment by 20k

10 hours ago

You have to be rather naive if you don't think these companies don't simply train on your prompts with or without your consent. They literally scrape everything - legal or not - and claim its fair use to train on, including straight piracy

The idea that they'll steal from everyone except you is just wishful thinking

And even if they are not doing it right now, they probably keep the history available for future training, "just in case".

This is why these companies paying subscriptions shoveling everything into Claude thinking "oh, they're not training on our stuff" is hilarious to me. Of course they are. They probably are extra super-duper sure not to have Claude admit to that in any way, but there's no way they're giving up training on the sum total of both open and closed source code out there.

If you have an enterprise contract with them, the legal protections for the consumer is much higher… is what I’ve been told anyway

  • It isn't, in both cases you're protected by exactly the same legal system

  • Not sure what you are saying exactly, but

    > the legal protections for the consumer is much higher

    You know you give away the right to file class action law suits against Anthropic when you accept their Terms Of Use, right? (at least the Americans ones)

  • If you have an enterprise contract with them you're still unlikely to ever know that you've been lied to unless some whistleblower at the AI company comes forward and even if you do somehow find out, it's too late. Once they have your data and have trained on it there's no taking it back. At most they'll pay out some tiny settlement that's a fraction of how much money they make in a week and they'll continue to profit from your data forever.

Pinky-promises, embarrassing. I'd point to tinfoil.sh. I'm not a shill for tinfoil, I haven't even used it or looked past its homepage really, but if we are able to legitimately secure privacy, then training data questions are moot (and the market for training data would probably shift/expose).

It would be catastrohic for any of the big labs if it came out that they were training on what was sold as private.

I get this cynical conspiratorial energy, it fits the internet well, but I can assure you most people with even mild business sense would be intensely opposed to this idea. Well, except maybe Zuckerburg, but they don't really do enterprise anyway.

  • No it wouldn't.

    It wasn't "catastrophic" for the largest of the 3 US credit reporting agencies when their entire dataset was breached. The company is 100% IP and the only value they have was completely copied. Their largest value is to verify identities by the things Americans know (KDB) and after that "single factor of identity" was 100% compromised, the company only got bigger and more contracts.

    When there are only 4 competitors in the large scale foundation model business and they all throw caution to the wind because they are racing to own the "$30 trillion TAM" they are all going to make critical security, RBAC, and segregation mistakes.

    Both ChatGPT and Claude threads marked for sharing have been indexed in Google at large scale. This is incredibly easy to tell Google crawlers via robots.txt not to crawl those URLs, but nobody at either of these uber unicorns could be bothered to add that one pattern to the one file.

    And all of the skepticism here is about verifiability. The foundation model companies are liable for potentially more the companies are worth if found to be violating copyrights of content used for training. They aren't going to make it easier for lawsuits against them by detailing their data ingestion into training pipeline.

    • The credit reporting agencies lost the data of normal people, they didn’t lose their customers proprietary internal data. The credit agencies didn’t loose or misplace their customers data, so obviously their customers don’t really care that much, and the credit agencies weren’t sued into oblivion.

      But I can guarantee you that if a companies internal data got leaked or misused, then every single enterprise customer of that lab would turn around and start suing them. As an enterprise customer you would be foolish not to, if only if figure out via discovery just how badly you got screwed.

      You want to see how nasty that can get. Just go and look at what Apple is doing to OpenAI at the moment. Do you really think Apple wouldn’t find a way to sue a lab into oblivion if they discovered a lab had secretly started training on their data?

  • exactly, I don't understand how HN doesn't understand this

    by this logic every business contract in tech is just a bunch of lies and means nothing and the only way to do anything is to have a server sitting next to you, otherwise it's "someone else's computer"

    • I mean, they've already violated the law in acquiring all their training data already, why would they be uncomfortable violating a contract to get more training data?

      1 reply →

  • It would be catastrophic if they violated confidentiality blatantly, but that doesn't exclude learning of any description. For example, an ordinary human being can't fork a subagent for a particular client and wipe it afterwards. Humans can't stop themselves learning, so confidentiality can't ban all learning.

    Instead, confidentiality includes not literally copying material, not using trade secrets or inventions, and not using knowledge of business dealings for your own purposes.

    So, while I fully expect that the big labs don't train on private material to the extent that they do those things in a blatant way, it would not be surprising if they pushed the boundaries. Humans push the boundaries all the time.

    Up until now, machines did not have judgement, so if you set up a machine in such a way that you hadn't ensured it couldn't violate contract, you were culpable. But now that they have some kind of judgement, maybe it's enough to avoid liability to tell it to obey the contract, even if you give it incentives not to. After all, that's how it works with human employees, isn't it?

    Perhaps now we have machines that understand language, someone somewhere is working on getting them to understand "a nod and a wink" as well.

  • It's going to be the same as PRISM, people will be outraged and business will go back as usual.

    (And they totally won't do it again they swear, the contract says so)

  • I mean the models have literally trained on:

    1. Child porn

    2. Stolen music

    3. Private github repos, before that was 'stopped'

    4. Illegally pirated books

    Them training on company prompts against the terms of service would be one of the least bad things that these companies have trained AI models on

    Why do you think a company - willing to break the law for child porn - won't break the law when it comes to your personal data?

    • These folks need to feel repercussions so hard their souls flee to the afterlife leaving only their sad, dead husks behind.