← Back to context

Comment by WarmWash

11 hours ago

It would be catastrohic for any of the big labs if it came out that they were training on what was sold as private.

I get this cynical conspiratorial energy, it fits the internet well, but I can assure you most people with even mild business sense would be intensely opposed to this idea. Well, except maybe Zuckerburg, but they don't really do enterprise anyway.

No it wouldn't.

It wasn't "catastrophic" for the largest of the 3 US credit reporting agencies when their entire dataset was breached. The company is 100% IP and the only value they have was completely copied. Their largest value is to verify identities by the things Americans know (KDB) and after that "single factor of identity" was 100% compromised, the company only got bigger and more contracts.

When there are only 4 competitors in the large scale foundation model business and they all throw caution to the wind because they are racing to own the "$30 trillion TAM" they are all going to make critical security, RBAC, and segregation mistakes.

Both ChatGPT and Claude threads marked for sharing have been indexed in Google at large scale. This is incredibly easy to tell Google crawlers via robots.txt not to crawl those URLs, but nobody at either of these uber unicorns could be bothered to add that one pattern to the one file.

And all of the skepticism here is about verifiability. The foundation model companies are liable for potentially more the companies are worth if found to be violating copyrights of content used for training. They aren't going to make it easier for lawsuits against them by detailing their data ingestion into training pipeline.

  • The credit reporting agencies lost the data of normal people, they didn’t lose their customers proprietary internal data. The credit agencies didn’t loose or misplace their customers data, so obviously their customers don’t really care that much, and the credit agencies weren’t sued into oblivion.

    But I can guarantee you that if a companies internal data got leaked or misused, then every single enterprise customer of that lab would turn around and start suing them. As an enterprise customer you would be foolish not to, if only if figure out via discovery just how badly you got screwed.

    You want to see how nasty that can get. Just go and look at what Apple is doing to OpenAI at the moment. Do you really think Apple wouldn’t find a way to sue a lab into oblivion if they discovered a lab had secretly started training on their data?

  • that was accidental? this would be straight up fraud.

    • You can structure it so that it becomes accidental.

        1. Ensure security barriers are weak or honor based.
        2. Put individual researchers under a lot of pressure.
        3. If you get caught, blame the weak barriers, or the individual researcher.
      

      Basically setup the incentive structure to incentivize researchers sticking their mittens in the private cookie jar while putting the cookie jar in a dark unmonitored/unsecured room with a sign on the door saying please don't enter.

exactly, I don't understand how HN doesn't understand this

by this logic every business contract in tech is just a bunch of lies and means nothing and the only way to do anything is to have a server sitting next to you, otherwise it's "someone else's computer"

  • I mean, they've already violated the law in acquiring all their training data already, why would they be uncomfortable violating a contract to get more training data?

    • Because violating the law was the prerequisite to starting their business, without it they're worth $0 and have no models. They've already survived Training on customer input may help the models but it isn't "bet the farm" helpful. Now, they have a thriving business, so they shouldn't risk their business for incremental data that they can buy.

      At this point, the reputation of the business matters too. Fable explicitly didn't support zero-retention usage, and it saw significantly lower adoption vs other flagship models, and their past releases. Being caught abusing enterprise contracts is really hard to dig out of.

It would be catastrophic if they violated confidentiality blatantly, but that doesn't exclude learning of any description. For example, an ordinary human being can't fork a subagent for a particular client and wipe it afterwards. Humans can't stop themselves learning, so confidentiality can't ban all learning.

Instead, confidentiality includes not literally copying material, not using trade secrets or inventions, and not using knowledge of business dealings for your own purposes.

So, while I fully expect that the big labs don't train on private material to the extent that they do those things in a blatant way, it would not be surprising if they pushed the boundaries. Humans push the boundaries all the time.

Up until now, machines did not have judgement, so if you set up a machine in such a way that you hadn't ensured it couldn't violate contract, you were culpable. But now that they have some kind of judgement, maybe it's enough to avoid liability to tell it to obey the contract, even if you give it incentives not to. After all, that's how it works with human employees, isn't it?

Perhaps now we have machines that understand language, someone somewhere is working on getting them to understand "a nod and a wink" as well.

It's going to be the same as PRISM, people will be outraged and business will go back as usual.

(And they totally won't do it again they swear, the contract says so)

I mean the models have literally trained on:

1. Child porn

2. Stolen music

3. Private github repos, before that was 'stopped'

4. Illegally pirated books

Them training on company prompts against the terms of service would be one of the least bad things that these companies have trained AI models on

Why do you think a company - willing to break the law for child porn - won't break the law when it comes to your personal data?

  • These folks need to feel repercussions so hard their souls flee to the afterlife leaving only their sad, dead husks behind.