← Back to context

Comment by biophysboy

20 hours ago

Why is it unlikely?

Because the models are trained on hundreds of billions of user conversations, across more than a billion different humans. The conversations are anonymized and not easily traceable back to a specific user.

It's unknowable and not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes.

We also don't know if the authors unintentionally provided data to OpenAI through alternate means, such as via alternate accounts or model feedback queries.

  • Anonymized doesn't mean there's no way to know whether it is in there. My ballot is anonymized, but it's known to be in the box because a checkmark was put next to my name when my ID was verified. OpenAI can trivially check their account settings to know what happened to their chats. The fact that they are being vague about this likely indicates that they have already done so and discovered that the data did go into the training set.

    Further, given that this is all in the open now, they can search the training data. No way somebody is using some specific unique cutting edge mathematical approach to solve a fluid dynamic problem 99.9% of people have never heard of and it's not locatable. Considering they spent $15,000,000 already on this, they could afford to grep around to be able to state that their hands are clean.

  • > It's unknowable

    Knowability and likelihood are almost orthogonal here. If I commit a crime and perfectly destroy the evidence, my deed may be unknowable. That doesn’t make it more or less likely.

    > not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes

    It may be. We haven’t seen the researchers’ transcripts. We don’t know what Buckmaster or his co-author uploaded to OpenAI or with what permissions (or if OpenAI actually respects those toggles).

    • I think OpenAI desperately wants to make a blanket denial that they didn't look at or train on Buckmaster and Alpöge's chat transcripts, but know they cannot, because the data is anonymized.

      The fact they can't make a blanket denial triggers everyone's bullshit detectors, and they're getting eviscerated over it.

      1 reply →

  • I understand that AI is not just cut and paste, but some documents will have more influence than others w/ power law scaling. I would be very surprised if this distribution were not extremely steep for arcane math

Their base model must have been trained with hundreds of trillions of tokens several months ahead, at this point of time, it is impossible to rule out the possibility the model had seen that session at one point of time, and it probably did, without any OpenAI personnels actually know about it.