Comment by tristanj
14 hours ago
Because the models are trained on hundreds of billions of user conversations, across more than a billion different humans. The conversations are anonymized and not easily traceable back to a specific user.
It's unknowable and not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes.
We also don't know if the authors unintentionally provided data to OpenAI through alternate means, such as via alternate accounts or model feedback queries.
Anonymized doesn't mean there's no way to know whether it is in there. My ballot is anonymized, but it's known to be in the box because a checkmark was put next to my name when my ID was verified. OpenAI can trivially check their account settings to know what happened to their chats. The fact that they are being vague about this likely indicates that they have already done so and discovered that the data did go into the training set.
Further, given that this is all in the open now, they can search the training data. No way somebody is using some specific unique cutting edge mathematical approach to solve a fluid dynamic problem 99.9% of people have never heard of and it's not locatable. Considering they spent $15,000,000 already on this, they could afford to grep around to be able to state that their hands are clean.
> It's unknowable
Knowability and likelihood are almost orthogonal here. If I commit a crime and perfectly destroy the evidence, my deed may be unknowable. That doesn’t make it more or less likely.
> not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes
It may be. We haven’t seen the researchers’ transcripts. We don’t know what Buckmaster or his co-author uploaded to OpenAI or with what permissions (or if OpenAI actually respects those toggles).
I think OpenAI desperately wants to make a blanket denial that they didn't look at or train on Buckmaster and Alpöge's chat transcripts, but know they cannot, because the data is anonymized.
The fact they can't make a blanket denial triggers everyone's bullshit detectors, and they're getting eviscerated over it.
> fact they can't
Again, see the Apple lawsuit. OpenAI has never been constrained by facts in what it can and can’t say.
To the extent anything is setting off my bullshit detector, it’s in the idea that this time is different (Moreover, the idea that we should assume this divergence without evidence.)
OpenAI doesn’t have the benefit of doubt. They shouldn’t for anyone who’s honest and reasonable. That doesn’t mean they’re automatically at fault. But when the twentieth person comes forward and says a pattern is continuing, I’m giving them the preliminary benefit of doubt. It’s a bit extreme to conclude based on that. But it’s far more baseless to swing to the other side and claim we need to clear the table for a serial offender.
I understand that AI is not just cut and paste, but some documents will have more influence than others w/ power law scaling. I would be very surprised if this distribution were not extremely steep for arcane math