Comment by nehan
18 hours ago
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
I think they should be able to unravel whether or not any sessions by Tristan or Levent went into the training data for this model.
If they could then it wouldn't be de-identified data...
Well you could search for elements similar to the proof / problem in the training data, even if it's de-identified, right? OpenAI can probably do better than Ctrl-f "Navier Stokes".
There are probably thousands of serious academics taking a crack at millennium problems using AI every day. All those attempts are in the training data. And in fact the two researchers benefited from those attempts as well.
The researcher could share a string from one of their conversations and OpenAI can confirm whether it exists in their training data.
Or OpenAI could just look at their code and say what it does (maybe have their AI do it if they're having so much trouble with this?)