Comment by tristanj
14 hours ago
The models are trained on the conversations of hundreds of millions of people. ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.
It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.
Do you not log the training data? Seems like you should be able to just check what was in the training data. To not keep track is just sloppy work and certainly unprofessional science.
That training data does not preserve provenance seems a "smoking gun" in terms of intent to plagiarize.
The burden of proof is on Buckmaster and Alpöge to reveal if they had the "Improve the model for everyone" setting enabled or disabled. OpenAI shouldn't be expected to reveal private user configuration data. You're asking them to perform a user privacy violation.
If the reason that OpenAI is unable to state whether they trained on this data is because they (as policy) do not reveal whether a given member has turned on/off the "Improve the model for everyone" setting, they can at least say so.
FWIW, publicly facing OAI docs are very unclear about whether this setting even applies to Codex conversations.
4 replies →
Does Anthropic, Grok, etc. log their training data? I had the impression it was rather a mess.
> ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.
How many of those trillion conversations were about Navier-Stokes you reckon?
I agree that it is not possible to prove if any one specific conversation (or derived RL tasks) was key to solving Navier-Stokes (at least without massive resource expenditure).
I don't really understand how the quantity of training data/rollouts used in training is relevant to the question of whether or not it was trained on these conversations.
I also don't really believe that whether or not this model was trained on these conversations is unknowable information.
[dead]
>It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.
If the conversation was in the training set, there's a high likelihood that the small set of conversations related to solving Navier-Stokes was used by the model. I get Astra to still quote some of my friends' books or blogposts nearly verbatim on certain niche issues.
Much more importantly, we _can_ determine whether a conversation was used in the training data. And if it was, it gives us a great idea whether that logic was captured in reasoning for a novel problem never yet solved.
Given that you don't see any of this as below the belt according to your other comments, maybe your contribution here is more for yourself than a fair conversation about attribution.
The flaw with this line of reasoning is that Buckmaster and Alpöge only had a partially completed proof of a weaker version of the Navier-Stokes problem. OpenAI's internal model solved the full, harder problem. This means the key information needed to bridge the gap was not present in Buckmaster and Alpöge's chat history.
You might retort that ChatGPT used the training data to copy their approach, but the approach Buckmaster and Alpöge chose was already published by Luis and Diego in 2023 and in every frontier model's training set.
This argument proves too much. By this standard, it wouldn't have counted as copying their approach if the researchers had just fed in Levent & Buckmaster's paper verbatim as a prompt into the swarm.
1 reply →
I don't believe people are denying that the model is impressive. The problem is that learning someone else is making progress on a topic using method X and then rushing to scoop them borders on academic misconduct. If, on top of this, their private conversations about X were used in the proof, I really don't see how its defensible...
2 replies →
I’m pretty sure you think you are doing a good job of defending your employer and you probably believe “Open”AI are the good guys here. I also acknowledge that they butter your bread so your financial future currently depends on their success.
However the way you are conducting yourself in public, while announcing yourself as an OpenAI employee is doing enormous harm to the greater and magnanimous aim of your organisation. Take a step back and read the temperature of the room. Being the smartest guy in the room will never protect you from alienating the rest of the room into a baying mob. Right now you are Icarus flying straight into the sun.
2 replies →
Sorry but this is a misconception: these models are both capable of complete novelty and of plagiarism. For a concrete example, image diffusion models have been shown to reproduce many existing images nearly 100% exactly, yet clearly, they can also create new ones.
A model being trained on lots of irrelevant information does not mean relevant information was not used.
The IP laundering machine strikes again.