← Back to context

Comment by tedsanders

17 hours ago

All of those statements sound true, based on what I've heard.

- "very little human" input feels ambiguous, and if someone spends a few days prompting a model to solve a super hairy problem requiring a 100-page proof, I can understand reasonable people interpreting that as both "very little" and "not very little" human input

- it's all true that a team worked on this, a bunch of compute was burned, and the problem was solved in stages and pieces

I'm not sure how any of this provides evidence that OpenAI took any of their work.

As evidence against, we never looked at any of their ChatGPT conversations and our model's proof is quite different from theirs.

(I work at OpenAI, but not on the team that did this proof.)

I'm confused, your employer very directly stated that they are unable to confirm that the model was not trained on the conversations.

  • The models are trained on the conversations of hundreds of millions of people. ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.

    It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.

    • Do you not log the training data? Seems like you should be able to just check what was in the training data. To not keep track is just sloppy work and certainly unprofessional science.

      8 replies →

    • > ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.

      How many of those trillion conversations were about Navier-Stokes you reckon?

    • I agree that it is not possible to prove if any one specific conversation (or derived RL tasks) was key to solving Navier-Stokes (at least without massive resource expenditure).

      I don't really understand how the quantity of training data/rollouts used in training is relevant to the question of whether or not it was trained on these conversations.

      I also don't really believe that whether or not this model was trained on these conversations is unknowable information.

      1 reply →

    • >It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.

      If the conversation was in the training set, there's a high likelihood that the small set of conversations related to solving Navier-Stokes was used by the model. I get Astra to still quote some of my friends' books or blogposts nearly verbatim on certain niche issues.

      Much more importantly, we _can_ determine whether a conversation was used in the training data. And if it was, it gives us a great idea whether that logic was captured in reasoning for a novel problem never yet solved.

      Given that you don't see any of this as below the belt according to your other comments, maybe your contribution here is more for yourself than a fair conversation about attribution.

      9 replies →

    • Sorry but this is a misconception: these models are both capable of complete novelty and of plagiarism. For a concrete example, image diffusion models have been shown to reproduce many existing images nearly 100% exactly, yet clearly, they can also create new ones.

      A model being trained on lots of irrelevant information does not mean relevant information was not used.

You don’t work on the team that did the proof yet you can with certainty make all of these claims?

It's unclear if you're suggesting that OpenAI did not train on their input or use their chats as inputs to training on a model that found the solution. Let's not provide an Elizabeth Holmes-esque interview where the question is dodged and words gain new meaning. The question can be answered with "Yes, we trained on their conversations" or "No, we did not train on their conversations".

I'm not coming from a place of distrust here. This should just be definitively answerable given the weight of the claims here. Surely between you, your lawyers, and other members of your team you can just clear this part up.

>> I'm not sure how any of this provides evidence that OpenAI took any of their work.

Sorry, but the burden of proof lies in the other direction: OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published. Otherwise OpenAI simply stole the glory and the spotlight (and I'm being charitable here).

  • That's entirely unreasonable. Allegations of malfeasance always need to be backed up by evidence.

    • But there is evidence, the blog post says: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."

      In other words, yes, they had been using ChatGPT, and yes, ChatGPT could very well have trained on their data. Now that there is evidence, we need an investigation: yes or no, was it the case?

      3 replies →

    • But the evidence is in the hand of the potential culprit. That's why allegations can be enough to force confiscation and intrusion to get evidence in safe hands before it is destroyed by the accused party.

      6 replies →

    • That is backwards. It is the responsibility of a researcher to do a thorough literature review and conscientiously avoid plagiarism or claiming false novelty.

  • > OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published.

    I don’t think they’re too concerned about appeasing you, enraged_camel.

    For most reasonable people, achievement in solving the other Millenium Prize problems at an unprecedented rate will be enough. At some point people will see models are capable of solving hard issues without whatever 0.00001% of the training data coming from irate individuals who believe their sample was the key component of the solution.