← Back to context

Comment by hexomancer

18 hours ago

So you definitely did train on their data, you just think it is unlikely that it impacted the final model significantly?

I have no idea if their data was trained on. For example, if they used ChatGPT, asked a math question, and clicked the thumbs up button, that could have provided a small reward signal. I highly doubt this sort of feedback made a difference to a problem like Navier-Stokes, but it's not something that's feasible for us to prove one way or the other.

Edit: Also, if they opted out of training, then we didn't train on it.

  • > it's not something that's feasible for us to prove one way or the other.

    This kind of question is exactly what a company named _Open_AI and founded as a nonprofit is supposed to be doing; open research on AI that helps inform, rather than obscure.

    Anyhow, you do have the data available about the documents in the user's accounts, what they opted into (or were forced into via non-negotiable ToS), and whether they pressed a "thumbs up" button. You can answer whether the data entered the training pipeline or not. Yes, how much influence it had is an open question, and one that would be good to have research on and better tools for exploring, but I'll accept that it can't currently be answered precisely.

    But whether the data entered the trianing pipeline can be answered. And how to provide better tools for quantifying and tracing this kind of thing is exactly what should be studied.

  • I think it should be incredibly easy to verify this. Just look at the training data and see if it contains any of the chats. It should be trivial for a company with tens of thousands of super-genius agents at their disposal.

    • Two steps would be needed.

      (1) We'd have to identify their chats. How would we do this? We'd need them to share their chats with us so we could look for matches.

      (2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.

      #1 requires their cooperation and a bit of work on our side. #2 is extremely expensive and not really feasible.

      3 replies →

    • Just because something is in the training data, doesn't mean it is the root of an LLMs output.

      Turn off web search and ask a model what a random redditor said about a random topic in 2015. You will only get hallucinations at best, even though that comment is definitely in the training set.

      1 reply →

    • What they're saying, and I think this was the clear implication of the blog post too, is that the training data definitely would contain these chats and the only question is whether it got encoded into the weights.

  • Presumably, given that you also operate in the EU, you would have asked for their explicit consent before you did, so you could just check for that?

That’s also what I understand. If true yet another disgusting behavior from the company