← Back to context

Comment by lambda

18 hours ago

Why can't they rule it out? Is even OpenAI unable to track the provenance of all of their training data?

This is one of the major problems with these enormous closed models, and even most open-weights models, which don't disclose their training process or training data. You can never be sure what went into its training. Did it come up with an idea originally, or is it just plagiarising its training data? Are there malicious inputs being used to train in particular behaviors when given certain trigger phrases? What are the characteristics of the RLHF data and what kind of biases are those embedding in the models?

With proprietary closed models, or even open weights models that don't have open training datasets, you just can't answer these questions.

To truly prove some incidental usage data made no difference we'd have to (a) identify any of their de-identified data that came from their usage of ChatGPT, (b) train a bunch of expensive giant models, and (c) ask them all to solve the Navier-Stokes Millenium problem until hitting some level of statistical significance. It's just not feasible to run experiments like this to prove whether a piece of data has an effect on model behavior.

As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.

There's no reason to believe that anything they did in ChatGPT led to our solution; it's just impossible for us to truly prove it. And knowing most of the recipes we use, there's really no reason to think such contamination happened. I've asked the team to make a clearer, less-lawyerly statement here - let's see what happens.

(I work at OpenAI.)

  • So, one way to prove that the data played no part is to trace and show that it wasn't used in the training process at all. If the data was never used in training, then it couldn't have played a part in the training process.

    You're right; if the data was used in training, then it gets much trickier; it would be very difficult to show whether some particular data had a significant effect on the outcome.

    This is one of the big problems with giant models like these; it becomes nearly impossible to discern what is and isn't plagiarism, or copyright violation.

    It would in theory be possible to have things like n-gram databases or rolling hashes of training data, somewhat similar to OLMoTrace (https://arxiv.org/abs/2504.07096), which would allow for detecting whether particular documents ended up in the training data or not (you'd have to keep this for every model used in the whole training chain, as synthetic data generated by earlier models could be influenced by training data that wasn't included in later models). I'm sure there are practical issues with providing such a tool, but I think that it's necessary if you want to be able to categorically say "no, this document has never been present in the training data of this model."

    Or look at it the other way: if your model wasn't influenced by things in your training data, why include them in the first place? Clearly, you train on all of these documents because they influence the model. Yes, it's hard to trace the exact influence of each one. But if they're not affecting the output, then why not just stop training on them? You could just not train on any private documents; only train on public, traceable data.

    But instead, you choose to train on these private documents, so you have to admit, your model and its outputs are influenced by them.

  • > As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.

    That reads as incredibly dismissive and condescending. What makes you think you’re in a position to communicate like that when engaging on such a sensitive topic?

    • I intended no dismissiveness or condescension. My hope was to explain why it's hard to prove whether something affects model behavior. In the case of the moon, we have a strong prior belief that it makes no real difference. But it's hard to prove, because what if there's an unexpected impact from tides, cosmic rays, grid voltages, holiday traffic, etc. Models trained under slightly different conditions could have slightly different weights and behave slightly differently when solving math problems. Similarly, I have a strong expectation that, for example, a thumbs up signal from a ChatGPT chat will not meaningfully affect long-horizon mathematics work in our latest model, but it's always possible that it could. I think the plausibility of the ChatGPT route is higher than the tides, but still incredibly low. I respect Tristan and Levant a great deal and I'm bummed that this controversy has erupted (I acknowledge this will ring hollow if you think it's our fault). It reminds me a bit of the Frontier Math controversy, where people on the internet boldly claimed over and over again that we had trained on the Frontier Math evaluation set, even though we had not.

      2 replies →

  • I think there is a much easier way to prove that the ChatGPT usage of Tristan Buckmaster and Levent Alpöge (possibly also the ChatGPT usage of Córdoba and Martínez-Zoroa, if they use it) had no influence on OpenAI solving the Navier-Stokes problem.

    If the internal OpenAI model is as capable as you claim (being able to solve a Millenium problem without using unpublished insights built on years of work from mathematicians), then it should be able to demonstrate this capability again.

    How about OpenAI solves another Millenium problem within the next two weeks, that doesn't coincide with the parallel discovery/solution of other teams of mathematicians, using ChatGPT for preliminary proofs & write-ups.

  • So you definitely did train on their data, you just think it is unlikely that it impacted the final model significantly?

    • I have no idea if their data was trained on. For example, if they used ChatGPT, asked a math question, and clicked the thumbs up button, that could have provided a small reward signal. I highly doubt this sort of feedback made a difference to a problem like Navier-Stokes, but it's not something that's feasible for us to prove one way or the other.

      Edit: Also, if they opted out of training, then we didn't train on it.

      10 replies →

    • That’s also what I understand. If true yet another disgusting behavior from the company

  • > identify any of their de-identified data that came from their usage of ChatGPT

    "de-identified" seems more of a euphemism than normal in this context, given the very unique work they were doing.

    • I wouldn't expect poking at millennium problems to be that rare in ChatGPT. They were uniquely successful - but it's probably not easy to check de-identified data for the presence of any of their work on the problem because it would blend into a haystack of less successful work on the problem.

  • Thanks for the details, it's definitely believable, but if the user had not consented to have their conversations used for training, then shouldn't it be straightforward to state that their conversations were never used for training?

    If you need to do a whole series of extensive experiments to check in that scenario, it implies there are pathways for your conversations to end up in training even though you opted out of that setting.

    Of course, this is assuming that the toggle was set to not consent to training. I can't know that of course, but if this is considered a possibility even after using an enterprise account or toggling off data retention, it's a bit concerning.

  • (a) identify any of their de-identified data that came from their usage of ChatGPT.

    You don't need his login information, you just need to identify if anyone was approaching the NS problem using his method. Nobody else on earth (presumably) besides him, his team, and at best OpenAI were approaching the problem this way.

  • Please answer this question: do you or do you not train your models on anonymized user data, where those users have opted out of such training?

    The blog post appears to imply the answer to this is yes, as otherwise I assume it would be impossible for this contamination to have happened.

  • Why wouldn't contamination be possible? I can believe the data is de identified so you couldn't simply prompt the model to "follow this guy's approach", but it's entirely plausible that there is a very tiny amount of data about this approach in your dataset, and it comes precisely from this researcher.

  • "There's no reason to believe that anything they did in ChatGPT led to our solution"

    do you think that the model's proof was unrelated to being fed a solution that was close to completion?

    any comment on openai allegedly trying to drop attribution for alpöge and then threatening buckmaster?

  • It is, perhaps worth considering that the reputational community might not care about the difficulty for the AI builder to verify pedigree.

    If OpenAI's answer to this problem is "We can't know," then the rational conclusion may very well be "If I seek to have my reputation attached to the discovery of the solution, it is not sane to use the AI as an assistive tool, lest it scoop me on my own work using my own work. After all, they don't know it doesn't do that..."

  • That's such a shit parallel example that it borders on dishonest.

    There are hundreds of incredibly strong scientific priors that would have to be disproven for the moon to contribute to the solution.

    If a model was trained on this data, even if it was trained using methods that lead you to believe it unlikely to have learned details about the proof (e.g., maybe it was only used to train some kind of reward model, which played a minor role in the overall training and would thus be very unlikely to transfer details of a proof), you wouldn't have to disprove large swathes of known science to be wrong.

  • If the model has access to the "anonymized" data from chats, and the model is capable of building its own context from data that it can search through, including this data. Then it looks pretty damning. An independent review of the data traces from CoT and tool use involved in producing the result should make it clear one way or the other. Seems like discovery in a civil lawsuit could be very productive.

  • > As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.

    The _gall_ to say something like this. Do you perhaps think we are all stupid?? This very blogpost claims not to know if their work was used as input for this model. I don't even understand how that is possible, surely you can know if something is part of the training data, even if you are in the dark about what impact it actually made, qualitatively. The moon....

    > Knowing most of the recipes we use, there's really no reason to think such contamination happened.

    Yeah sorry but I don't trust you. I don't trust people or companies that have shown themselves to be dishonest before. Especially when the previous paragraph is comparing plagiarism and training data contamination with, _the phases of the moon_.

    Might even be you're actually telling the truth, but the boy that cried wolf and all that.

    -----

    As an aside, I would bet very good money at how most (all?) these companies are flouting their ZDR.

  • [flagged]

    • One of the wild things about how these models work is how often things that aren't sampled directly end up a variable in the model via secondary signal.

      They aren't keying queries by phase of the moon. But if, for example, more people talk about camping outdoors when the moon is full, and they're using conversation topic and timestamp as signal in what eventually becomes training data, it's not impossible the model has learned something about moon-phases.

      That's the kind of thing that's hard to prove had no impact on an answer.

A careful reading of "we cannot rule out that de-identified data derived from their usage of our products helped improve our models" could be saying that yes they trained on it but they don't know if that training data resulted in an "improvement" to the model. That is, they can't rule out that the only reason the model found this solution was because it had been trained on this approach.

The term ruled out is very open ended and gives them significant flexibility of meaning. They may have the information to determine exactly what happened, but they haven't looked so they can't "rule it out".

> Why can't they rule it out? Is even OpenAI unable to track the provenance of all of their training data?

Probably? I have a few hundred TB of training data for various small scale models and I can attest that I have _no idea_ what's in them. As in, literally zero. Half is scraped from GitHub and other hosting sites, other than that, I couldn't tell you anything else.

At OpenAI's scale their entire pipeline is likely 100% automated.

  • Yeah, I'm sure it's completely automated.

    But that doesn't preclude being able to index and track what the sources of data are. For your data sets, I would hope you are including source information for where the data came frome. And at OpenAI's scale, I would presume they are doing some amount of rolling hashing or similar to weed out duplication, training on too much duplicate data can cause problems.

    AllenAI have at least attempted to add some amount of traceability to their models with OLMoTrace (https://arxiv.org/abs/2504.07096), by letting you find n-gram matches from the outputs in their training data. It's not the most useful, there's a reason that LLMs use full fledged attention mechanisms and not just n-grams, a lot of times the n-gram matches it finds aren't all that related to the given output, it might be better to supplement this index with a vector search or other ways of keeping track of what training data would have most influenced particular parts of the output.

    But anyhow, this is something that is an important question, and the big labs should be working on to make their products more trustworthy. Instead, they are hiding information about how they train, hiding their reasoning traces, and just producing output with no information on what might have influenced the training.

    • Attributing training data seems pointless for trustworthiness. The way you trust a model is the same way you trust a human; you ask it to:

        1. Provide a chain of reasoning from agreed premises. These days LLMs can even do this airtight with proof assistants.
      
        2. Cite data sources for non-agreed premises. I don't care where the model learned a fact. It might not have ever read a document directly from the primary source. I want it to link directly to either widely agreed facts (e.g. standard textbooks, and if necessary school syllabi demonstrating that the text is standard) or primary sources (e.g. datasets). 
      

      Training provenance is irrelevant. It's neither necessary nor sufficient to deal with truth.

  • The question is not "does OpenAI know", it's "can OpenAI attest that the usage of their products for confidential data is not going to cause that sensitive data to become known to their models". And right now the answer I'm reading is that OpenAI can't attest to that.

  • Aye, but do they train on user data in these circumstances or not? If they do, then almost certainly the model was influenced by the input of the allegedly plagiarised material.

OAI could check whether those accounts enabled training data. If "yes", OAI could trace whether that data was used in any related training process. If either of those answers comes out to be "no", then that's sufficient to conclude training data independence.

We wouldn't need a full ablated re-training and solution attempt, contra tedsanders in a sibling comment.

  • > could trace whether that data was used

    The point of de-identifying data is to ensure you can't trace who it came from. It would be a serious privacy violation if they could.

    • If the model includes unique data from a person then that person can identify the data - the allegedly plagiarised material - and so re-identify it. There doesn't need to be a privacy breach to close that loop as it requires the person to identify the information is associated with them first.

Good chance their whole training pipeline is vibe coded so yah they probably don't actually know.

At the scale at which these models are now, regardless of whether they are proprietary or open weight or list their training datasets, there are hundreds of billions of works that have gone into trillions of parameters, each one providing tiny perturbations in some tiny fraction of the weights. It is probably impossible to attribute provenance to any specific input (which is also why the courts' finding of Fair Use is reasonable.)

Which is why, as I said in a recent comment (https://news.ycombinator.com/item?id=49530864) inadvertently leaking ideas to models is a grave risk for Intellectual Property.

> The risk with IP, however, is a lot more grave. You may not even need to memorize the details of the IP verbatim, just the broad idea may be enough. It may lurk encoded in the weights forever, just waiting to be activated by the right prompt to start a chain of thought that unlocks further details. Heck, it may even appear as if the model suggested the idea itself.

However, from a quick skim of the timelines, the specific discoveries, and all the he-said-she-said, so far it seems unlikely that OpenAI's model cribbed from the NYU / Anthropic pair, even if it would be impossible to prove.

Maybe what might help is a timeline of when the other two were using Codex for their work, whether they had opted out, and how long it takes for user data to make it to the training of their internal models. That last bit may be considered sensitive information however, as it could give away a lot about their internal processes.

  • There are two different things:

    - was item X in the training data

    - did the inclusion of X in the training data lead to Y

    I understand why the second is hard, but why is the first one hard?

    • Yep, the last part in my post was suggesting some ways we could determine if "item X was in the training data" (as well as some potential blockers for that from OpenAI's perspective.)