← Back to context

Comment by mayakacz

7 hours ago

I'm not one to comment often but this really pisses me off.

OpenAI looked at user data, stole world class researchers' work, and then tried to threaten those researchers to do what would make their corporation profit (which they would anyways!).

Imagine you have been working on a terribly difficult math problem for a decade. This is a result you have spent years on, and what you will likely be remembered for. And to have some punk from OpenAI lie to you, threaten you, and tell you that they are willing to go on the record that you "deserved" it? What is this, the Godfather?

If OpenAI solved Navier-Stokes, that is an astounding result! - yet they'll still be remembered as those who thought credit was more important than results. That winning was more important than collaboration. If this is true, they're burning any trust left with academia.

I'm stunned that people are taking this accusation as a fact.

OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around on some kitchen table somewhere and 2. I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business.

There are things that Buckmaster alleged and things that he speculated. The entire training data thing is speculation. If this is pissing you off, then you ought to evaluate how you ingest information.

  • He didn't even make that accusation!

    > I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.

    The shocking/interesting thing would be if it was trained on the sessions. I think it's very implausible that they gave the model access to someone else's sessions as input. That would be a huge privacy violation and would probably blow up a large proportion of their enterprise business.

    Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.

    • Parse that statement more carefully.

      > I was told the model did not look up user data.

      The naive way to read this is "Nothing you guys did influenced the way our model got to the solution".

      The less naive way to read this is "Of course the model isn't looking up your user data. I (the guy trying to blackmail you to remove the Anthropic employee from credit on your paper) looked up your sessions, and tipped our model off on how to solve this problem".

      1 reply →

    • >> Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.

      How "fast" does it have to be? Buckmaster and Alpoge have been working on this for just a day short of a year. See Alpoge's tweet announcing his collaboration with Bukmaster dated 9/19/25:

      https://x.com/__alpoge__/status/2097206973418611054

      It takes a few months to train a model these days but not a whole year. OpenAI had all the time to train on Buckmaster and Alpoge's results of just a few months earlier at which point they must have been well on the path to their result.

    • It would be shocking if it wasn’t trained on sessions. Have you read the ToS parts for both openai and anthropic that talk about it? It’s so obviously a weaselly way to say "no we do not train on your exact chats but we talked with legal and we think a cleanroom reimagining of your convo is probably fine and frankly where else are we going to get such a treasure trove of training data?"

      There’s potentially trillions on the line, do you seriously expect those companies to adhere to laws and regulations any more than, say, uber?

      The only unlikely part is the timeline - your sessions from a week ago probably haven’t made their way into the model. It’ll just take a while longer, and will be massaged just enough so that it isn’t really your exact session word for word so you can’t sure as easily.

      1 reply →

    • It wouldn't be shocking at all. They stole human data to train the first models and they've been stealing it ever since to train new models. Stealing mathematicians private chats and private research and taking credit for it would absolutely be par for the course.

      Enterprises are well aware of it and are fully on board. You didn't think every corporation in America has an OpenAI subscription because the models were good, did you?

      The whole reason they have subs is to train them on YOUR WORKFLOWS lol

      3 replies →

  • Whenever I see comments defending AI companies, I look at the account's creation date, and interestingly almost all of them were created post 2024.

  • This is how internet discourse works on Reddit/Twitter/HN and the rest. Someone said something which confirms your biases so it’ll now be treated as a fact and repeated endlessly in the echo chamber.

  • He asked whether they used their chats as training data and received no response. Any speculation here seems quite appropriate?

  • I don’t think you understand how brazen big tech companies are in practice.

> OpenAI looked at user data, stole world class researchers' work

This doesn't seem to be clear and is very implausible for a large company. Be as cynical as you want, but a normal researcher will simply not have access rights to this data, which will be siloed away somewhere else.

It might very well be somewhat unfair to catch wind of a promising approach and then try to frontrun them by throwing compute at the problem, but this isn't really the same.

  • No, plausible given AI companies want/need session data to train their next models. Probably not someone peeking an eye to sessions directly, but probably not so hard to find the useful sessions in anonymized training data to post train a model on. As stated in the paper, OpenAI did not explicitely denied the researcher sessions were not used for training the model. So either they don't know, or don't want to tell

    "I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer."

    Let's see what statement OpenAI will come up with for their side of the story

    EDIT: precised my thought on user data vs session data

This conclusion is flawed. It's unclear at this point if OpenAI's model or employees actually looked at or stole the author's data. Having worked at large companies before, I'm leaning towards no, since very few employees have access to that data.

And simply knowing a problem can be solved is half the battle.

  • From Buckmaster's text:

        The route to the Clay problem through a
        smooth force, options c and d in Fefferman’s statement of the problem, is the
        route Luis and Diego opened and the one Levent and I had quietly chosen to
        attack. Almost nobody else I know of was working on it. It is not the direction
        one arrives at in a few days by giving a model the problem statement. When I
        heard “forced,” it was a bright red flag.
    

    This is much more than the knowledge than the problem can be solved, it's also the specific, non-obvious approach to solving it. That's much more damning for OpenAI, if confirmed.

    • That's a stretch. The Luis and Diego paper was published in 2023 and is included in every frontier model's training dataset. An AI model could independently choose the same path route as Luis and Diego, without access to Buckmaster and Alpöge’s work.

      And the article states "an insane amount of compute had been used," which implies OpenAI brute-forced their way to a solution. I.e. they searched for every paper published on Navier-Stokes and exhaustively attempted every approach. Such an approach would lead them to a solution.

      There is not enough information at this time to reach a conclusion. The best option is to wait for statements from both sides, then reevaluate.

      5 replies →

    • I want to point out that almost all previous AI discoveries in math were made in almost the same way. The ideas were there in the community, but weren't considered mainstream/worth pushing forward. Read Tao's comments on the unit distance problem, for example (sry I can't find a link right now).

      OpenAI said there [1]: > The method by which the problem was solved is also notable. The proof brings unexpected, sophisticated ideas from algebraic number theory to bear on an elementary geometric question.

      [1] https://openai.com/index/model-disproves-discrete-geometry-c...

  • > And simply knowing a problem can be solved is half the battle.

    Have you done any mathematical research? If not, then no, knowing that a problem is solvable is not “half the battle”.

    Homework problems are all designed to be solvable, yet they can vary greatly in difficulty. Research mathematics is even more extreme, because, unlike with homework, you don’t know that it is solvable with the extant mathematics, and you might need to invent new maths.

    • You're taking the phrase too literally. The point is that knowing a solution is possible gives you the conviction to actually find that solution. The hardest part of solving a problem is often a lack of conviction to see it through, and quitting too early. Once you know a solution exists, you can commit maximal effort towards solving it and know that your efforts are not in vain.

      If not for the rumors that A/ had already solved NS, OAI would likely never have pursued solving the problem with such fervor. The rumors drove OAI to assemble an entire team to crack this.

  • How is it unclear? The entire point of deploying models across corporate America is to train on your workflows. Eventually replacing you with digital you is why they're doing it!

Things are more entangled than that. The contribute made from both OpenAI and Anthropic models to solve these problems are clear, now it really hard to quantify which one contributed more, if the role played by the human is major or minor.

OpenAI tried to collaborate and share the results together with a fixed timeline, to avoid this mess but it was inevitable. There is a conflict of interest, where the other researcher works at Anthropic, who will also try to take credit.

Where they may be in the wrong is if they took user data regarding the problem, how will we know if they did or not?

  • They offered to collaborate by asking to drop a coauthor.

    That is not collaboration, and is not an academic norm.

Yes, but only if you take this one sided statement at face value.

  • Why would have they rushed the publication if this was not true? Are you also suggesting that he fully invented the call with Open AI?

    • The results being true, the 'deal' that was made being true doesn't mean some of the implied accusations here are true, for example - that Open AI used their Codex logs to drive their breakthrough.

      11 replies →

  • I’m open to evidence, but just using Bayesian reasoning, OpenAI is one of the most dishonest companies in history. They’re currently being sued for a dozen employees stealing Apple hardware! I don’t understand why I should give them any grace.