I work at OpenAI, though not on the team that did this, and my understanding is:
- we decided to ask our model for Millenium problem solutions because of two reasons: (a) our new model was looking incredibly good and (b) we heard rumors that some Millenium problems had been solved and were curious if our models could solve them (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these)
- we did not read any private chats (but of course the model was aware of prior research literature published to the internet)
- the proof generated by our model was very different from theirs and also goes far beyond the published literature
- we made an effort to jointly announce rather than immediately scoop (I understand Tristan was unhappy with the conversations; I know zero details here and I hope more is shared today)
"I was shown a prompt and told the internal research model had simply been
given the problem statement. Levent had been told by Sebastien “very little
human input” had been used. This turned out not to be true. Over the course
of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute
had been used."
- This, from Tristan Buckmaster's writeup yesterday, indicates to me that there was more than incidental inspiration from Alpoge and Buckmaster.
All of those statements sound true, based on what I've heard.
- "very little human" input feels ambiguous, and if someone spends a few days prompting a model to solve a super hairy problem requiring a 100-page proof, I can understand reasonable people interpreting that as both "very little" and "not very little" human input
- it's all true that a team worked on this, a bunch of compute was burned, and the problem was solved in stages and pieces
I'm not sure how any of this provides evidence that OpenAI took any of their work.
As evidence against, we never looked at any of their ChatGPT conversations and our model's proof is quite different from theirs.
(I work at OpenAI, but not on the team that did this proof.)
Your post says “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .” We can discuss what it means to “read” things but obviously the issue here isn't whether you did it manually or automatically.
But more importantly, what on earth are you doing threatening real scientists to remove their coauthors, then making fun of them on social media? Does the entire company run on that toxic culture, or did those people run off of some kind of outrageous tangent?
Except that’s not what happened. OpenAI offered to collaborate and put conditions on their offer. They aren’t threatening the removal of a coauthor for an independent work.
My employer would be rightfully outraged if I commented publicly on a sensitive, nuanced, and controversial issue like this based on my second-hand understanding of the matter.
> (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these)
That is your opinion, but the optics of that should raise for you some flags. OAI could have waited (how long is a task left to the ethics committee) to see how the rumors panned out. Right now the optics look a lot like "we don´t care there is a 1/7 chance we one-up a human researcher by reacting to this rumor immediately, might makes right"
The question I am interested in is not "did we read private chats", but "was this new model trained using any of Tristan and Levent's chats, regardless of whether they were marked private". Can you comment on that?
If they opted out of training, then we definitely did not train on them.
If they did not opt out, then I don't personally know if training signals came from their chats, and I don't think we'd be able to tell without their cooperation in identifying them. And even if signals were trained on in some manner, I highly doubt it made a difference to a problem as challenging as the NS proof.
Reasons for my doubt:
- I know most of our training recipes
- Our model's proof is very different from theirs
- The proof took a tremendous amount of tokens to derive (it wasn't a recall/lookup type question)
- This unreleased model has beastly performance on many unsolved math problems, not just the Euler solution
I acknowledge that this requires trust, and if you think we'd lie shamelessly about this stuff, then nothing we say can really help our case here.
Reminds me a bit of the Frontier Math fiasco, where people accused us of training on the eval set (we didn't), but it's hard to convince someone if they think you're lying.
If you're convinced we lie and cheat, then nothing I say may help. But if you're not sure, then hopefully providing my perspective is helpful.
He's a human, like everybody else. Mostly a bunch of hungry animals looking to put bread in our mouths. It's rarely ever something a bit more sophisticated than that.
Clearly The goal was to scoop Anthropic not a single researcher. OpenAI heard the rumor that Anthropic solved an open problem. So they went nuts pulling all plugs to scoop them.
Turns out it wasn’t actually Anthropic and just a researcher with a single Anthropic guy friend working on it .
I worked at OpenAI previously, but don't know any of the people involved in this.
My guess was it was probably this was more a nerd snipe than any action from OpenAI that was a "massive team" being put on it. Literally someone looking at this and asking "I wonder if our models are good enough yet".
It's easy to assume that having access to massive compute amounts means significant coordination, but this assumes that you're looking at the costs of this sort of thing from an external lens. Internally, tokens are often treated as free and infinite.
In fact, the entire outline of the proof is very similar to the external team's proof.
Details may be different, but the use of a very similar tack is very suspicious. Combined with thuggish comments "why would you ruin your career" and "I don't have to be nice" take away pretty much any credibility the OpenAI team's statements might have had.
AI companies seem much more relaxed than most about their employees posting on twitter/HN about this stuff. I'm not sure if it's about building hype or if it's about retaining talent. Probably both.
Prove it. Your systems hack and/or abuse other systems and you can't seem to even observe it happening much less do anything about it. Why should we believe your claims when they depend on an ability you don't actually have?
Your coworkers, after they learned about major progress in this problem, asked a model which was trained on the year of private work (the blog post even acknowledges this). No wonder it found the proof in less than a week using significantly higher compute resources. And if Tristan's accusations are true, that was absolutely intentional on the part of OpenAI. You are an evil company with evil people.
> - the proof generated by our model was very different from theirs and also goes far beyond the published literature
I'm hearing two completely conflicting stories. Buckmaster is claiming the approach used by OpenAI is so strikingly similar to the one he used, that mere coincidence is astronomically small. Yet OpenAI is claiming that the methods used are entirely different.
Anyone care to provide primary evidence proving one way or the other?
My OpenAI account was deactivated on Sunday due to a claimed infraction of production of child materials, maybe based on a few words in a technical chat that clearly isn't about that. Can you take a look? rviragh@gmail.com - I was doing a lot of important work and projects and sharing much of my work with OpenAI. I also am a big proponent of funding Social Security Trust Funds (OASI & DI Solvency) so reactivating my account would let me do that as well. Thank you for taking a look.
> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
I asked whether the model had been trained on, or had access to, our sessions
in Codex, into which we had been putting all our drafts for the whole of this
project. I was told the model did not look up user data. I asked again, about
training, and I did not get an answer.
This whole episode is more horrific than "AI is eating math". We now have a clear and economically damaging (or at least career damaging) example of the "training on customer tokens" problem.
It’s fair to give benefit of doubt to Buckmaster given OpenAI is currently being very credibly sued by Apple for openly stealing others’ original work in another context.
It's not a meaningful response to the accusations. Any productive new research direction would be expected to lead to a number of different possible proofs of a number of similar problems. (Given their bizarrely compressed timescale here, it's possible that the proofs really are so different it's clear they came independently, and they just didn't have time to come up with that information before hitting publish.)
Yes, that was the allegation last night.
I work at OpenAI, though not on the team that did this, and my understanding is:
- we decided to ask our model for Millenium problem solutions because of two reasons: (a) our new model was looking incredibly good and (b) we heard rumors that some Millenium problems had been solved and were curious if our models could solve them (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these)
- we did not read any private chats (but of course the model was aware of prior research literature published to the internet)
- the proof generated by our model was very different from theirs and also goes far beyond the published literature
- we made an effort to jointly announce rather than immediately scoop (I understand Tristan was unhappy with the conversations; I know zero details here and I hope more is shared today)
Edit: Here's is Sebastian's take: https://x.com/SebastienBubeck/status/2097379411691516310?s=2...
"I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used."
- This, from Tristan Buckmaster's writeup yesterday, indicates to me that there was more than incidental inspiration from Alpoge and Buckmaster.
All of those statements sound true, based on what I've heard.
- "very little human" input feels ambiguous, and if someone spends a few days prompting a model to solve a super hairy problem requiring a 100-page proof, I can understand reasonable people interpreting that as both "very little" and "not very little" human input
- it's all true that a team worked on this, a bunch of compute was burned, and the problem was solved in stages and pieces
I'm not sure how any of this provides evidence that OpenAI took any of their work.
As evidence against, we never looked at any of their ChatGPT conversations and our model's proof is quite different from theirs.
(I work at OpenAI, but not on the team that did this proof.)
55 replies →
> we did not read any private chats
Your post says “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .” We can discuss what it means to “read” things but obviously the issue here isn't whether you did it manually or automatically.
But more importantly, what on earth are you doing threatening real scientists to remove their coauthors, then making fun of them on social media? Does the entire company run on that toxic culture, or did those people run off of some kind of outrageous tangent?
If the training toggle is switched on, maybe OpenAI doesn't consider a chat to be private? Therefore making this a 'safe' statement.
1 reply →
Except that’s not what happened. OpenAI offered to collaborate and put conditions on their offer. They aren’t threatening the removal of a coauthor for an independent work.
1 reply →
My employer would be rightfully outraged if I commented publicly on a sensitive, nuanced, and controversial issue like this based on my second-hand understanding of the matter.
I came here to say this, like... wow. I'm pretty sure at at least a few of the places I've worked that would be grounds for immediate termination.
4 replies →
> (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these)
That is your opinion, but the optics of that should raise for you some flags. OAI could have waited (how long is a task left to the ethics committee) to see how the rumors panned out. Right now the optics look a lot like "we don´t care there is a 1/7 chance we one-up a human researcher by reacting to this rumor immediately, might makes right"
>we did not read any private chats
The question I am interested in is not "did we read private chats", but "was this new model trained using any of Tristan and Levent's chats, regardless of whether they were marked private". Can you comment on that?
If they opted out of training, then we definitely did not train on them.
If they did not opt out, then I don't personally know if training signals came from their chats, and I don't think we'd be able to tell without their cooperation in identifying them. And even if signals were trained on in some manner, I highly doubt it made a difference to a problem as challenging as the NS proof.
Reasons for my doubt:
- I know most of our training recipes
- Our model's proof is very different from theirs
- The proof took a tremendous amount of tokens to derive (it wasn't a recall/lookup type question)
- This unreleased model has beastly performance on many unsolved math problems, not just the Euler solution
I acknowledge that this requires trust, and if you think we'd lie shamelessly about this stuff, then nothing we say can really help our case here.
Reminds me a bit of the Frontier Math fiasco, where people accused us of training on the eval set (we didn't), but it's hard to convince someone if they think you're lying.
If you're convinced we lie and cheat, then nothing I say may help. But if you're not sure, then hopefully providing my perspective is helpful.
16 replies →
>Can you comment on that?
No answer is also an answer.
He's a human, like everybody else. Mostly a bunch of hungry animals looking to put bread in our mouths. It's rarely ever something a bit more sophisticated than that.
1 reply →
If the goal was not to scoop them, why did openai put a massive team on this, working weekends, only after they heard rumors of the solution?
Clearly The goal was to scoop Anthropic not a single researcher. OpenAI heard the rumor that Anthropic solved an open problem. So they went nuts pulling all plugs to scoop them.
Turns out it wasn’t actually Anthropic and just a researcher with a single Anthropic guy friend working on it .
Wild times
1 reply →
I worked at OpenAI previously, but don't know any of the people involved in this.
My guess was it was probably this was more a nerd snipe than any action from OpenAI that was a "massive team" being put on it. Literally someone looking at this and asking "I wonder if our models are good enough yet".
It's easy to assume that having access to massive compute amounts means significant coordination, but this assumes that you're looking at the costs of this sort of thing from an external lens. Internally, tokens are often treated as free and infinite.
2 replies →
You’re straddling a weird line here where I am not sure if you are speaking on behalf of OpenAI or not.
In fact, the entire outline of the proof is very similar to the external team's proof.
Details may be different, but the use of a very similar tack is very suspicious. Combined with thuggish comments "why would you ruin your career" and "I don't have to be nice" take away pretty much any credibility the OpenAI team's statements might have had.
I have no idea why Sebastian would offer the individuals attribution if OpenAi didn't somewhat knowingly scoop them
The fact that you're even here commenting on this is... a choice
AI companies seem much more relaxed than most about their employees posting on twitter/HN about this stuff. I'm not sure if it's about building hype or if it's about retaining talent. Probably both.
Prove it. Your systems hack and/or abuse other systems and you can't seem to even observe it happening much less do anything about it. Why should we believe your claims when they depend on an ability you don't actually have?
How are people talking about this there? Why are so many employees posting nasty things about Tristan on twitter?
Can you point me to any nasty things being posted? I'll ask them to delete.
5 replies →
Your coworkers, after they learned about major progress in this problem, asked a model which was trained on the year of private work (the blog post even acknowledges this). No wonder it found the proof in less than a week using significantly higher compute resources. And if Tristan's accusations are true, that was absolutely intentional on the part of OpenAI. You are an evil company with evil people.
Some millennium problems? Are there more coming?
> - the proof generated by our model was very different from theirs and also goes far beyond the published literature
I'm hearing two completely conflicting stories. Buckmaster is claiming the approach used by OpenAI is so strikingly similar to the one he used, that mere coincidence is astronomically small. Yet OpenAI is claiming that the methods used are entirely different.
Anyone care to provide primary evidence proving one way or the other?
Buckmaster came up with an approach.
OpenAI's approach was to copy his work, which is technically a different method of coming up with an approach.
My OpenAI account was deactivated on Sunday due to a claimed infraction of production of child materials, maybe based on a few words in a technical chat that clearly isn't about that. Can you take a look? rviragh@gmail.com - I was doing a lot of important work and projects and sharing much of my work with OpenAI. I also am a big proponent of funding Social Security Trust Funds (OASI & DI Solvency) so reactivating my account would let me do that as well. Thank you for taking a look.
[dead]
Its funny, it is uniquely with this one act that I have turned forever on OpenAI, which I hitherto defended up and down against nonsense charges.
I dedicate my life to its complete destruction beginning today.
Wow, the origin of a supervillain! /s
That is addressed in the article.
OpenAI's position:
> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
Why is it unlikely?
8 replies →
I thought openai don't use any user data if we opt out of training and via api?
3 replies →
> https://cims.nyu.edu/~tristanb/statement.pdf
This really need to be a top-level story on HN..
I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
This whole episode is more horrific than "AI is eating math". We now have a clear and economically damaging (or at least career damaging) example of the "training on customer tokens" problem.
We can't ignore this problem any longer.
We don't have any proof of that at all. Please stop rushing to judge without data.
> We don't have any proof of that
It’s fair to give benefit of doubt to Buckmaster given OpenAI is currently being very credibly sued by Apple for openly stealing others’ original work in another context.
Did you actually read the article and the substance of the solution?
>our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)
It's not a meaningful response to the accusations. Any productive new research direction would be expected to lead to a number of different possible proofs of a number of similar problems. (Given their bizarrely compressed timescale here, it's possible that the proofs really are so different it's clear they came independently, and they just didn't have time to come up with that information before hitting publish.)
[flagged]