My take home from this entire drama is that one should not use LLM services for confidential or proprietary information as they all seem to be run by assholes. And you’re sending them everything you are doing. Would you send your lab notebook to an asshole? Hell no.
I say that as a mathematician (on paper) who perhaps surprisingly doesn’t give a crap about the problem itself.
Their privacy policy for normie subscribers says in plain English they use your Personal Data for research. I think it’s pretty unreasonable to use the service and expect otherwise.
My doctor's privacy policy is a bit more abusive than OpenAI's. It exists mostly because of a $%^&&* legal framework rather than malice, and I've grown accustomed to "if I don't want to die then I sign away these rights." Despite my having theoretically signed my soul away, my doctor isn't selling personal information to my exes or to life insurance companies (though they could in the US; that extremely personal information is no longer mine). OpenAI is engaging in the "technically legal maybe we'll see but obviously unintended" side of this transaction, and maybe that works out for them, but I wouldn't personally choose to be a shill for "it's unreasonble to expect somebody with 'legal' permission to do something other than the maximum 'legally' permitted" if I were in your shoes.
> My two favourite hypothetical questions regarding this used to be:
> If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe how.)
> If I brainstorm with ChatGPT about potential new directions for my company, what's the chance that information might be exposed to a competitor in six months' time who asks "what might company X plan to do next"?
> My new preferred hypothetical for this is:
> If I use ChatGPT to help me partially solve a Millennium Prize problem, what are the chances that my work will influence training such that a later model helps someone else solve it first?
>While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
What do you mean, as OpenAI employee, you cannot tell that his work has entered the training data ?
But also correct me if I'm wrong, if the two mathematician were really close to finish this problem, and their conversation were used by OpenAI,
shouldn't the Agent have succeeded way faster/efficiently instead of using "4.9 million messages and used about 300 billion output tokens."
> ... we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors ...
I've observed this exact effect last week. I made a discovery regarding a stepwise performance improvement in a codebase. I shared the benchmark results with a peer and within 12 hours they replicated the same. We had both been looking for this for years.
I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days. Competition is a hell of a drug, and frontier LLMs aggressively compound that energy.
> I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days.
If you read the history of major scientific discoveries, this has been the case for a long time. There are many things that were independently discovered by different people at nearly the same time. Once people know something is solved or solvable, it gets a relentless amount of focus.
This seems to be the norm rather than the exception.
On a tangent, the genius of people like eg Einstein is not so much that he came up with all these things: other people were close, but that he was a singular individual that did all of these discoveries, instead of five different guys all making some breakthrough here or there.
What's missing from the story I think is the part about "...then had a breakthrough on August 15th. The mathematical rumour mill kicked into gear...". If only the two of them were working on the problem in secret, how did their breakthrough become a rumor?
People talk. If you have a breakthrough solving one of the most famous problems outstanding, you're going to tell people.
You'll say, "Don't tell anyone", which they will ignore because they get a rush and perceived status by sharing it. So then they tell someone, along with "Don't tell anyone", etc.
It's a small enough world (both in academic math, one at Anthropic) that you get to OAI in very few hops.
The reality is also that most of your colleagues have no interest in stealing your work: they have their own work to do anyway, and having a colleague effervescing about whatever they are working on is kind of the norm in pure research. Just because they are making progress it doesn't mean they are about to do anything interesting.
So in general it's pretty safe to talk generally about whatever you're doing.
LLM’s seem very good at solving mathematical problems of which there is an enormous amount of exisiting work/attempts in their training data. This is an amazing capability, but does not convince me that these models are «thinking» or «reasoning» in the way a human does. A human mathematician could in theory categorize/discover an entirely new field of mathematics tomorrow, based purely on their «human intelligence», I wonder if we will see similar examples by LLM’s soon. It seems to me currently impossible that LLM’s can replace human mathematicians, because of their (assumption) likely dependence on human input in the sense of enormous amounts of pre-existing attempts/data.
If an entirely new problem, within a new field of mathematics were to appear tomorrow, I highly doubt an LLM would be useful at all on their own. Is this the «ultimate ASI test»?
Extreme temperature levels (>2.0) can push a GPT of its manifold, essentially producing predictions barely distinguishable from random noise (it flattens the probability distribution of the next token). In theory this could predict anything including the next field of mathematics (infinite monkey theorem) but realistically that would never happen.
However, how to we know the next field of mathematics isn't a novel combinations of several other sub-fields? That level of mathematics would be indistinguishable from magic to most people and so in their eyes the GPT did something truly inventive.
It’s reasonable to wonder about what chat usage data gets into models (to be honest probably quite little - carefully curating training data and creating higher quality synth data seems to be the current approach) and the implied risk to privacy and creativity (every new patent filed this year probably touched a model before filing).
What I cannot reconcile is the timeline and the concern in this specific case.
I don’t think training pipelines are anything close to the level of continuous training needed to incorporate Aug 15th ideas into a model that generates a breakthrough early Sept. Either OpenAI nakedly had someone with mathematical understanding dig into a specific user’s chats (a massive red flag) or this really is poor handling of a more classic parallel discovery situation (with one party clearly having worked on it longer)
OpenAI can easily identify these outstanding human behind their accounts. Human in OpenAI constantly check their logs for breakthrough. When they find something interesting, they brute force the result using their massive computing power.
I run a small SaaS[1], like so many others, that uses AI to generate and optimize SQL. Getting this to perform optimally has been a lot of work and now I wonder if OpenAI is outright stealing this knowledge, which without a doubt is highly valuable to them.
SQL is so ubiquitous and the use case so obvious, there's no way they have not already been tracking performance and benchmaxxing on SQL queries for years.
I think this drama was blown up a bit out of proportion. The entire discourse I am seeing online seems to revolve around this:
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
I mean... yeah? What do you expect? What else can they say? How could you prove a negative in this case? I do not want to comment on specific OAI employee chat messages, but on the actual OAI discovery here.
A very simple "these two pipelines don't connect up in our architecture, here's our internal high level network diagram that we will stand by in court" as opposed to "yeah, we don't even entirely know how our own customer facing systems are connected to our training pipeline, but it probably didn't happen".
Yeah, this was pretty shitty by OpenAI. Not surprising, sadly.
Them being assholes, trying to exclude an author just because he worked at Anthropic, shows the kind of culture within (that part of) their organization. The focus wasn't on supporting academics or expanding research. It was on getting great marketing.
If they had to burn millions of dollars solving a problem _that they thought was already being solved_ to do so, they'd do it.
The concept that "knowing something has been done" allows others to find the solution to an previously unsolvable problem is an old proven one. For example when Germany launched a rocket (V2 prototype) the British knew its rough trajectory and from spying it's rough size. Although they had previously believed that ballistic missiles weren't possible because no engine could provide the necessary thrust to weight ratio, given the obvious German launching of one, they went through all the known chemical compounds to arrive at the combination (Ethanol and LOX) used. [Story from RV Jones "Most Secret War"].
I’ll go back to the point about authorship. I’m Not a mathematician but I am in academia. if you are fucking around with authorship you are immediately suspect.
That aspect alone would/should be unthinkable to any serious academic. Authorship reflects who did the work and changing it for business competition reasons should be a red flag for multiple different reasons. They include, the sheer tactlessness of treating a major theoretical advancement as a competitive posturing first, the norms of academia second, and all the misunderstandings of the culture of the disciplines culture that people will now suspect are hiding beneath the visible surface (insert topography joke).
Math as a field is fairly unique even in how they list authorship. It was long the norm that authorship to be alphabetical because the idea of first, second, senior etc authorship is harder to define than many other fields.
“The stated rationale for alphabetical order is that it treats co-authorship as intellectually joint work: every listed author’s name carries equal weight, and no one has to negotiate, or be seen to negotiate, over billing. That is a genuine advantage over position-coded conventions, where disputes over who is “first author” are one of the most common sources of authorship conflict in fields that use them” [0]
That norm is changing, slowly, but one option people are pursuing is notable: randomized author order. Their is a perception that alphabetical is too biased…that’s the world OpenAI is stepping into when they make that offer of authorship to one scholar with a demand that he exclude his partner.
I can’t speak to the facts of anything else in this, but if a grad student came to me and said someone made them that offer, I would tell them to run and if they were brave report it.
OAI's side of the story is that they discovered the approaches (and indeed solved problems - Euler equations vs NS equations) differed. They then offered Buckmaster lead authorship of OAI's proof, without Alpöge. But they never demanded that Alpöge be stripped of coauthorship on resolving the regularity of the Euler equations. At least that's the claim.
My take home from this entire drama is that one should not use LLM services for confidential or proprietary information as they all seem to be run by assholes. And you’re sending them everything you are doing. Would you send your lab notebook to an asshole? Hell no.
I say that as a mathematician (on paper) who perhaps surprisingly doesn’t give a crap about the problem itself.
Their privacy policy for normie subscribers says in plain English they use your Personal Data for research. I think it’s pretty unreasonable to use the service and expect otherwise.
My doctor's privacy policy is a bit more abusive than OpenAI's. It exists mostly because of a $%^&&* legal framework rather than malice, and I've grown accustomed to "if I don't want to die then I sign away these rights." Despite my having theoretically signed my soul away, my doctor isn't selling personal information to my exes or to life insurance companies (though they could in the US; that extremely personal information is no longer mine). OpenAI is engaging in the "technically legal maybe we'll see but obviously unintended" side of this transaction, and maybe that works out for them, but I wouldn't personally choose to be a shill for "it's unreasonble to expect somebody with 'legal' permission to do something other than the maximum 'legally' permitted" if I were in your shoes.
>implying most users read them
If that is the case, why on earth would you use it in any professional setting?
2 replies →
You'd think theft would still be illegal regardless of what a privacy policy says.
1 reply →
> My two favourite hypothetical questions regarding this used to be:
> If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe how.)
> If I brainstorm with ChatGPT about potential new directions for my company, what's the chance that information might be exposed to a competitor in six months' time who asks "what might company X plan to do next"?
> My new preferred hypothetical for this is:
> If I use ChatGPT to help me partially solve a Millennium Prize problem, what are the chances that my work will influence training such that a later model helps someone else solve it first?
LLM can't be trained that easily. More like actual human are checking your logs and stealing valuable things from you.
Maybe just a rumor of a high value target having their API keys accidentally consumed in the context...
>While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
What do you mean, as OpenAI employee, you cannot tell that his work has entered the training data ?
But also correct me if I'm wrong, if the two mathematician were really close to finish this problem, and their conversation were used by OpenAI, shouldn't the Agent have succeeded way faster/efficiently instead of using "4.9 million messages and used about 300 billion output tokens."
It's almost as if it was actually found by manually written brute-force algorithm running on OpenAI's massive computing power.
> ... we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors ...
I've observed this exact effect last week. I made a discovery regarding a stepwise performance improvement in a codebase. I shared the benchmark results with a peer and within 12 hours they replicated the same. We had both been looking for this for years.
I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days. Competition is a hell of a drug, and frontier LLMs aggressively compound that energy.
> I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days.
If you read the history of major scientific discoveries, this has been the case for a long time. There are many things that were independently discovered by different people at nearly the same time. Once people know something is solved or solvable, it gets a relentless amount of focus.
This seems to be the norm rather than the exception.
On a tangent, the genius of people like eg Einstein is not so much that he came up with all these things: other people were close, but that he was a singular individual that did all of these discoveries, instead of five different guys all making some breakthrough here or there.
[delayed]
It's something that happened before LLMs - multiple discovery. Calculus is a classic example.
If they wanted to solve a Millenium prize problem so much, why did they not try to solve P-NP instead? It boggles the mind.
What's missing from the story I think is the part about "...then had a breakthrough on August 15th. The mathematical rumour mill kicked into gear...". If only the two of them were working on the problem in secret, how did their breakthrough become a rumor?
People talk. If you have a breakthrough solving one of the most famous problems outstanding, you're going to tell people.
You'll say, "Don't tell anyone", which they will ignore because they get a rush and perceived status by sharing it. So then they tell someone, along with "Don't tell anyone", etc.
It's a small enough world (both in academic math, one at Anthropic) that you get to OAI in very few hops.
The reality is also that most of your colleagues have no interest in stealing your work: they have their own work to do anyway, and having a colleague effervescing about whatever they are working on is kind of the norm in pure research. Just because they are making progress it doesn't mean they are about to do anything interesting.
So in general it's pretty safe to talk generally about whatever you're doing.
They weren't working in secret?
LLM’s seem very good at solving mathematical problems of which there is an enormous amount of exisiting work/attempts in their training data. This is an amazing capability, but does not convince me that these models are «thinking» or «reasoning» in the way a human does. A human mathematician could in theory categorize/discover an entirely new field of mathematics tomorrow, based purely on their «human intelligence», I wonder if we will see similar examples by LLM’s soon. It seems to me currently impossible that LLM’s can replace human mathematicians, because of their (assumption) likely dependence on human input in the sense of enormous amounts of pre-existing attempts/data.
If an entirely new problem, within a new field of mathematics were to appear tomorrow, I highly doubt an LLM would be useful at all on their own. Is this the «ultimate ASI test»?
Extreme temperature levels (>2.0) can push a GPT of its manifold, essentially producing predictions barely distinguishable from random noise (it flattens the probability distribution of the next token). In theory this could predict anything including the next field of mathematics (infinite monkey theorem) but realistically that would never happen.
However, how to we know the next field of mathematics isn't a novel combinations of several other sub-fields? That level of mathematics would be indistinguishable from magic to most people and so in their eyes the GPT did something truly inventive.
Time to repeat the same angry mob style discussion again. Great job to the mods.
It’s reasonable to wonder about what chat usage data gets into models (to be honest probably quite little - carefully curating training data and creating higher quality synth data seems to be the current approach) and the implied risk to privacy and creativity (every new patent filed this year probably touched a model before filing).
What I cannot reconcile is the timeline and the concern in this specific case.
I don’t think training pipelines are anything close to the level of continuous training needed to incorporate Aug 15th ideas into a model that generates a breakthrough early Sept. Either OpenAI nakedly had someone with mathematical understanding dig into a specific user’s chats (a massive red flag) or this really is poor handling of a more classic parallel discovery situation (with one party clearly having worked on it longer)
I would assume a lot of codex data goes back into training. A well steered session is extremely valuable data.
LLMs can't contribute good code to some of the good OSS math libraries, How is it even solving these problems?
That's a good question.
It is able to contribute code, but maybe not good code.
It's the same in math: it's able to solve problems, but not necessarily in a good way with a human readable code.
Math papers are a lot like software:
- theorems are like API
- lemmata like internal/private function API
- definitions are like types
- the proofs are the implementation
The proofs of ChatGPT are not necessarily readable or maintainable.
search, verifiability and compute
Maybe someone on the NYU team forgot to opt out of “improve the model for everyone”.
Basically no new info here, not really sure why this post needed to be written tbh.
On the contrary, a level-headed summary that gathers information from all the different sources is necessary.
Sounds like a great use case for an LLM
Agreed. This person just loves to shill their blog and their pelican benchmark.
There’s no shilling here. The person who posted the link isn’t the person who wrote the blog.
1 reply →
Yes, you have ability to get information very quickly, but not everyone does.
It could be much worse.
OpenAI can easily identify these outstanding human behind their accounts. Human in OpenAI constantly check their logs for breakthrough. When they find something interesting, they brute force the result using their massive computing power.
No LLM is even needed.
I run a small SaaS[1], like so many others, that uses AI to generate and optimize SQL. Getting this to perform optimally has been a lot of work and now I wonder if OpenAI is outright stealing this knowledge, which without a doubt is highly valuable to them.
[1]: https://www.sqlai.ai
SQL is so ubiquitous and the use case so obvious, there's no way they have not already been tracking performance and benchmaxxing on SQL queries for years.
Can’t wait to see the human verifying the results and then figure out that the AI model actually cheated and the results are not correct.
the proofs were verified in Lean, so unlikely.
As long as the proofs itself are correct. How long they were this time?
Edit: at least ~600,000 lines
https://stanfordtechreview.com/articles/openai-buckmaster-na...
I think this drama was blown up a bit out of proportion. The entire discourse I am seeing online seems to revolve around this:
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
I mean... yeah? What do you expect? What else can they say? How could you prove a negative in this case? I do not want to comment on specific OAI employee chat messages, but on the actual OAI discovery here.
A very simple "these two pipelines don't connect up in our architecture, here's our internal high level network diagram that we will stand by in court" as opposed to "yeah, we don't even entirely know how our own customer facing systems are connected to our training pipeline, but it probably didn't happen".
If they have zero-retention, then it is not possible.
So what they are saying is that they don't have zero retention.
But why is this news? This is clearly described in their ToS and in the Settings to improve their models.
Yeah, this was pretty shitty by OpenAI. Not surprising, sadly.
Them being assholes, trying to exclude an author just because he worked at Anthropic, shows the kind of culture within (that part of) their organization. The focus wasn't on supporting academics or expanding research. It was on getting great marketing.
If they had to burn millions of dollars solving a problem _that they thought was already being solved_ to do so, they'd do it.
All your datum are belong to us!
The concept that "knowing something has been done" allows others to find the solution to an previously unsolvable problem is an old proven one. For example when Germany launched a rocket (V2 prototype) the British knew its rough trajectory and from spying it's rough size. Although they had previously believed that ballistic missiles weren't possible because no engine could provide the necessary thrust to weight ratio, given the obvious German launching of one, they went through all the known chemical compounds to arrive at the combination (Ethanol and LOX) used. [Story from RV Jones "Most Secret War"].
Is there a good tldr on this topic?
This article is the tldr.
I’ll go back to the point about authorship. I’m Not a mathematician but I am in academia. if you are fucking around with authorship you are immediately suspect.
That aspect alone would/should be unthinkable to any serious academic. Authorship reflects who did the work and changing it for business competition reasons should be a red flag for multiple different reasons. They include, the sheer tactlessness of treating a major theoretical advancement as a competitive posturing first, the norms of academia second, and all the misunderstandings of the culture of the disciplines culture that people will now suspect are hiding beneath the visible surface (insert topography joke).
Math as a field is fairly unique even in how they list authorship. It was long the norm that authorship to be alphabetical because the idea of first, second, senior etc authorship is harder to define than many other fields.
“The stated rationale for alphabetical order is that it treats co-authorship as intellectually joint work: every listed author’s name carries equal weight, and no one has to negotiate, or be seen to negotiate, over billing. That is a genuine advantage over position-coded conventions, where disputes over who is “first author” are one of the most common sources of authorship conflict in fields that use them” [0]
That norm is changing, slowly, but one option people are pursuing is notable: randomized author order. Their is a perception that alphabetical is too biased…that’s the world OpenAI is stepping into when they make that offer of authorship to one scholar with a demand that he exclude his partner.
I can’t speak to the facts of anything else in this, but if a grad student came to me and said someone made them that offer, I would tell them to run and if they were brave report it.
[0] a to the point lay description of the history of math authorship can be found here: https://casrai.org/guides/mathematics-alphabetical-authorshi...
OAI's side of the story is that they discovered the approaches (and indeed solved problems - Euler equations vs NS equations) differed. They then offered Buckmaster lead authorship of OAI's proof, without Alpöge. But they never demanded that Alpöge be stripped of coauthorship on resolving the regularity of the Euler equations. At least that's the claim.
https://xcancel.com/SebastienBubeck/status/20973794116915163...
Yes, there is no doubt about scientific misconduct. I think the entire math community agree on that.