Comment by sashank_1509

1 day ago

Both things can be true:

1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation.

2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and have little to do with some specific technique mentioned in its chat.

The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths, and is basically solving anything you throw at it. We’ll know soon enough, but I’m inclined to believe this is true. Maths is a fully verifiable domain amenable to self play, massive scale RL can develop a search agent far better than any human and I’m inclined to believe OAI would have solved these conjectures without any of this chat data in its pre-training.

On your second point: there is a more plausible explanation which David Bessis calls the "overhang". The short version is that there is a large amount of relatively low hanging fruits in mathematics, because no human has broad enough knowledge and enough time to try them all. AI is not constraint by that, and therefore can systematically pluck all those low hanging fruits.

Quote: "The Overhang consists of the unrealized capital gains of past mathematical creativity, the latent value from connecting the dots in the existing corpus. It is a dividend of canonization. Mathematician X states problem A, mathematician Y crafts concept B, then mathematician Z notices that B trivially solves A and “captures” the social reward. But in the process of capturing the reward, Z usually introduces new concepts and new open problems, reinjecting latent value into the Overhang.

LLMs can be trained on the entirety of the mathematical corpus. Thanks to their phenomenal memorization and pattern-matching abilities (without always being able to map out their associative logic and attribute due credits), they are in a unique position to harvest the Overhang. By contrast, professional mathematicians have typically read a few hundred articles in their career, out of millions of existing references, less than 0.1% of the total.

This will lead to great discoveries, which is unambiguously exciting. But it could also lead to a sad new deal, where human slaves painfully curate the Overhang while AIs systematically beat them at the finish line."

source: https://substack.com/inbox/post/183753276

  • >Quote: "The Overhang consists of the unrealized capital gains of past mathematical creativity, the latent value from connecting the dots in the existing corpus. It is a dividend of canonization. Mathematician X states problem A, mathematician Y crafts concept B, then mathematician Z notices that B trivially solves A and “captures” the social reward.

    I've made an entire career out of being 'jack of all trades, master of none'. Being able to synthesize connections from relatively trivial knowledge in a bunch of domains is SOP for many humans as well. I think AI just has deeper knowledge and better pattern matching to make up for it's (at least now) lack of strength in cognition and 'ex nihilo' creativity.

    (Which probably isn't 'ex nihilo' at all, and has more to do with the plethora of modalities that humans live in vs. large language models. For example, why do we pick the color red for notating important things and why do we say a schedule 'slips'...these are informed by a shared human experience borne of distinct physical sensation deep in our wiring that LLMs can only infer from what we write.)

    • A college advisor I had 20 years ago was a firm believer that interdisciplinarity was the future, that generalist skills and the ability to make connections between different fields would be paramount in advancing science. I suppose he was right in the big picture, even if the career prospects for human generalists aren't looking so rosy.

      1 reply →

    • How do you thrive in an environment of specialists? That's is the problem I seem to have. I'm spread a little across a few of the domains involved with what I do. Because of that, I have a bit more insight, so am very often the person pointing out relatively fundamental problems, usually caused by either not understanding the problems from a "first principles" perspective, resulting in, or being caused by, categorical type errors, where they've boxed a problem into a tiny space it doesn't belong.

      I've been trending "quiet" lately, because I don't like the "friction"/convincing aspect of it all. It's hard to get people to see things from a different angle, or even convincing them there's a problem to begin with!

      The last project required a complete redesign from a problem I pointed out during the first review, and second, and third, but now I'm seeing even more friction.

      Maybe this is just corporate life, after a group gets large.

      Any tricks/advice?

      1 reply →

  • > Quote: "The Overhang consists of the unrealized capital gains of past mathematical creativity, the latent value from connecting the dots in the existing corpus. It is a dividend of canonization. Mathematician X states problem A, mathematician Y crafts concept B, then mathematician Z notices that B trivially solves A and “captures” the social reward. But in the process of capturing the reward, Z usually introduces new concepts and new open problems, reinjecting latent value into the Overhang.

    That overhang seems like a precious resource for AI companies. They can exploit that overhang to inflate the impression of AI's capabilities, and hopefully that exploitation will discourage the next generation of mathematicians from pursuing math. If they play their cards right, OpenAI and Anthropic can dominate the field even if they ultimately can't replicate the creativity of human mathematicians, because they'll have driven their competition out.

    What we should be trying to achieve is a ladder-breaking maneuver: knock out the lower rungs so no person can reasonably climb to the top-reaches of mathematical skill anymore. That may ultimately result in stagnation, but it's what's best for AI, so it's what should be done now.

    We need to do everything we can to create the greatest-possible dependence on AI tools.

  • Great essay, thanks for sharing.

    When I was a software library developer, I came to resent application developers. I noticed a pattern. Libraries solved hard problems and did so carefully, thoughtfully, in a way that others could reuse. Apps would come along and carelessly, recklessly glue together several high quality libraries into a piece of software targeting a general audience. The apps would then harvest all the credit.

    What's happening in mathematics right now feels similar. Applications (theorems) were always how one built objective reputation, but libraries (concepts, definitions, boring lemmas) were also rewarded socially within the mathematics community. And individual mathematicians often managed to both build their own libraries, and use them to prove an important result. And then those libraries were sometimes of use in other results.

    Bessis asks whether AI Lean proofs will land in Mathlib or Mathslop. Or in my framing: will they be libraries, or applications?

    At present they're mostly Mathslop. The proven result is perhaps useful, but the methods employed aren't novel or reusable. I worry that this trend will only worsen, because applications make headlines, and the libraries they used do not. We are not properly incentivizing library development in OSS, or in math, or in infrastructure writ large. There's a serious credit assignment problem here.

    What might change this? Once the low hanging fruit is picked, will citation count rise in relative status again? Will we get result fatigue and start to reward legibility — no one cares unless the paper has an accompanying ELI5 tiktok video? A labeling regime that certifies the proof was produced sustainably, organically, by local artisans with no AI additives?

  • There is also "sexy proof", people want nice math that can be printed in t-shirt. Not super hard grind, where you need several years of studying, just to understand the question (that is before even trying to solve it).

    Many problems are solvable, but require months of work, and thousands of pages of proof. So people do not even try to create or verify the proof. AI changes that, it can verify and perhaps even simplify it, to more digestible form.

  • The overhang, being defined as the Cartesian product of existing knowledge — randomly combining existing knowledge.

    (I mean actually randomly, not asking an LLM to do the randomness.)

    Most of the output would be incoherent (like many dreams), but occasionally you would get a gem.

  • > no human has broad enough knowledge and enough time to try them all.

    The other part is, humans don’t really want to fund other humans doing this.

    Very few want to be a math major; and of those that do, fewer complete a grad degree; and for those that do get grad degrees, there’s scant few research jobs; and for those who do get jobs there’s hardly any research funding to go around.

    There does seem to be unlimited money for ai researchers to use ai to solve these problems though.

    We’ve turned education into job training, so because there’s no jobs in solving math problems, few aspire to do it. If there were more opportunities for people, more people would do it, and more low hanging fruit would be plucked.

    • I’m assuming the reported 22 million dollars worth of tokens used to solve this particular problem is far more than what humans have paid to solve it previously. So I think you’re correct.

      1 reply →

  • It's like AlphaGo but playing against all living mathematicians. (Overhang being low hanging fruit is what allows this comparison, of course the general moot point is the skepticism that LLMs are also innovative etc.)

    • We are not seeing those incredible moves yet. The approach used in N-S was conjectured to work after B&L’s initial breakthrough. See a post by Tao. So on one hand the proof is an amazing accomplishment. On the other hand, humans have not yet discovered any superhuman moves in the proof. Just $MM grind.

OpenAI said they sicced this agent army on Navier-Stokes on Sept 1st, while only a couple of days earlier OpenAI's Noam Brown happened to reply to a tweet saying that they had already tried to solve all the Millennium Prize problems and failed... So, it seems either the previous attempt didn't have the training to succeed, or was just not given the compute to do so.

Once OpenAI heard that Navier-Stokes was solved, this caused them to immediately revisit the problem and throw a ton of compute at it, apparently using a more (very) recent model than what they had tried before. What we don't know is just how recent this model was, and therefore what it may have been trained on. Buckmaster/Levant had apparently been working towards this for at least a year, and made their "forced" blow-up breakthrough on August 15th.

Presumably any anonymized prompts that are being trained on are part of pre-training, so older, but once OpenAI had heard that Navier-Stokes had been solved and wanted to revisit it, it seems possible they may have done a few weeks of incremental RL training on anything Navier-Stokes adjacent they could come up with, in addition to then throwing unlimited compute at it, now confident that there was something to find.

  • OpenAI have come out and said:

    >The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”

    >The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”

    https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...

    • Is there a reason they scoped that so narrowly to Buckmaster/codex/2 months

      two people worked on this for a year before the breakthrough. Perhaps that earlier work reduced the search space sufficiently to brute force the problem with 10,000 agents?

      3 replies →

    • OK, good to know (if they can be trusted - Altman clearly is a liar), but it doesn't really change the big picture much.

      1) OpenAI by their own admission, only re-tackled Navier-Stokes because they heard it had already been solved (but not yet published). This isn't advancing science or helping the mathematical community, this is just being a dick.

      2) OpenAI, specifically Sebastien Brubeck, then threaten to "not be nice" and "ruin the career" of one of the mathematicians whose work they had succeeded in duplicating, unless he agreed (which he refused to do) that his collaborator, an Anthropic employee, was not named. This is not only against mathematical norms of credit assignment, it is also being a pathetic human being.

      OpenAI would have you believe this result shows how powerful their mystery better-than-Astra model is, but the reality here is that this model needed 10,000 agents, $20M of compute, and the assistance of a whole team of people at OpenAI, to replicate (then exceed) the work that just took two people, with some academic grants as an AI spending budget to achieve (a few $100K - listed below).

      https://cims.nyu.edu/~tristanb/

      I'd say advantage humans this time. Better luck next time OpenAI - and if you don't want unfavorable comparisons then maybe choose to work on problems that have not been solved yet, and that humans are NOT making nice progress on.

      18 replies →

    • Apart from the well-known dubious position of OpenAI wrt truth, the prompts/inputs do mot include the outputs.

      You can train on a sequence of outputs. In the end, OpenAI outputs are OpenAI's property.

      You can learn a lot from a single side of a conversation.

      3 replies →

    • It seems logical since if one used chats in train, one would expect that there would be a delay before their use to get them the form appropriate for batch learning.

      The only way the chat could have been used would be for Open AI to baldly violate their policies.

      That said, sometimes it take very little information to point someone in a given direction, "I'm working on Navier-Stokes" said by someone with a given specialization might itself be very useful information.

  • And conceptually novel approaches to outstanding problems are the sort of thing that a retrain should pick up on, because they would be hard to compress into what it already knows.

  • > What we don't know is just how recent this model was, and therefore what it may have been trained on.

    OpenAI's statement says that they began training their new model on August 28.

    • omitting when training concluded

      edit: ffsm8 makes a great point below, it doesn't matter. I'm not great with dates, sorry.

      1 reply →

  • Openai said that a new model became available to them during this. But that could mean anything from a big new base model to a LoRA, fine-tuned on a few dozen prompts...

Even OpenAI's own publication [0] on Navier-Stokes from two days ago appears to contradict "basically solving anything you throw at it". The chart shows a pass rate of ~0.5 (vs. Astra's ~0.2) on "a curated set of open math problems". (Based on the timelines and events described in the publication, I presume that the "Internal Model" in the publication represents OpenAI's latest and greatest model. Evidently, this pass rate may improve in the future.)

[0] https://openai.com/index/navier-stokes-solution/

I feel that we don’t praise Lean enough. AFAIU it’s what enables LLMs to brute force those problems

  • The brute-forcing is a good, old-fashioned generate-and-test approach like in Simon and Newell's Logic Theorist, which was presented in the Dartmouth convention in 1956, where AI was named by John McCarthy. Logic Theorist caused a big stir by (re) proving several of the theorems in Principia Mathematica by Russel and Whitehead.

    There was much excitement, then, as now, for this kind of approach and there were several systems that followed along the same lines, e.g. Automated Mathematician by Doug Lenat.

    Eventually it became clear that this approach is limited by what it can generate: you may have a sound and complete verifier, but if the generator, i.e. the first step in the generate-and-test pipeline, is incomplete, then the entire thing will run out of steam sooner or later.

    The difference with LLMs is that they are... well, large. They are the most powerful generators ever created. That means their limits are not in sight and it will probably take us a very long time to find them.

    Which is all to say that, yes of course, automatic verification is indispensable. But without an LLM generating an unprecedentedly large number of plausible theorems, there would be no AI mathematics, or in any case AI mathematics wouldn't have gone as far as it has.

  • True, but could humans cross pollinating lean x prolog x A* ( or any search algorithm) could have solved such math problems with super computer ?

    • I cannot say, math research isn’t my domain of expertise, I’m just trying to follow along :)

      But I find it interesting that Lean, a validator/compiler made by humans, is what enables those discoveries. But somehow all the praise goes to the models

      2 replies →

    • I don't think so. People have been trying things like this with evolutionary algorithms for a very long time already. LLMs can interleave symbolic manipulation with empirical experiments and simulations and charts and thinking/reasoning text, and an LLM will much more efficiently search the space of candidate ideas than any handcrafted mutation algorithm. Any task with a cheaply verifiable goal that requires fanning out across a massive search space is ideal for contemporary LLM technology to make progress with.

Both can be true:

1. OpenAI couldn't have solved the problem without the researchers' private data for training.

2. OpenAI models can solve math problems

  • Very likely.

    These mathematicians’ prompts are not like “hey chat, please solve Navier-Stokes for me”. They add real expertise and intuition from the cutting edge of their field.

  • Anthropic isnt getting enough scrutiny for their unprofessionalism:

    1. Anthropic employee working on monumental problem but didnt receive/ask for the full backing of the company's resources

    2. May or may not be mixing unreleased Claude output with Codex without zero data retention agreement

    3. Victory lap on Twitter and giggling around the city before they finished the job, sparking rumors for competitors

    • How dare employees do something without asking for the full backing of the company's resources. Incredibly unethical!

    • Dr. Buckmaster sounds unsanitary.

      Recklessly prompting OpenAI without a care to the safety of their knowledge.

      And after that trying to cast aspersions at OpenAI?

      Hopefully we get some better facts, because OpenAI are disliked enough that a smear campaign could work against them.

      Edit: also the narritive is getting framed as OpenAI versus Anthropic. A highly political extremely capitalist fight is going on, and facts are victims.

  • You forgot possibility 3: OpenAI solved the problem without using any private training data from the two researchers.

    Everyone in this thread seems to have made up their mind about OpenAI's guilt though.

    • If the new model is that good, and is chewing through open problems at an unprecedented rate, the smart move would have been to let the humans have their W on this one and present solutions to those other problems.

      Especially if there really is a long list of them.

      "Here are a few hundred proofs" is far more convincing than "We really Navier Stokes and coincidentally someone else did too but we don't know the details or anything, who us, definitely not."

      It's a PR fiasco, and a cynic might wonder if it's entirely about the IPO.

      I'm consistently entertained by how these companies, with the most advanced models on the planet, consistently do the most idiotic things.

    • Extraordinary claims require extraordinary evidence.

      An article post that wouldn't even amount to a white paper + the LEAN proof is not evidence of how they got to produce it.

If your rumor is true, what we are witnessing is a giant paradigm shift rather than individual incidents. Mathematicians were the first victims of super-intelligence.

Of course it’s not an endless source. They had to burn millions of dollars to solve a single problem.

  • >They had to burn millions of dollars to solve a single problem

    I'd like to adjust that to "They had to burn a lot of energy (create a lot of entropy) to solve a single problem. As we go into the super-intelligence age the current paradigm of money as humans understand it may break at some point. For example to a paperclip-maximizer money at best is a short term instrumental goal, hard power of matter conversion machines is what it wants and once it has those money no longer has purpose.

  • They "burn" a lot when they do benchmarks, while these runs can become valid roll outs for training. Perhaps less efficient than other data creation, but hardly burned in the same way.

  • > were the first victims

    Spinning it negatively like that doesn't do anybody good.

    Were mathematicians the "victims" of calculators? of Matlab?

    Were writers the ""vIcTiMs"" of word processors?? (apparently yes, according to old TV shows about computers during the 1980s, that you can see on YouTube)

    > "tHiS iS nOt ThE sAmE" — Everyone every time.

    No, just look it up. Look into old magazines and TV shows or newspaper articles from whenever a disruptive new technology came out.

    • What you say is true but ... This is qualitatively different than calculators or computers.

      I'm a professional mathematician and all the better mathematicians I know are in crisis mode. Most of us hadn't taken this sufficiently seriously and don't know how to use these models effectively but we play with them and immediately see that the entire way we've worked all our professional lives has to change. We worry less about ourselves than about the younger folks. I've got good ideas ai still doesn't know about ... Younger folks may not get the chance.

      1 reply →

    • It’s not the same. AI potentially completely replaces intellectual work without creating any* new jobs (*almost any - there will be some extra jobs for building data centers but that’s negligible).

      3 replies →

> The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths

Obviously these are unbiased and trustworthy sources.

If they have solved hundreds of open problems in math, why are they publishing results for the ones other mathematicians happen to be working on at the same time? Why not the others?

  • Well I'm sure if they find a millennium prize problem that no mathematician has worked on recently they will get right on publishing that.

  • You think other mathematicians are currently working on very little subset of relatively low-hanging fruit problems?

The leakage wouldn't be from training, but from other uses of Personal Data.

As far as I understand it, users can opt out from the training aspect, but they cannot stop their conversations (“User Content”) being used “[t]o improve and develop our Services and conduct research, for example to develop new features”.

The big question is whether OpenAI is training on "de-identified" sessions that are marked as "do not use for training"

The answer is almost certainly yes, and this is a problem for most users.

> We’ll know soon enough, but I’m inclined to believe this is true.

I mean, we’ll know as soon as they decide they want to provide verifiable proof. Really dragging their feet on this front so far.

I’m inclined to believe this is false.

The Cult tells us the AI is almight andpowerful; unfortunately, the cult cant actually describe the indescribable.