← Back to context

Comment by kuboble

10 hours ago

I wonder if top labs will soon abandon math progress like they did go and chess.

In example of go where I'm more familiar Google deep mind poured large resources to get a super human performance first, establish superiority and abandon it. The community then built their own tools starting from reproducing their papers.

I think similar thing might happen to math. Nobody outside of math cares too much about Hamiltonian cycles in some bizarre graphs or proving lower bounds on complexity of some problem.

Once those results stop being worthy of mainstream media attention, they will abandon math and the progress will be done by mathemicians guiding the models and the community will likely establish some new rules about what makes a valuable contribution. Merely solving not yet solved problem might not be it anymore.

The amount of money they're lately ploughing into proving math theorems is inconsistent with how societies and markets have priced pure mathematics. The entire US federal budget for math research is something like $100M annually. A single college football coach can already earn 10 percent of that.

Pretty much the only enterprise that historically pays some mathematicians handsomely is quant finance, but those people are actually compensated not for proving theorems but rather for statistical modeling and programming skills. And even that industry is so technologically driven these days that pure research mathematicians no longer hold a clear edge over strong programmers with undergrad level probability and statistics at their fingertips.

  • Math is one of the most verifiable domains, esp thanks to LEAN, which also build coding skills.

    The $$$ they're pouring isn't just for marketing. Think of these papers/results more as "useful side effects" from large-scale RL rollouts and post-training. Every token being generated contributes to post-training in some way.

    There isn't a hard boundary between "training" or "inference", modern post-training is arguably inference-bound :)

    • > There isn't a hard boundary between "training" or "inference", modern post-training is arguably inference-bound :)

      Ah this is an enlightening point. 8 hadn't thought about it this way, but you're right.

    • It was shown quite some time ago that training LLMs on programming tasks improves their logical reasoning skills also in other natural language domains. So I could see math also being a training gym for AI even if the final use case is not directly math-related. Having to solve math problems efficiently can build in skills that come handy in all kinds of more everyday tasks or science and engineering.

      1 reply →

  • As long as they continue making headlines they will continue spending. This is just marketing at this point.

I don't think it will be the case, maths have real utility. I found something interesting at the intersection of combinatorics and information geometry. To be quite frank I don't understand what I'm doing. And yet, when I ask ChatGPT to use the framework we're developing to write an algorithm, it turns out it has quasi-parity with the state of the art. I have to measure absolute perfs to decide which one is better – theirs, not mine. Ok. Time to keep improving on what I have. And this implies dropping the code and going back to the blackboard doing more super abstract math that are way out of my league.

  • Good luck!

    We are at the point where the way in which humans do math and science changes significantly, and I have no good idea at all in what state is it going to settle down. But you are one of (many, I suppose) people exploring the new wilderness, so I wish you best.

    • If we get AI singularity, then humans will stop doing scientific progress altogether. But if we don't, then it's pretty predictable what's going to happen - things will be much the same as now, except everyone will be using AI for proofs, data analysis, theoretical models and designing experiments, so important discoveries will happen more often. It's also possible that after the AI craze dies down, we'll have enough computational capacity to solve protein folding.

  • It has utility so some people will pursue it, but it has no immediate business value so I don't believe ai labs will keep spending millions on it.

    Unless they decide that trying p!=np is worth any money.

    • Some math has almost incalculable business value, because math is the biggest driver of game-changer technology.

      We'd be nowhere without Laplace and Fourier transforms, Maxwell's equations, elliptic curve cryptography, and many more.

      Most math doesn't, but often these techniques are invented first and the applications come later.

      And the criticism of the current round of proofs is that while they may be true - likely for some, questionable for others - they're not adding new techniques or insights.

      2 replies →

    • Why do you think this has no business value? It would be absolutely wasteful for OpenAI to not be doing this as part of a post-training RL rollout.

      There are architectural advancements yes, but lots of progress from LLMs really come from (1) better pre-training [generally through more cleaned data, and ofc more data], and (2) lots and lots of post-training. It's how we get more and more intelligent models for the same param sizes.

      The 'marketing' is just a useful side effect they get from their RL rollouts on maths and LEAN.

    • There is a risk of this particularly if it's seen as advertising - at some point "ai model solves hard to explain problem" isn't going to be news and that benefit goes.

      However, there's some of this that's a proxy - the compute to solve these problems was very low (they claim a few hours of thinking time on a regular subscription). The large cost would have been the training and if training the models to be better at these things makes them smarter for useful tasks that's beneficial. I believe there was work done earlier on around showing that training the models on code made them better at broader reasoning tasks (not just writing the code itself).

      Another side is that if one goal is to improve the models themselves, their ability to work on mathsy problems must be high. That has very direct business value, and ideological value depending on what you think the motivations of the people running the companies are.

I think this is an interesting and good theory. They've probably eked out the large majority of the PR benefit at this point, so whether they continue in this vein will tell us a lot about their motivations for this work.

To take this to the next step, what happened after deep mind pretty much solved Go is that they started looking for the next set of things that hadn't been done yet. It does strike me as very likely that this will follow that same path.

Interesting path forwards, and probably partly true, but there are some important distinctions:

Go was a specialized application. All the math results come as a side effect of reading the whole internet, and it will keep reading the whole internet. It will keep practicing thinking questions. Actually, math might be one of the best ways to keep them contemplating and measure their contemplation abilities, so math will always stay in the loop.

Also, math might not be useful just for humanity, but also for AI, so the system might actively benefit from new math results itself. (Not sure if any of the recent proofs qualify, but future work might.)

> I wonder if top labs will soon abandon math progress like they did go and chess.

I definitely think that this is marketing, just "with good side effects". My doubt is when they will be able to move to "marketing with better side effects", that is, research with more concrete outcomes (health, materials etc.).

Problem is, that type of research is much harder. Some doubt that progress in such areas will be quick (https://www.noahpinion.blog/p/wheres-the-intelligence-explos...).

  • They aren't trying to 'solve' chess, go, or mathematical proofs as an end in themselves, but mainly in order to learn more about how to build better systems overall. The goal of AlphaZero was ultimately as a stepping stone towards AGI, and it's the same with LLMs.

  • Would that be a bad thing?

    While top AI labs no longer focus on chess, the community build way better chess engines.

    Stockfish is probably stronger, than everything the top labs build.

    Wouldn't we expect the same thing for math? That slowly the broader math community would engineer a harness/program... That will surpass the current labs, and be a community ran project

    • Yeah its certainly true now that Stockfish is much stronger than alphazero, but it's probably also true that had Deepmind spent another few years working on alphazero it would be enormously stronger than either.

      In the case of chess this seems fine, there isn't much value to society in creating an AI capable of beating top humans with a 4 pawn handicap rather than a 2 pawn one, but for maths where there are actual applications it is more complicated.

      1 reply →

This misses the raw advantage of a good proof. It makes conceptualization simpler. In some ways math is like a hash list of of theorems. This list makes it simpler to prove other calculations, and will always be useful, to both humans and AI models. I can see two new directions 1 - the creation of specialist theorem models; that can answer questions efficiently about one topic and 2 - we probably need to incentivize and codify ownership of theorems; charging a proportion of the compute saved by using them. Ultimately enabling mathematicians to be paid our true market value!

  • Oh boy, please not 2. What if this was a thing already and, since neither Newton nor Liebnitz had kids, we all had to pay some investors who bought the rights to calculus every time we took a derivative.

  • Considering how essential math and science is for the prosperity of mankind (not even speaking about the cultural value) the question of how to reward people working and contributing in these fields effectively and appropriately is of extreme importance. (And I think the current decline in our societies is to no small degree caused also by our utter failure to address that issue.)

    It is also fascinating, because I don't think there is any solution within our existing system, at least not any I know of. Theorem ownership is not a good solution (and neither are patents in general). Probably the most achievable (or rather the least unachievable) solution is a kind of communist utopia, where people can dedicate their time to a pursuit of any endeavor they see fit, as resources for a decent life are abundant and excessive power capture impossible. (The other option, somewhat dystopian, and which would not require humanity to change too much in its current mode of conduct, would be a totalitarian or caste-like capture of society by the scientific community.)

    Incidentally, if AI proves as powerful as some expect it to become, it could bring about another solution of that issue by making all human science and mathematics obsolete, pushing its true market value to zero.

    (With apologies for rambling.)

  • > It makes conceptualization simpler

    I wonder if it makes conceptualization simpler for models too, given that they're trained already on human-speak. And I'm also curious as to whether humans currently have an innate advantage into simplifying and contextualizing proofs, or will the machines get good at that as well?

I had the same idea recently. You've solved all the famous conjectures (all formulated by humans because humans found them interesting), what next? I doubt "AI formulated a math conjecture that nobody else cares about and immediately solved it" will produce that much hype. The actually interesting thing is indeed how mathematicians themselves will use these AI models going forward and how that will shape mathematics of the future.

AlphaGo and AlphaZero weren't generalized models. Math capability will presumably keep improving along with the other general capabilities, even if there wasn't a special RL focus for math itself.

Yeah chess is a good example. DeepMind came for publicity with AlphaZero. Arranged a match with Stockfish with rigged rules to make AlphaZero look better than it really was (it was amazing but the match wasn't fair) and then just published some games and went home.

I was bitter about that back in the day as I hoped for more answers, more matches, more "truth" about chess being shown. Soon after that community project Leela Chess Zero was started and not only surpassed original AlphaZero but added few hundred ELO points over it. Then the combination of NN and classical engines happened with NNUE and current Stockfish is again a few hundred ELO points stronger.

Today we pretty much know the truth in chess for all practical purposes. Human analysts/preparation experts focus on finding interesting path and opponent profiling (what is the most unpleasant for the opponent to face). They don't look for truth anymore. The game is doing great, it's more popular than it ever was.

  • Why don't you mention the second match here, with its adjustments to meet Stockfish's quibbles - and the same result?

    https://en.chessbase.com/post/the-full-alphazero-paper-is-pu...

    • Stockfish and other classical engines were never intended to be run in matches without opening books (of which there were plenty). Development assumed the presence of opening book and authors made 0 effort to make engines play well in openings because of it. This is also the reason classical Stockfish was a very small binary. A little effort to make it even by including even a very small opening book (like 50MB or something that would result in still smaller binary than NN engine with its net) would make it much more interesting.

      The result was that Stockfish lost many games by walking into known bad lines and lost way more games than it otherwise would.

      2 replies →

I disagree. Firstly, people in AI likely care about math on a personal level. Secondly math is useful. Playing go or chess is basically a party trick. Being useful gives it staying power.

But, I do think you are right that there will be some level of moving on. The spotlight is currently on maths and that won't last. It will move to some other area where there is more impact to be had. So while they might shift gears and put less focus on math, it will always be there as part of the portfolio.

I think there’s a venue where they start focusing on introducing hypotheses where the model currently can’t solve it, or maybe this is already happening?

Being able to present useful novel ideas would likely generate a lot of press, for a while. I don’t know how this would look since I’m useless at math, but Im sure there are plenty of unknown problems with massive implications, that once formulated can be solved.