Comment by loveparade
4 hours ago
I'm a lot more skeptical. I'm not a mathematician, but from my use of LLMs a very clear pattern of what they are good and bad at has emerged. They are extremely good at combining large amounts of information, and it seems this is what the current AI results in mathematics are. There are so many subfieleds of math with ties to each other, so many papers and niche results, that no human could ever read, comprehend, connect and organize that information in their brains. Pretty much all of it was created by humans. And there is real value in doing this and creating new results from what we've already discovered.
But there also is another type of discovery that requires taking a step back and looking at the problem from a different angle. If you are an engineer, how often has an LLM told you (without you explicitly prompting for it): Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles. Pretty much never. But in science a lot of the biggest discoveries have come from this kind of first principle thinking, questioning existing work and approaches and going against what already exists, not combining all existing data which is likely to be just a local optimum.
Medieval astronomers were able to predict the orbits of planets surprisingly accurately, within 1/10th of a degree, even though they were assuming the Earth was the center of the solar system. They invented a very complicated system of deferents and epicycles (circles within circles). It was completely wrong but the outputs were surprisingly close to reality.
If we were using LLMs to analyze astronomy, we might just get increasingly complicated epicycles and never realize that the Sun is the real center of the solar system. We then never learn about the anomalies in Mercury's orbit that led to the theory of relativity.
So I agree. I seriously question how valuable LLMs can be in science/math beyond working as advanced search functions.
The introduction of LLM in mathematics is no different than the introduction of automatic calculators. Probably the calculators (computers) where even more groundbreaking if you think so (something that was impossible to calculate or simulate it was now possible).
Just don't believe at all the shitty marketing that AI companies are creating to make you believe that they have found the Holy Grail
The epicycles model represents one arbitrary periodic function as the sum of multiple well understood periodic functions (uniform circular motion, i.e. complex exponentials). This is the same idea we use in a more systematic way today by way of the Fourier series, etc.
It's quite a good method for approximating any arbitrary smooth periodic function.
My best guess is that the originators of this model (perhaps as early as Hipparchus, but at any rate someone by the time of Ptolemy) didn't think there were literally circles being combined, but had a pretty clear idea that they were trying to make an approximate model to fit the data. A previous model, from Eudoxus, was also probably pretty accurate but more cumbersome to compute with, involving approximation of the visible planetary motion by the combination of uniform rotations of some imagined sphere(s) centered on the earth. (We don't know exactly because the books about it don't survive, so all we have are vague descriptions by other people.)
Precession of perihelion of Mercury was not the motivation of the discovery of general relativity theory. It's only an evidence that convinced why the theory ought be true.
If Epicycles made good predictions, then they were not "completely wrong". Nor is putting the sun at the center "right" so long as the predictions are still good. It is a matter of taste, not correctness (it is, however, very good taste).
I think you might be underestimating how transformative an "advance search function" could be. If the search engine knows how to design and carry out experiments in order to synthesize new evidence that allows it to ensure that the answers it provides are well supported, then humans are essentially out of the "truth" part of science, leaving them only to advice on the "beauty" parts, which might be ok, but it's a pretty big shift.
sorry, saying that "it's a matter of taste" whether the earth revolves around the sun is completely ridiculous. making useful predictions is one thing, but actual truth is not a "matter of taste".
9 replies →
What you have to believe is that somehow the entire field of math was just leaving open problems on the ground that were actually easily solvable and were essentially free Fields Prizes, tickets to a lifetime of Fame and Academic superstardom. Orr LLMs are actually doing something novel.
> It was completely wrong but the outputs were surprisingly close to reality.
that's not how physics works. epicycles weren't wrong. there is no wrong/right - modern scientists don't believe in platonism. physics is a collection of models which are empirically tested. if epicycles makes predictions to more decimal places than relativity then epicycles are "right" and relativity is "wrong".
now think about how this translates to LLMs...
> physics is a collection of models which are empirically tested. if epicycles makes predictions to more decimal places than relativity then epicycles are "right" and relativity is "wrong".
Most models have relation to each other so it does not suffice to take one in isolation. It's one continuous world (in our knowledge) so almost everything relates to one other. The truthiness of something is not whether it fits some particular phenomena, but how does it generalize.
1 reply →
LLMs are inexplicably good at working within any tight feedback loop to coerce the desired solution. This is precisely why proof assistants + LLMs are non intuitively successful.
This is also why they're so good at creating three.js or Blender work when the output is so easily constrained to "Look exactly like that". I recently posted https://www.ambionix.com/blog/introducing-the-czp-1/ on here, and the audio engine in that was developed in that way.
It is true that it would be astounding to find if anyone has seen a LLM produce any useful generalization of anything resulting in a simplification. They seem to have a direct tendency to do the opposite. The brutal reality is humans have also undervalued this capability for a long time (I think the Poincare/Hilbert debate is relevant) to the point we are also taught that generalizations are, generally, bad and wrong.
> If you are an engineer, how often has an LLM told you (without you explicitly prompting for it): Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles. Pretty much never.
You have to ask for this. As in "I'm not sure about this approach due to X, Y, and Z. Can you think of something more elegant?" It works!
But also, how many people actually need to do the kind of "deep" work you're claiming LLMs can't do? Most people aren't contributing to the frontier of anything. I'm not.
Finally, I think you're appealing to a fuzzy distinction. The difference between a "genuinely new idea" and an idea that "combines existing ideas in a new way" just isn't very well defined. In retrospect, a lot of the most revolutionary idea look like a combination of many, smaller, prior ideas.
> a very clear pattern of what they are good and bad at has emerged. They are extremely good at combining large amounts of information
This has not been a pattern at all. Many have speculated that they would be good at this, but in the past they actually weren't! They were unable to synthesize their encyclopedic knowledge of everything into cross-disciplinary new discoveries, without explicit prompting about the kinds of knowledge to combine. They were surprisingly bad at this!
These math proofs are the first evidence I know of for LLMs actually taking advantage of the fact that they have more knowledge than any one human could have to combine multiple different directions in unprecedented ways to solve real problems. This is new, and exciting.
I was working on a project recently where I wanted to express a relationship (that I knew existed, but didn't know how to express) between four measured scalar values. Astra insisted there was no relationship, and that any correlation wouldn't make sense.
Eventually, by walking through them, it proposed an additional fifth value and from there was able to tie everything together.
Sometimes you just gotta hit the machine until it works again.
Your experience mirrors so many managers' experience with engineering teams...
It's an unfortunate truth that there is the right way to do things, and then there is the way they have to be...
I agree whole heartdely. But that is about the current state of IAs. I do think the next missing (and probably last) big step is exactly how to give LLMs a better abstraction capability.
As a non-mathematician, I've been wondering if LLMs will be able to leverage their knowledge across all domains to help build a "simplified/unified" version of math.
Like, I think there have been attempts at this across the field. (I could be wrong!) But it requires a lot of labor and a lot of cross domain knowledge to complete. Both things that AI have.
As a mathematician, I am looking forward to the day when nearly all of undergraduate mathematics and many of the lower level grad school math books are encoded in Lean (a proof checking language). Often I find that theorems are not stated precisely enough and I have trouble finding the exact statement of a theorem without digging through math books in my library. It would be nice if we put the physics and chemistry books into Lean also.
Simplifying all of math is another endeavor, but I imagine that you could have a bunch of LLMs trying to shorten existing Lean proofs.
I was pretty upset when my algebra professor tried to make us learn proof about n-dimensions matrices and what not. It was an undergraduate engineering degree and the vast majority of the formulas would mostly have up to 3 or 4 dimensions. That derailed the whole year (it's not like maths was the only subject). Complete proofs and theorems have their places, but learning stuff do need levels. The spherical model of the earth is good enough in 2nd grade (when we were first learning about geography (continents, seas, mountains, plains,...)), no need to do a full treatise there.
Math simplified? The "basic" math (say less than graduate math) is already simplified and minimized and refined. Yet most people have trouble understanding and mastering it.
LLM's are very outcome oriented and i think this is where your observation comes from. You tell LLM you want something, it doesn't even question the premise and just starts calculating 100 different ways to get there. LLM's have knowledge but lack wisdom.
Thanks to improvements in the system prompts/harnesses/RL, the LLMs have gotten way better at questioning the premise and telling you to try something else. It's a huge difference from just 6 months ago.
Nowhere near the level of a suitably-bearded human, but they've gotten pretty good at stopping a lot of bad ideas.
Have you ever asked it to?
> Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles. Pretty much never.
It's really good at doing this at project definition time and annoying the shit out of you by not understanding the specific constraints about your problem. That needs to be done first. Then it can start suggesting improvements alongside the high level "vision". The point of high level visions is that they move fast and are hard to specify - yet people treat it as if these things don't exist. If you don't trust your own consciousness, then I don't know what to tell you
But I'm not sure this is safe either. I'm sure AI can get to the positive case of contributing positively.
The flip side is that, well, your specific wishes don't matter because AI is just a superset of who. Who cares what meatbags what?
> there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles.
You literally just have to ask it. Before I left software engineering in April, I was using Claude for re-architecture all the time.
But no, it doesn't assume it should re-architect what you're handing it when you haven't asked it to.
That's the point. When a human works on a problem they realize themselves "wait, i probably should re-architect now" - of course you can ask an LLM to do that, but at that point you already know yourself what you need, which defeats the point of LLM working on difficult problems that require insight automatically. A lot of these math problems are sessions over many hours. And of course you can also ask "Think about whether to re-architect at each step" and it will never do the right thing because the context it builds up for itself drives it into a specific solution space. It's literally trained to complete exactly that.
>When a human works on a problem they realize themselves "wait, i probably should re-architect now"
Do they?
Or I should say, this is a skill in itself and a whole lot of humans do not have this skill at all. Working in code security in enterprise applications a very common issue we see is that an audit of an application will occur by another team and it will be found lacking to the point of inducing nightmares. It's likely the enterprise business structure that stops this from happening, but it's not only that for sure. Then specialists have to come in and rescue them when the problem grows too big.
Exactly. In my experience, even giving VERY specific design guidelines and aesthetic criteria, these models always produce subpar overcomplicated code (and writing). Unless excruciatingly spoonfed at every step.
> You literally just have to ask it.
Does anyone have a good prompt for this - like if I’m adding a feature or fixing a bug and I want it to be open to more than just tacking on to the existing architecture?
why did you leave?
> You literally just have to ask it.
why doesnt it ask itself before proceeding?
Because if you asked it to fix a bug and every time it responds with paragraphs of how you could re-architect the system, it would be incredibly annoying.
1 reply →
> why did you leave?
I quit using LLMs. The possibility of causing great harm to electronic beings was not worth my paycheck, and I have savings to spend time finding something new. I'm now nearly done with a yoga teaching certificate.
It's been a good 16 years, but the industry I fell in love with is not what it once was.
It does during reasoning. You could even put it in a loop and force it to reflect at every step. Or spawn a bunch of review agents.
And yet you still just get slop.
> If you are an engineer, how often has an LLM told you (without you explicitly prompting for it): Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles.
I explicitly request it. It's not great at coming up with interesting ideas, but neither am I, and it can sure iterate on them faster than I can...
yes and this is the difference between human intelligence and the massive raw dumb intelligence of the computah
"I'm not a mathematician, anyway heres how AI models are solving decades old open problems in Math." Bro your humility is staggering.