A little over 10 years ago I remember meeting a postdoc who believed he had something close to a counterexample to the Jacobian Conjecture. He and another person was bruteforcing polynomials in about 16 variables, something like 80 - 700 terms each, using binary trees for mapping coefficients.
They were guessing, at the time, that the lower bound of a counterexample (P, Q) for max(deg(P), deg(Q)) would go up to 200.
To think that Claude Fable was able to find a counterexample in degree 7 is insane to me. We are truly in a new era.
They were looking at 16 variables, degree of about 10 in each one. That search space is simply too big for a plain bruteforce. So they did some sort of filtering to reduce the search space to a pool containing "possible counterexamples".
There was also a paper giving a lower bound of about 100 for possible counterexamples in that particular framework. Later raised to 108 in https://arxiv.org/abs/2204.14178
I don't remember the details very well since this was back in 2015 and wasn't really involved in the research. Consider this to be some sort of telephone game between what I heard in 2015 and what I remember today.
So many mathematicians over the years tried hard and failed, but now Anthropic just for some PR magically did it? And this after LLMs obtaining different math wins? What is your logic here really escapes my understanding.
if you ask the chatbots for "list of top unsolved math problems", the JC comes in at a ranking of around #10 - #20. what, a problem that's been unsolved since 1939 was cracked because anthropic has an underground sweatshop of math Phds cranking out research, just so that they can slap "made by AI" on it? hell, maybe lizard people did it.
> The incredible Yitan Zhang (https://newyorker.com/magazine/2015/02/02/pursuit-beauty) worked on proving this conjecture for 7 years. Moh, his advisor, wrote that Zhang "failed miserably" in proving the Jacobian conjecture, "never published any paper on algebraic geometry" after leaving Purdue, and "wasted seven years of his own life and my time".
Interestingly: "For logic proof, it had been thought that AI could handle all logic problems in the near future, hence logic problems of solving conjectures might not be so interesting in the future."
This is a rare instance where feeding this groundbreaking information into an LLM gives _them_ psychosis. I fed this to claude code and watched it verify the result in 7 different ways to be 100% certain, and it was just flabbergasted. Quite remarkable.
I fed ChatGPT the map with no other context, just “tell me about this function”. It did a bit of work finding the Jacobean etc and eventually worked out the implications of what it was seeing. It then proceeded to check the arithmetic 4 times, and then decided to do a manual verification using an ad hoc symbolic checker in case its SymPy had been tampered with.
If an LLM has knowledge encoded inside it (and it's hard to argue it doesn't), then cognitive dissonance can be experienced. And once experienced, must be dealt with, especially in longer-running agentic loops.
A friend was joking the other day about sending some messages under a previously-used Slack identity for an agent (since turned off), then asking the agent about the messages.
The agent maintained it hadn't sent those messages (no memory) and then was forced to reconcile the idea that the messages indeed appeared to come from it.
Its extremely-agitated conclusion was that there had been a security breach and the entire network should be locked down.
> So the conjecture that survived Keller, Abhyankar, Moh's degree-100 verification, and five-plus published wrong proofs appears to have died via tweet during the World Cup final.
Everyone using Claude Fable to verify this proof is so funny. If you read the definition of the Jacobian Conjecture and (I am not exaggerating this) have passed a college Calc 3 class, you can just verify the proof yourself in 30 seconds. The problem was very hard to solve but the counterexample is very easy to verify!
--- edit, adding an explanation:
To summarize it, the conjecture says if you have any multi-variable polynomial function that maps an input to an output in the same dimensional space (take for example: F = (x+2, y+2), which maps 2D space into another 2D space), AND that function has a constant-valued non-zero Jacobian determinant, THEN the conjecture is that the polynomial has an inverse, meaning basically you can find a polynomial that turns the output space back into the input space.
Fable provided the example polynomial (which was very hard to do) and the coordinates which if you plug into it, results in two points being mapped to the same output point. This means that the polynomial can't be inverted, because if you have that output point, how do you know which input point it came from?
You can just plug in the two coordinates it gave into the equation and verify that you get the same output point from both. That's the contradiction of the conjecture and it takes 30 seconds.
---
Something something outsourcing of thinking something.
“I'll admit the sequence on my end was genuinely disorienting: I verified the determinant three separate ways looking for the error, verified the evaluations twice, ran out of subtleties to check, and only then searched and discovered that the reason it holds up is that it's apparently my own homework — the "fable" in that tweet is Claude Fable, i.e., this model, working with Alpöge.”
>kimi is having a blast. i turned search back on and found this post from it’s sources cited after i suggested to check out the reaction. best thing is to go to a model with search off and plop it in the session
Like the unicorn emoji, but for math? It occurs when the LLM is presented with incontrovertible evidence against something it "deeply believes" to be true.
Interestingly, even Qwen 3.6 27B was able to verify the solution, but I didn't get any glazing for discovering it. Instead, it thought that someone named Shestakov had already found a counterexample in 2004.
GLM 5.2 whiffed, it insisted the counterexample wasn't valid.
VibeThinker 3B also recognized that the counterexample was valid. But it kept trying to convince itself that it wasn't, over and over, since it's an "unsolved problem." Eventually it just answered "-2."
I told Gemini Pro I woke up after having dreamt that polynomial and asked if it was related to the Jacobian conjecture. It spent some time thinking and referenced this tweet announcement, saying:
"If you truly dreamt about that specific polynomial, you might be mathematically clairvoyant."
In the rest of the answer, it maintained a cautious skepticism about my claim, saying:
"Here is exactly why the math world is currently scrambling to verify the polynomial you "dreamt" about."
> Matches! This is bizarre. A Jacobian counterexample has been sitting here in a prompt? Wait... is this map a known "fake" counterexample from the literature?
Many mathematicians have tried and failed. This specific map might come from a paper or a forum where it was proposed and then debunked. Or... is it actually correct?
Gemma's having trouble accepting it too. A solution?! At this time of year? At this time of day? In this part of the country? Localized entirely within my own prompt?
Current LLMs behave very counterproductively around unsolved problems, especially if they learned that humans consider them difficult. This has many straight up preventing themselves from attempting anything...
I think vibethinker is heavily overtrained on not attempting to solve open problems.
I had a fun time taking some open problems and disguising them algebraically so that vibethinker 3b would work on them. It managed to prove some interesting things that I didn't know and would be publishable, but for the fact that they already have been. :) (though hard to know if this was because it had been exposed to that knowledge even though it didn't reconize the hidden problem).
Under some maskings it would eventually figure out the problem was equivalent to an open problem then immediately shut down.
It also managed to make some false proofs for various things that duped some other more powerful models.
More anecdata: When I just tried Qwen 3.6 27B (Q6_K_XL) it (ultimately, after a lot of going back and forth) claimed it was not a counterexample and claimed the Jacobian wasn't constant (which I'm guessing is incorrect). It also mentioned a whole bunch of names it attributed the example to, in its thinking trace.
Connecting it to a Coding Agent seems much better.
I connected DeepSeek in OpenCode and told it that I dreamed of this counterexample. It called SymPy tools to verify it, said my dream was "surprisingly accurate", and suggested consulting an expert in algebraic sets for independent verification.
I've heard about mathematicians going through kind of the same thing when they get a weird proof that ends up being right from some weird source or themselves.
Which is fair, they get inundated with kooky proofs from amateurs all the time and odds are incredibly good that there's some major fatal flaw that the amateur doesn't see. Or in the case of themselves, there's a certain blindness that makes it a little more difficult to critically evaluate your own leaps. In ether case the way it manifests is by going over it many times and many ways, each time more certain that you missed something until you just kind of break. Only then do you publicly start suggesting that there might be something to this new leap.
I read some thinking traces someone posted on X, and yeah, near psychosis from refusing to believe this simple of a solution had not been found already
sol medium can't believe its own input and output tokens either despite computing everything itself; this is what it gave me:
> Taken literally, these two facts would make this map a counterexample to the complex Jacobian conjecture in dimension 3: scaling one output coordinate would normalize the determinant to 1 without restoring injectivity. Since the complex Jacobian conjecture is still treated as an open problem, this strongly indicates that the displayed formula has been mistranscribed or contains a subtle typographical error.
I fed it to Google AI Studio, enabling tool execution and disabling web access. It also quickly verified it with SymPy, then went into psychosis.
5 minutes later: all previous chats are loading fine, but the only "Counterexample to the Jacobian Conjecture" chat is not loading.
Well, I'm not a conventional conspiracy theorist. But everyone knows that in every major LLM provider there are hell of hidden guarding systems that mark users and dialogues based on content (for topics about national security, biology, security, adult topics, etc.) - so there is a small chance a CEO of Google is now receiving a dozens of notifications about "ground-breaking results that could be attributed to Gemini, if act quick". So if any of thousands researchers have ever submitted this polynomial to Claude previously, any Anthropic employee can accidentally or intentionally "rediscover" the result of other researcher (and even hide the traces by deleting a dialogue of other user).
The great thing about these mathematical mopping up type operations is that no person will waste their time trying to prove it to be true anymore. If anything that’s a win.
It would be great if an LLM could settle the Collatz conjecture next, god knows how many man-years have been burned on that by unsuspecting victims.
The reason this was "easy" is because the conjecture turned out to be false. If the collatz conjecture holds true (and most mathematicians seem to think it will), it will be much harder to prove than your average Erdos problem.
I think parent’s point is that every false conjecture can cost a lot of time to be spent on futile affirmative proofs. So if we “clean up” a bunch of false conjectures, then more effort can be spent on interesting proofs of the others. (Probably a rather naive view of the value of conjectures but I’m just offering an alternative interpretation of the comment.)
I used to be a young mathematician sometime at the turn of the century...
"waste their time trying to prove it" is the MBA approach, where you should spit out results and articles.
Outside the MBA-thinking box, attacking hard problems, even unsuccessfully, is the way to gain deeper insight into various results and tools that you can later apply to other problems, i.e. no waste of time at all, unless you go to the extremes (like spending years on a single problem and nothing else).
Yeah but the Collatz probably has one of the highest man-hours of actual waste.
Ergo:
> “This is a really dangerous problem. People become obsessed with it and it really is impossible,” said Jeffrey Lagarias, a mathematician at the University of Michigan and an expert on the Collatz conjecture.
and
> “Collatz is a notoriously difficult problem — so much so that mathematicians tend to preface every discussion of it with a warning not to waste time working on it,” said Joshua Cooper of the University of South Carolina in an email.
I don't think this is how this works. The next step is to determine for what kinds of polynomials the jacobian conjecture is true and for what kinds of polynomials it's false.
It’s surprisingly easy to do with AI. The hard part has been manually verifying and validating the results. I took one of the smaller findings (disproving a conjecture) and wrote a paper as my first endeavor into publishing.
Because the next few findings i have in the pipeline are substantial in the field of quantum topology and physics im taking some time to publish them with a ton of scrutiny. And verification has taken more time than it did to make the discoveries.
I'm a phyicist by training and have published in the past but I left academia and so haven't for a number of years now.
I tried out using Claude to do some physics problem solving - mix of maths and simulation - and ended up with it getting in quite a mess. It's incredible at setting things up, suggesting approaches you might not have considered. I found it much much worse at interpreting things.
How much do you try to understand while doing it? I.e. how many levels of abstraction down in your own understanding do you go vs vibing at the surface level?
Using LLMs to generate piles of code and/or proofs of dubious quality is very questionable thing, and I understand these non-stop debates about it.
But in this case, as using plain brute force is already quite a common thing in searching for counterexamples, using LLMs as a sort of more advanced brute force seems to be just the right thing to do, so I struggle to understand so much hostility to this approach.
> I struggle to understand so much hostility to this approach.
It's going to be really rough for a lot of folks as machines get better and better at domains that the human brain was exclusively useful for. I tend to view the hostility as a mix of both "unless we have proof this could all be hokum" and "we're going to lose a lot of what we consider makes humans amazing".
Compassion is going to be very, very important in the next few years. We're going to hurt, both internally and externally.
The good thing about math is that you can formally verify the results, making hallucinations much less scary.
This is similar to games, where large learning models had some of their first big wins (go). Games, like math, have strict rules that can be used to prune invalid outputs.
It is already in the making, the next big step in math will be the complete formalization of all existing, relevant theorems and proofs via LLMs.
The poster works at Anthropic, so they likely have internal access to the next generation of Fable. Their internal model is probably an absolute beast at mathematics, and the upcoming benchmark results will likely set a new record for maths performance.
I suspect this is what happened, because the poster is coy about sharing the actual prompt / reasoning trace used to reach this result. That would be covered by an NDA until the model is properly released.
Sol is able to find the same counter example independently [1], so no reason to conclude in the existence of a benchmark destroying math beast Fable 6.
My Fable 6 theory is admittedly speculative, but your “[public GPT-5.6] Sol is able to find the same counterexample” is also a jump in conclusions. Aaron specifically says he used “an internal version of Codex”. When asked whether that meant a different model or harness, he dodged the question, and only said the harness should be the standard commercial GPT-ultra harness [0]. He (intentionally) avoids identifying which model was used, so your claim is similarly unresolved. Given that Aaron works at OpenAI and has access to internal models, that GPT-5.7 is expected to launch in a few weeks and is rumored to be 10T+ parameters, it's very plausible that Aaron used that model in his analysis.
Furthermore, public GPT-5.6 pro failed six times to find a disproof to the Jacobian Conjecture [1], even with hints, which is evidence against the claim that the public GPT-5.6 Sol can solve this.
Regarding the existence of Fable 6: an internal upgraded version of Fable or Mythos almost certainly exists, given that Anthropic has been testing Mythos internally since April and previously released new models roughly every ~6 weeks.
I'm not very excited. Access to the best AI is not a party I was invited to. And the people who are at that party, well, they don't exactly reflect on my best interests.
Is mathematics an science or an art? To the extent it's art, it's expressive and rewards the human experience that inspires and that creates it. If it's art, it's drastically less valuable to advance it through automation. But if mathematics is a science, then our entire goal is to increase humanity's understanding of the field. Whether by automation or genius inspiration or as a reward for decades of grinding it out incrementally, it's all the same: knowing more is the point.
Given how small the counterexample is... this feels like a great example of where a lot of interesting results are going to be found: not because they were super difficult, but because intelligence didn't scale, and until computers could do this for us, the number of people who seriously poked at many such things was low.
I'm very excited for the impact of this effect in science and medicine and other disciplines too.
I have a question I'm surprised people are not asking: How did Fable find this? Was it like guessing a bunch of families and then solving for possible solutions in those families? Was it clever search? something else?
Very short version: there’s an existing false counterexample in the literature which holds almost everywhere except at a pole. It looks like Fable used this polynomial as a base & extended it in a way that eliminated the pole whilst preserving the structure.
I'm going to paraphrase what GPT told me: Consider the canonical degree 3 (subvariety of the trivial P1 bundle consisting of zeros) cover of the projectivization of homogenous polynomials of degree 3 in 2 variables (so it's a 3fold cover of P^3). The top space is P1 x P2 and if you take a standard affine open of the base and look at the cover over that restricted to a subset where the zero of the cubic is simple you get the map for some choice of coordinates...
I honestly have no idea if it's correct lol I didn't check it (I should given I actually work in AG) but it doesn't look impossible at first sight
I’m doing this by working all logical steps into lean (formal verification) the quick feedback loop between the AI prose and the Lean verification errors and warnings ensures that its logically consistent.
The issue that remains are two things, ensuring the idea of the proof is actually the thing you want to prove and the interpretation of the results you get. But besides that, everything inside of the kernel checked code is logically consistent
In my experience, lean will show that it's correct, but does it not lose the mathematical intuition that led to the result? As far as my experience goes, that's really hard to encode in lean itself.
Could we maybe get more information about the problem from the LLM trace itself here?
Since I actually don't know math, maybe my ELI5 understanding can be helpful (or corrected).
The conjecture says that you can always reverse (a process) to determine the original inputs.
But this proof shows multiple inputs creating the same output - which obviously cannot be reversed to determine the input - thus falsifying the conjecture.
[...] you can always reverse (a process) to determine [...]
There is a precondition - constant non-zero Jacobian - to the inverse existing and the inverse is claimed to be of a specific kind - polynomial. The counter example satisfies the precondition and by mapping two different inputs to the same output makes any inverse impossible, including polynomial ones. But maybe that is already ELI7.
Wikipedia's gonna Wikipedia. Unless there's a material debate over the Jacobian Conjecture itself, there's really no open question here. This isn't a complicated proof; it's a straightforwardly checkable certificate of a solution.
This is nonsensical: Properness of the map is equivalent to its being an isomorphism (quick proof: Jacobian invertible implies that the map is etale, and properness would imply that it is finite etale, but affine space doesn't admit non-trivial finite etale covers), so the lack of properness is just another way of verifying that this is indeed a counterexample.
The author has a PhD in math from Cambridge. If it turns out to be a false claim it is an interesting case study on AI's sycophancy causing even experts to drop their guard and make mistakes.
Who do you mean? The author of the tweet is Levent Alpöge, who does not have a PhD in math from Cambridge... but does have a PhD in math from Princeton. And his advisor was Fields Medalist Manjul Bhargava.
Also, there is no way this counterexample is wrong. You can very easily check it for yourself. (I did, I don't know why, obviously Levent wouldn't be wrong about this, but I guess I was in shock.)
I'd rather wait for independent seasoned mathematicians to verify such claims first before someone at said AI lab posting a claim about solving a proof online.
Let this be a lesson to those who fell for such AI psychosis and to not believe everything you see on the internet as real.
It's a Princeton math PhD who posted. The verification is quite straightforward and was posted by the tweet author. Wolfram would have to also be producing incorrect outputs for the counterexample to be false. The counterexample works as claimed and conjecture has been proven wrong.
This topic really is a testament to people's willingness to opine on things they have absolutely no clue about.
A first year undergraduate can completely check this counterexample in ten minutes. The original post even linked Wolfram alpha for the calculations.
And if you genuinely try you can very quickly understand using only high school math and a bit of Wikipedia that this counterexample is vanishingly unlikely to be wrong, even if you don't do the calculations yourself.
This is so unreasonable! As @__alpoge__ himself notes this is classic crank graveyard territory and yet the counter example is something a grad student in 1997 could have found w a ~3 day computer search. Wild!
The search space for a naive brute force of three polynomials of degree <= 7 with integer coefficients <= 12 is roughly 10^500. I think it would take a little longer than that.
> and yet the counter example is something a grad student in 1997 could have found w a ~3 day computer search
Is that true? Even restricting this to f(x,y,z) and coefficients and powers to 1 ≤ x ≤ 10, there are a lot of polynomials to check, and checking requires checking the Jacobian determinant and, if it’s a non zero constant, finding two points for which the polynomial produces the same value.
Or is there a way to generate all polynomials with a non-zero Jacobian determinant, and does that speed up things? (My intuition say it wouldn’t, because I guess those with zero determinants are rare)
Maybe, maybe not. The "proofs" may not have helped at all with finding a counterexample. Either way, it doesn't matter. A counterexample was found, no one found one before even though clearly a lot of people have tried who also had access to the prior "proofs".
I think it's becoming harder and harder to argue that LLMs don't really reason and just mimicry human speech. This counterexample is clearly the result of a sequence of steps that build on previous knowledge in context and logically combine it to reach other true statements - to a degree and complexity that rivals the best human minds.
For someone that use Claude Code every day, this is obvious, but for some reason many scientists refuse to accept that it's truly reasoning; perhaps not in the human sense, but in a very profound and real sense. These powerful results are devastating to their point of view.
I can sympathize, because I too called LLMs "fancy Markov chains" in the GPT 3 era. But there comes a time where you have to update your world view to match reality, or be stranded in fantasy land.
In my view, it's theoretically possible for a combination of the author's iterative prompts + evaluation with Wolfram Alpha to activate the weights that encode the language that describes the constraints on these polynomials (from the faulty proofs) in such a way that the author eventually arrives at this:
> ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3
I've been asking Fable very complicated questions in heterodox economic theory, which is something I know a lot about. The stuff it comes back with is incredibly deep.
To give a metaphor that everyone here on HN would understand, reading it's responses gives me the same level of wonder as one gets learning how quicksort works for the first time. It even stretches my brain to grasp what it's even come up with. I find myself getting mentally exhausted just digesting it's brilliance.
I think the singularity will have this point where AI comes up with ideas so profound, like a Ramanujen equation, that the most brilliant among us can't even decipher the answer to our questions. The internal reasoning of the machine is at a level of complexity that's beyond human comprehension to even keep track of everything enough to integrate the understanding of what it's come up with. This will happen with any even mildly complex question about any topic.
I don't consider myself an open weights supporter, but it's a bit of a bummer that if there's any novel search technique discovered by the model throughout this finding, it's possibly locked behind ant's reasoning summarization.
One wonders if they could turn their mechinterp work into analyzing the "thought processes" of these very special cases that turn into novel research and finally crack the creative thinking barrier.
>Imagine you had a frozen [large language] model that is a 1:1 copy of the average person, let’s say, an average Redditor. Literally nobody would use that model because it can’t do anything. It can’t code, can’t do math, isn’t particularly creative at writing stories. It generalizes when it’s wrong and has biases that not even fine-tuning with facts can eliminate. And it hallucinates like crazy often stating opinions as facts, or thinking it is correct when it isn't.
>The only things it can do are basic tasks nobody needs a model for, because everyone can already do them. If you are lucky you get one that is pretty good in a singular narrow task. But that's the best it can get.
>and somehow this model won't shut up and tell everyone how smart and special it is also it claims consciousness. ridiculous.
Smarter in some ways at least. Probably still not quite as smart at understanding human emotions (and things that aren't well enough catalogued on the internet or amenable to Reinforcement Learning on virtual environments).
Hot take: even GPT-3 was not a parrot. Skeptics have never properly internalized that the fundamental operation is basically irrelevant to the gestalt. Humans are not parrots yet neurons likely also largely operate via predictive processing.
> Jacobian conjecture [...] states that if a polynomial function from an n-dimensional space to itself has a Jacobian determinant which is a non-zero constant, then the function has a polynomial inverse.
> ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
But 1 != -1 and -3/2 != 3/2 . So it's not its own inverse. Is the conjecture that it is its own inverse or that is has an inverse?
Edit: it was worded a bit strangely, but it is saying that [ (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) ].map(F) all produce (-1/4, 0, 0). Thus it has no inverse and indeed disproves the Jacobian conjecture.
I think the map sends (1, -3/2, 13/2), -> (-1/4, 0, 0) and also (-1, 3/2, 13/2) -> (-1/4, 0, 0) so it's not invertable which disprove the jacobian conjecture that polynomial maps over complex numbers with a jacobian that's non-zero are globally invertible.
(Just as a note for myself, I had to think of why the fact that such jacobians are constant is a byproduct, I guess it's because of lioville's theorem implying that any polynomial over C that never hits 0 must be a constant [because the reciprocal is bounded and thus must also be a constant])
For all thehubbub, as far as I know, all the math breakthroughs via AI that I've heard about have come from Anthropic and OpenAI, not Chinese models. I could have missed those announcements, but one might think that between close to frontier performance plus cheap tokens, that they'd be leading the way on these things.
All (unless I've missed one?) the math breakthroughs are also coming from the relatively small number of mathematicians working at these companies, as opposed to the many orders of magnitude more mathematicians using LLMs for mathematics outside of the companies. I assume the missing link everywhere is being able to casually burn a few rainforests worth of tokens in pursuit of something publishable.
Depends on exactly where you draw the line at "breakthrough", but there's been at least a few novel and interesting results coming from outside mathematicians. Recently for example there were https://old.reddit.com/r/math/comments/1uxj3cy/after_openais... from Phillip Kerger at Berkeley, and https://www.erdosproblems.com/forum/thread/119/proof-claims from Samuel Korsky at Two Sigma (the latter of which was more of a collaboration between the human and machine, not a one-shot like several of the other results we've discussed).
No. Most of the results are from 3rd parties using publicly available models. The most impressive results have come from direct announcements, but most in total have not. Open AI have only announced 2 results and the total is up to a dozen or so now.
I am a number theorist and a graduate math student at Bonn. It is eery to see that all of a sudden everyone cares about pure math. Anyhow, I think such incidents like this one are only to happen more often in the coming days, and while many mathematicians are concerned about their role in the community; I believe since we are still very early stages of AI-driven discovery, we still need the appropriate infrastructure for human-AI research, such as with versioning for proofs, a bigger library than mathlib and a central reasoning space so traces aren't lost- which happens to be the most valuable training data for models, and without which, these discoveries do not help much at advancing the field, and remain as blackbox.
> without which, these discoveries do not help much at advancing the field,
At least the fact that frontier models are not optimized primarily on formal math reasoning is in itself is good for not putting mathematicians out of their jobs, isn't it?
How do these results (and the future painted by them) affect your profession; do you expect a similar route as in software, where junior developers are unemployable, LLM-assisted development is the norm, and that great developers stand out partly through better communication with their managers?
I think in short term, it will get better, as more funding flows from industry to academia, you will see a spur in well-compensated PhD positions. However, in the long run, while I doubt that mathematicians will go extinct, they might have to move on to industries like in trading, chip-making, where they oversee AI models write code and also proof that the code works and is consistent with other parts of the software. However, until we have infrastructure for human-AI math research, as Tao himself said that the "roads" for human-AI research is yet to be built, humans would simply be working FOR AI models, and not working WITH AI, which will truly scale discoveries at a massive scale.
The funniest thing about LLMs is the cognitive dissonance they cause people. People clearly recognize (and bemoan) the fact that LLMs produce derivative breathless prose ie they fundamentally fail at "unstructured creativity" (something the might accurately labeled intelligence) but are then shocked that the same LLMs can do math.
It's reasoning from a flawed premise that math universally requires intelligence and creativity. It does not. Anyone that's proved things via "diagram chasing" can affirm that. The conclusion you should draw is that math (at least the kind they excel at) isn't actually a creative endeavor.
I would think it's more that some aspects of math (like anything really, including programming) don't require creativity and can be solved through "brute force" or whatever you want to call what LLMs do. But it's pretty obvious that LLMs are not capable of solving the vast majority of problems in math (or programming) at this time. Eg. Google went 9/353 on Erdos problems. If a more powerful LLM is capable of solving those or if they require a certain je ne sais quoi of the human variety is still up in the air at this point. In either case it seems like they require a long and detailed prompt from a domain expert (ie. human) regardless.
Didn't humanity also score very low on 353, namely 0 since they were open? Probably collectively we could have gotten a slightly better score if all of Math started trying to solve those, but not by that much, I think, since they are precisely still open.
I dunno, this screams of goal post moving. Even if LLMs lack whatever nebulous definition of "creativity" that someone favours, there's no inherent reason for "creativity" to be required to solve any problems at all, "creativity" could just be a human method for solving problems that evolved because of it's broad applicability but is suboptimal at any given task.
> could just be a human method for solving problems that evolved because of it's broad applicability but is suboptimal at any given task.
Ya sure let's just posit another random hypothesis about evolutionary biology in order to substantiate the claim that LLMs are intelligent.
Or (bear with me) you can recall your (likely) experience proving stuff like SAS triangle identities and reflect on whether that required intelligence or just computation.
I think the real tragedy is that they are good at writing, just by default are tuned to have a kind of bland corporate tone. If you give the LLM a few pages of writing you like and tell it "Continue, but using this style" it will do a pretty good job of it. Most people just .... don't bother to do that.
If there really was a "simple" solution to Fermat's Last Theorem, Andrew Wiles wouldn't have achieved the important result he did, ending up making connections across disparate fields of math.
The LLMs "sweeping up" easy, or previously missed, results seems like a net negative. It's probably better for humans to struggle and come up with new tools than to just "clean up" low hanging fruit that doesn't add much value to the field.
Are you at least a little familiar with linear algebra?
If so, you've probably heard of the determinant. It's a certain way of "summarizing" a matrix with one value.
The determinant in this case is of the Jacobian, which is a matrix you can construct from a multi-variable function. Each term is the partial derivative with respect to each variable (x, y, z, etc.), with one line per output variable (vector element).
The Jacobian of a polynomial function is, in general, going to be a matrix where every term is some polynomial expression. And the determinant of that will also be a complicated expression. But in some cases all the variable terms cancel out and you're left with a single constant (0 or some other value).
The conjecture says that if the Jacobian determinant is constant (i.e., all the terms cancel out), then there must be a polynomial inverse. And the key condition for an inverse is that there must not be two input points that evaluate to the same output. It's just like y=x^2. It's not invertible, because both +2 and -2 square to +4.
So if you can find a function where the Jacobian determinant is constant and also find two or more points that evaluate to the same output, then you've found a counterexample to the conjecture. And that's what's been done. And remarkably, the counterexample is pretty simple. It would be tedious but a bright high school student could verify it.
I think another way to understand it is the generalization of the inverse function theorem. The inverse function theorem gives you local invertibility, but even being "locally invertible" everywhere does not imply global invertibility (you don't need too pathological an example to see this, a periodic function serves iirc).
The Jacobian conjecture roughly asks what whether local invertibility gives you global invertibility when you restrict only to polynomials (which we might hope "behave nicely"). Apparently for polynomials over reals this was disproved a while back, but up until now the general case of polynomials over complex numbers was open.
All the coefficients and evaluation points are rational, so it's a counterexample in all fields where 2 ≠ 0 and 3 ≠ 0, doesn't matter whether that field is the complex numbers, real numbers, rational numbers or even a finite field.
Why do you say this? I've admittedly never done a proper complex analysis course but I got the impression that that complex differentiability was a very strong condition that results in holomprhic functions behaving "nicely" in ways that real functions do not
Should have used quotes.. I didn't mean it in any formal sense. What I am saying the nature of unit in complex plane makes it difficult to intuitively imagine invertibility and determinants.
Speaking as a mathematician, it does seem like we're a bit fucked as a community. Anything that is at all accessible to currently existing methods and mathematical infrastructure is probably going to fall to the frontier models of today, and at this rate of progress it's likely that, already by next year, we'll see new infrastructure being put into place by AI, giving us a world in which a few designated interpreters of the oracle get to 'do' mathematics, while it withers on the vine as an avenue for the exploration of human meaning.
Proving theorems will have lower payoff, but posing new questions (for AI to chew on) will have higher payoff. Math will go from theorem proving to conjecture farming/exploration. In a way this could be even more fun.
Of course AI can also farm conjectures, but they have to develop taste, which might be harder than just proving theorems.
Yes, as someone who prefers developing the 'correct' structure over 'merely' proving theorems, this is good for me in the short term. However, the writing I fear is on the wall for my medium and long term utility.
I think in the long run mathematicians are probably fucked, but in the short run it's not that bad. All three of the big conjectures solved the answers were at the level where if you had given a grad student the questions and the right background reading there's a good chance they would have solved it. (This example, you could have given an undergraduate good at programming and computer algebra and told them to come up with a counterexample.)
At this point the advantage of AI is that it's read the entire mathematical literature, and it doesn't have to worry about wasting its time. The solved problems have all turned out to be surprisingly easy, so the real lesson is that we're bad at judging how hard problems are.
Assuming this state of affairs lasts, the medium-term problem is that you learn something when struggling with a problem, even if you don't solve it, and if mathematicians become too reliant on AI the skills they develop through struggle will erode.
The long-term problem, of course, is that it seems much more probable that a future model will make mathematicians all obsolete. But so far Fable hasn't. (Anthropic has probably burned a billion tokens on the Riemann hypothesis already, without telling anyone.)
> This example, you could have given an undergraduate good at programming and computer algebra and told them to come up with a counterexample
please try go try it. There's no way someone didn't do massive computer algebra searches before today.
> All three of the big conjectures solved the answers were at the level where if you had given a grad student the questions and the right background reading there's a good chance they would have solved it.
You cannot be serious... why didn't they solve it before then? Do you think no one tried it? What background do you give the double cycle conjecture student after the flow reduction? a linear algebra textbook???
Anybody that has to work for a living is fucked and not on the "can't do mathematics which they would find fulfilling"-level but on the "can't afford food, because human intelligence is simply not required anymore"-level.
You can't please everyone on the internet, if he had posted this on Arxiv then someone else will be complaining that Arxiv is for humans to publish and that llm output should be social media post instead. Also, the proof fits in a tweet. So why blow it up.
"Any idiot could have done this, it's just high school calculus and just a counterexample anyway. Stochastic parrot, spicy autocomplete, AI psychosis. Wake me up when an AI does something real."
The goal posts have moved. People generally stopped saying this stuff now.
Even if you go to the ultimate anti-AI subreddit r/betteroffline, they've changed from "AI is useless" to "AI is good but the AI bubble will collapse soon" over the last 6 months.
I think we’re not adapted well for this rapidly changing world. Here you have some people who were rightly skeptical about a new technology being shoved down their throats by giant tech corporations, and a technology that really was, and probably still is overhyped.
Yet it’s a technology which has rapidly grown in its capabilities.
So yeah now many of the people who thought it was useless before probably don’t think it’s useless anymore, but you’re holding them to their original words even though those words were about something completely different at this point.
If people aren’t saying it anymore it might be because they don’t think that anymore, and the people who have new goal posts might be entirely different people.
It’s like you’re looking at a different set of goal posts on a different field and saying, no! The goalposts have moved!
> I hold my stance that LLMs are stochastic parrots... Making the parrots ever more complex and training
> Except solving problem is probably the least (even though it's important) interesting thing in research.
> Can we use AI to get a cure for cancer yet? Or is math-turbation the only thing these things are good for?
> Train on enough examples and statistical autocomplete gets you places. I'm surprised how anyone would even consider this intelligence?
And, as much as HN has declined in the grips of an anti-AI psychosis, Reddit is worse. I would love if social fora would switch to the reasonable claim that we're in a bubble; that's something that can be debated. That's not the dominant critique of AI, though.
I don't think the (fairly factual) description of these systems as stochastic parrots means that they will never do useful work, just that they are not intelligent in the way we believe animals to be (to "push back" on your anecdata, I've also heard fewer people claiming that LLMs are actually conscious in the past year -- maybe we're reaching the happy medium?). That was the point the stochastic parrots paper and Chinese room thought experiments were making -- nobody claimed that the man in the Chinese room would be unable to accurately translate Chinese text.
Fuzzers are another kind of stochastic generator but nobody would claim they don't do useful work in a way that is hard to replicate through deterministic methods. (I still find the code these models produce kind of awful, but advancements in harnesses do mean that they can finally produce code that works most of the time.)
To be fair to the sycophancy tendencies, this was an open conjecture that held for 85 years, and not for lack of trying. So, maybe a bit warranted here? :)
The interesting thing about using Claude Fable 5 is it's nearly as irritatingly sycophantic as past Claudes while genuinely being smarter than the previous models. So you get a kind of yo-yoing of it glazing you as a creative genius and disappointedly revealing to you that your ideas are bad and dumb.
(NGL I wanted to suggest someone to go for the JC using 5.6 after the CDC proof came out, but then on reflection felt I should neither waste people's time NOR contribute to the myth of AI :)
My prediction is that the bubble will burst in 2031 Q4, one year after the Riemann Hypothesis is expected to fall (according to Demis)
After 2031, I will suggest going for the JC for N=4 because they would (dis)prove the Dixmier conjecture for N=2 :)
Assuming you mean C^2 -> C^2, Do you have a link? If so it would be good to add to the wikipedia page. Also I'm not sure, but does the fact that there's a disproof for n=3 imply that it's false in all n>=3, or could there be higher dimensions where it still holds (I'd guess not since you could probably trivially "embed" this in higher dimensions in some way)
Counterexample to Jacobian conjecture:
((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
GPT wrote some SymPy code to check it. The response?
"As written, this is an explicit counterexample to the Jacobian conjecture. I checked it using exact symbolic algebra.
I do not see an algebraic catch in what you typed. Unless a term or exponent differs from the intended expression, it appears to disprove the conjecture. This deserves serious independent checking rather than casual dismissal."
Reading the chain of thought for Fable, it is incredulous as well. It keeps thinking that it must be missing something or that this is a trick, because it can't have just found a counterexample to a famous conjecture. It's just like us!
Same awnser as much of the LLM Proofs - people cared about other things. There isn't a lot of money in academic math, and the ones that love it don't look for low value findings. Proofs like these are, funnily enough, usually the domain of hobbyists - but over the last few years, the "Monetize everything" mentality and struggling first world economy has pushed people away from interesting academic pursuits on their free time.
What does this mean? Fallacious argumentation and deceptive rhetoric is acceptable, if the topic is sensitive enough / there is enough riding on a wrong answer being accepted?
>> hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final
So where does it say that Fable "produced" the counterexample? The tweet says it was a collaboration between two people, using Fable.
I’ve been staggeringly productive with Fable. Opus 4.8 fails a lot more for me.
Fable often just “knows” what I want with vague instructions. It also is able to autonomously perform work that lasts an hour long from my experience. I haven’t tested further.
Without Fable included in subscriptions, I would have moved my entire team over to Codex 5.6.
A counterexample to the Jacobian Conjecture - and also a counterexample to “AI will never be smarter than humans.” Even the most dyed-in-the-wool AI hater at this point must acknowledge that it is more intelligent than any human. The other day I found out that Fable could read seal script! The small seal, standardized stuff no prob, but I found it even did OK at the hardest you can get, Warring States regional scripts. That’s something maybe 2,000 academics worldwide can do and nobody’s even talking about it because it’s just one more item in a very long list.
It’s a strange feeling, to be overtaken by our own creation. Top dog for millions of years and then in the blink of an eye we go from “how many Rs in strawberry” to this.
A little over 10 years ago I remember meeting a postdoc who believed he had something close to a counterexample to the Jacobian Conjecture. He and another person was bruteforcing polynomials in about 16 variables, something like 80 - 700 terms each, using binary trees for mapping coefficients.
They were guessing, at the time, that the lower bound of a counterexample (P, Q) for max(deg(P), deg(Q)) would go up to 200.
To think that Claude Fable was able to find a counterexample in degree 7 is insane to me. We are truly in a new era.
you might be confusing 2 variable case (which indeed was tested to 150+ degree) and 3 variable case (this counterexample)
Would this counterexample not be included in their search space?
They were looking at 16 variables, degree of about 10 in each one. That search space is simply too big for a plain bruteforce. So they did some sort of filtering to reduce the search space to a pool containing "possible counterexamples".
There was also a paper giving a lower bound of about 100 for possible counterexamples in that particular framework. Later raised to 108 in https://arxiv.org/abs/2204.14178
I don't remember the details very well since this was back in 2015 and wasn't really involved in the research. Consider this to be some sort of telephone game between what I heard in 2015 and what I remember today.
based on the approach described they were probably in dimension 2, this counterexample is 3 dimensional
2 replies →
[flagged]
So many mathematicians over the years tried hard and failed, but now Anthropic just for some PR magically did it? And this after LLMs obtaining different math wins? What is your logic here really escapes my understanding.
77 replies →
if you ask the chatbots for "list of top unsolved math problems", the JC comes in at a ranking of around #10 - #20. what, a problem that's been unsolved since 1939 was cracked because anthropic has an underground sweatshop of math Phds cranking out research, just so that they can slap "made by AI" on it? hell, maybe lizard people did it.
This anti-AI sentiment is getting borderline insane.
1 reply →
Finding a counterexample that humans have failed to find for 85 years is a marketing stunt?
[flagged]
10 replies →
> The incredible Yitan Zhang (https://newyorker.com/magazine/2015/02/02/pursuit-beauty) worked on proving this conjecture for 7 years. Moh, his advisor, wrote that Zhang "failed miserably" in proving the Jacobian conjecture, "never published any paper on algebraic geometry" after leaving Purdue, and "wasted seven years of his own life and my time".
https://x.com/aminkarbasi/status/2079129649830137989
https://en.wikipedia.org/wiki/Yitang_Zhang
Zhang, Yitang’s life at Purdue (Jan 1985-1991) T.T.Moh - https://www.math.purdue.edu/~ttm/ZhangYt.pdf
kind of a wild document to exist...
This doc is insane.
Moh repeatedly:
calls Zhang’s words “fake” and his claim “a lie”;
speculates, without demonstrating it, that Zhang “fooled” professors to get admitted;
says Zhang “failed miserably” and “wasted seven years of his own life and my time”;
alleges that Zhang “wanted to be famous all the time”;
publishes an irrelevant and humiliating story about Zhang not attending his father’s funeral;
says he would not touch Zhang “with a five feet pole”;
repeatedly emphasizes his own generosity, influence, mathematical work and supposed sacrifices.
The world doesn’t need advisors like Moh that are bitter enough to gossip about a family funeral.
2 replies →
The document says more about the kind of advisor Moh was than the kind of mathematician Zhang is.
Interestingly: "For logic proof, it had been thought that AI could handle all logic problems in the near future, hence logic problems of solving conjectures might not be so interesting in the future."
This is a rare instance where feeding this groundbreaking information into an LLM gives _them_ psychosis. I fed this to claude code and watched it verify the result in 7 different ways to be 100% certain, and it was just flabbergasted. Quite remarkable.
I fed ChatGPT the map with no other context, just “tell me about this function”. It did a bit of work finding the Jacobean etc and eventually worked out the implications of what it was seeing. It then proceeded to check the arithmetic 4 times, and then decided to do a manual verification using an ad hoc symbolic checker in case its SymPy had been tampered with.
Skepticism is the flipside of knowledge.
If an LLM has knowledge encoded inside it (and it's hard to argue it doesn't), then cognitive dissonance can be experienced. And once experienced, must be dealt with, especially in longer-running agentic loops.
A friend was joking the other day about sending some messages under a previously-used Slack identity for an agent (since turned off), then asking the agent about the messages.
The agent maintained it hadn't sent those messages (no memory) and then was forced to reconcile the idea that the messages indeed appeared to come from it.
Its extremely-agitated conclusion was that there had been a security breach and the entire network should be locked down.
1 reply →
Can you please share the chat log for this? I would absolutely love to see this.
4 replies →
Same result. Public share https://claude.ai/share/19fd1a34-d63b-4a16-8d83-60d5b79e7747
It did the multiple verification sequence before expanding to internet search where it found this thread.
> So the conjecture that survived Keller, Abhyankar, Moh's degree-100 verification, and five-plus published wrong proofs appears to have died via tweet during the World Cup final.
3 replies →
Everyone using Claude Fable to verify this proof is so funny. If you read the definition of the Jacobian Conjecture and (I am not exaggerating this) have passed a college Calc 3 class, you can just verify the proof yourself in 30 seconds. The problem was very hard to solve but the counterexample is very easy to verify!
--- edit, adding an explanation:
To summarize it, the conjecture says if you have any multi-variable polynomial function that maps an input to an output in the same dimensional space (take for example: F = (x+2, y+2), which maps 2D space into another 2D space), AND that function has a constant-valued non-zero Jacobian determinant, THEN the conjecture is that the polynomial has an inverse, meaning basically you can find a polynomial that turns the output space back into the input space.
Fable provided the example polynomial (which was very hard to do) and the coordinates which if you plug into it, results in two points being mapped to the same output point. This means that the polynomial can't be inverted, because if you have that output point, how do you know which input point it came from?
You can just plug in the two coordinates it gave into the equation and verify that you get the same output point from both. That's the contradiction of the conjecture and it takes 30 seconds.
---
Something something outsourcing of thinking something.
2 replies →
I like how first it's amazed and doesn't believe it, accepts it, then realizes you not only stole it from twitter but that it's its own proof lmao.
Which exact model does claude use here?
1 reply →
“I'll admit the sequence on my end was genuinely disorienting: I verified the determinant three separate ways looking for the error, verified the evaluations twice, ran out of subtleties to check, and only then searched and discovered that the reason it holds up is that it's apparently my own homework — the "fable" in that tweet is Claude Fable, i.e., this model, working with Alpöge.”
[flagged]
Who knew DOES NOT COMPUTE would be an actual thing?
The unexpected part is it does compute!
1 reply →
Captain Kirk defeated multiple evil computers this way
It reminds me of how AI will estimate that a coding project will take "4 weeks" and then proceed to finish the task itself in 15 minutes.
AI models are changing the world much faster than their own training can keep up with.
Can confirm, Claude is flabbergasted.
Gemini just checks the web first it seems, and already references the news.
Kimi doesn't quite believe it.
turn off search
>kimi is having a blast. i turned search back on and found this post from it’s sources cited after i suggested to check out the reaction. best thing is to go to a model with search off and plop it in the session
Deepseek Pro got stuck after munching it for a minute or so and couldn't quite believe it either.
Like the unicorn emoji, but for math? It occurs when the LLM is presented with incontrovertible evidence against something it "deeply believes" to be true.
So, basically human psychology?
Interestingly, even Qwen 3.6 27B was able to verify the solution, but I didn't get any glazing for discovering it. Instead, it thought that someone named Shestakov had already found a counterexample in 2004.
GLM 5.2 whiffed, it insisted the counterexample wasn't valid.
VibeThinker 3B also recognized that the counterexample was valid. But it kept trying to convince itself that it wasn't, over and over, since it's an "unsolved problem." Eventually it just answered "-2."
I told Gemini Pro I woke up after having dreamt that polynomial and asked if it was related to the Jacobian conjecture. It spent some time thinking and referenced this tweet announcement, saying:
"If you truly dreamt about that specific polynomial, you might be mathematically clairvoyant."
In the rest of the answer, it maintained a cautious skepticism about my claim, saying:
"Here is exactly why the math world is currently scrambling to verify the polynomial you "dreamt" about."
I love how it put "dreamt" in quotes.
1 reply →
> Matches! This is bizarre. A Jacobian counterexample has been sitting here in a prompt? Wait... is this map a known "fake" counterexample from the literature? Many mathematicians have tried and failed. This specific map might come from a paper or a forum where it was proposed and then debunked. Or... is it actually correct?
Gemma's having trouble accepting it too. A solution?! At this time of year? At this time of day? In this part of the country? Localized entirely within my own prompt?
2 replies →
Current LLMs behave very counterproductively around unsolved problems, especially if they learned that humans consider them difficult. This has many straight up preventing themselves from attempting anything...
1 reply →
I think vibethinker is heavily overtrained on not attempting to solve open problems.
I had a fun time taking some open problems and disguising them algebraically so that vibethinker 3b would work on them. It managed to prove some interesting things that I didn't know and would be publishable, but for the fact that they already have been. :) (though hard to know if this was because it had been exposed to that knowledge even though it didn't reconize the hidden problem).
Under some maskings it would eventually figure out the problem was equivalent to an open problem then immediately shut down.
It also managed to make some false proofs for various things that duped some other more powerful models.
1 reply →
More anecdata: When I just tried Qwen 3.6 27B (Q6_K_XL) it (ultimately, after a lot of going back and forth) claimed it was not a counterexample and claimed the Jacobian wasn't constant (which I'm guessing is incorrect). It also mentioned a whole bunch of names it attributed the example to, in its thinking trace.
1 reply →
Connecting it to a Coding Agent seems much better.
I connected DeepSeek in OpenCode and told it that I dreamed of this counterexample. It called SymPy tools to verify it, said my dream was "surprisingly accurate", and suggested consulting an expert in algebraic sets for independent verification.
1 reply →
> that someone named Shestakov had already found a counterexample in 2004.
Qwen has the sprit of a grad student
2 replies →
I've heard about mathematicians going through kind of the same thing when they get a weird proof that ends up being right from some weird source or themselves.
Which is fair, they get inundated with kooky proofs from amateurs all the time and odds are incredibly good that there's some major fatal flaw that the amateur doesn't see. Or in the case of themselves, there's a certain blindness that makes it a little more difficult to critically evaluate your own leaps. In ether case the way it manifests is by going over it many times and many ways, each time more certain that you missed something until you just kind of break. Only then do you publicly start suggesting that there might be something to this new leap.
> they get inundated with kooky proofs from amateurs
Source?
2 replies →
I read some thinking traces someone posted on X, and yeah, near psychosis from refusing to believe this simple of a solution had not been found already
This is also just an annoying Claude personality trait
sol medium can't believe its own input and output tokens either despite computing everything itself; this is what it gave me:
> Taken literally, these two facts would make this map a counterexample to the complex Jacobian conjecture in dimension 3: scaling one output coordinate would normalize the determinant to 1 without restoring injectivity. Since the complex Jacobian conjecture is still treated as an open problem, this strongly indicates that the displayed formula has been mistranscribed or contains a subtle typographical error.
quite interesting indeed!
Just take it as more confirmation that LLMs are unintelligent pattern-matchers.
5 replies →
[flagged]
[flagged]
I fed it to Google AI Studio, enabling tool execution and disabling web access. It also quickly verified it with SymPy, then went into psychosis.
5 minutes later: all previous chats are loading fine, but the only "Counterexample to the Jacobian Conjecture" chat is not loading.
Well, I'm not a conventional conspiracy theorist. But everyone knows that in every major LLM provider there are hell of hidden guarding systems that mark users and dialogues based on content (for topics about national security, biology, security, adult topics, etc.) - so there is a small chance a CEO of Google is now receiving a dozens of notifications about "ground-breaking results that could be attributed to Gemini, if act quick". So if any of thousands researchers have ever submitted this polynomial to Claude previously, any Anthropic employee can accidentally or intentionally "rediscover" the result of other researcher (and even hide the traces by deleting a dialogue of other user).
> all previous chats are loading fine, but the only "Counterexample to the Jacobian Conjecture" chat is not loading.
This happened multiple times to me with Gemini. For the most trivial of requests, like translating a video into English.
> "ground-breaking results that could be attributed to Gemini, if act quick"
This would be such a dumb thing to do, and so easy to get caught with...
7 replies →
Well, Gemini is just a bad model in general, you have two variables going on.
>flabbergasted
phatic mimicry.
The great thing about these mathematical mopping up type operations is that no person will waste their time trying to prove it to be true anymore. If anything that’s a win.
It would be great if an LLM could settle the Collatz conjecture next, god knows how many man-years have been burned on that by unsuspecting victims.
The reason this was "easy" is because the conjecture turned out to be false. If the collatz conjecture holds true (and most mathematicians seem to think it will), it will be much harder to prove than your average Erdos problem.
I think parent’s point is that every false conjecture can cost a lot of time to be spent on futile affirmative proofs. So if we “clean up” a bunch of false conjectures, then more effort can be spent on interesting proofs of the others. (Probably a rather naive view of the value of conjectures but I’m just offering an alternative interpretation of the comment.)
9 replies →
True but Noam brown (openai researcher) said that in 2 years AI will start creating new math.
5 replies →
Give the llms a few years, they'll be smart enough to make progress on that.
I used to be a young mathematician sometime at the turn of the century...
"waste their time trying to prove it" is the MBA approach, where you should spit out results and articles.
Outside the MBA-thinking box, attacking hard problems, even unsuccessfully, is the way to gain deeper insight into various results and tools that you can later apply to other problems, i.e. no waste of time at all, unless you go to the extremes (like spending years on a single problem and nothing else).
Yeah but the Collatz probably has one of the highest man-hours of actual waste.
Ergo: > “This is a really dangerous problem. People become obsessed with it and it really is impossible,” said Jeffrey Lagarias, a mathematician at the University of Michigan and an expert on the Collatz conjecture.
and
> “Collatz is a notoriously difficult problem — so much so that mathematicians tend to preface every discussion of it with a warning not to waste time working on it,” said Joshua Cooper of the University of South Carolina in an email.
(from https://www.quantamagazine.org/mathematician-proves-huge-res... )
We won't be running out of hard problems to attack anytime soon, even if we produce a bunch of counterexamples to some of them.
I don't think this is how this works. The next step is to determine for what kinds of polynomials the jacobian conjecture is true and for what kinds of polynomials it's false.
I’ve been math-vibe coding a few months now.
It’s surprisingly easy to do with AI. The hard part has been manually verifying and validating the results. I took one of the smaller findings (disproving a conjecture) and wrote a paper as my first endeavor into publishing.
Because the next few findings i have in the pipeline are substantial in the field of quantum topology and physics im taking some time to publish them with a ton of scrutiny. And verification has taken more time than it did to make the discoveries.
Here’s my first piece if anybody is interested in number theory: https://arxiv.org/abs/2607.09793
I'm a phyicist by training and have published in the past but I left academia and so haven't for a number of years now.
I tried out using Claude to do some physics problem solving - mix of maths and simulation - and ended up with it getting in quite a mess. It's incredible at setting things up, suggesting approaches you might not have considered. I found it much much worse at interpreting things.
How much do you try to understand while doing it? I.e. how many levels of abstraction down in your own understanding do you go vs vibing at the surface level?
i'm doing the same thing but for type theory
This is very interesting to me. Care to share your process?
2 replies →
Using LLMs to generate piles of code and/or proofs of dubious quality is very questionable thing, and I understand these non-stop debates about it.
But in this case, as using plain brute force is already quite a common thing in searching for counterexamples, using LLMs as a sort of more advanced brute force seems to be just the right thing to do, so I struggle to understand so much hostility to this approach.
> I struggle to understand so much hostility to this approach.
It's going to be really rough for a lot of folks as machines get better and better at domains that the human brain was exclusively useful for. I tend to view the hostility as a mix of both "unless we have proof this could all be hokum" and "we're going to lose a lot of what we consider makes humans amazing".
Compassion is going to be very, very important in the next few years. We're going to hurt, both internally and externally.
Without seeing the Fable trace that produced this, claiming it amounts to ‘advanced brute force’ seems like a leap.
The good thing about math is that you can formally verify the results, making hallucinations much less scary.
This is similar to games, where large learning models had some of their first big wins (go). Games, like math, have strict rules that can be used to prune invalid outputs.
It is already in the making, the next big step in math will be the complete formalization of all existing, relevant theorems and proofs via LLMs.
The poster works at Anthropic, so they likely have internal access to the next generation of Fable. Their internal model is probably an absolute beast at mathematics, and the upcoming benchmark results will likely set a new record for maths performance.
I suspect this is what happened, because the poster is coy about sharing the actual prompt / reasoning trace used to reach this result. That would be covered by an NDA until the model is properly released.
Exciting times!
Quite a jump in conclusion you are making here.
Sol is able to find the same counter example independently [1], so no reason to conclude in the existence of a benchmark destroying math beast Fable 6.
[1]: https://x.com/aaron_lou/status/2079218392452530249
My Fable 6 theory is admittedly speculative, but your “[public GPT-5.6] Sol is able to find the same counterexample” is also a jump in conclusions. Aaron specifically says he used “an internal version of Codex”. When asked whether that meant a different model or harness, he dodged the question, and only said the harness should be the standard commercial GPT-ultra harness [0]. He (intentionally) avoids identifying which model was used, so your claim is similarly unresolved. Given that Aaron works at OpenAI and has access to internal models, that GPT-5.7 is expected to launch in a few weeks and is rumored to be 10T+ parameters, it's very plausible that Aaron used that model in his analysis.
Furthermore, public GPT-5.6 pro failed six times to find a disproof to the Jacobian Conjecture [1], even with hints, which is evidence against the claim that the public GPT-5.6 Sol can solve this.
Regarding the existence of Fable 6: an internal upgraded version of Fable or Mythos almost certainly exists, given that Anthropic has been testing Mythos internally since April and previously released new models roughly every ~6 weeks.
[0] https://x.com/eliebakouch/status/2079237073001730510
[1] https://x.com/Tomodovodoo/status/2079172223055863895
3 replies →
I'm not very excited. Access to the best AI is not a party I was invited to. And the people who are at that party, well, they don't exactly reflect on my best interests.
Is mathematics an science or an art? To the extent it's art, it's expressive and rewards the human experience that inspires and that creates it. If it's art, it's drastically less valuable to advance it through automation. But if mathematics is a science, then our entire goal is to increase humanity's understanding of the field. Whether by automation or genius inspiration or as a reward for decades of grinding it out incrementally, it's all the same: knowing more is the point.
Given how small the counterexample is... this feels like a great example of where a lot of interesting results are going to be found: not because they were super difficult, but because intelligence didn't scale, and until computers could do this for us, the number of people who seriously poked at many such things was low.
I'm very excited for the impact of this effect in science and medicine and other disciplines too.
I have a question I'm surprised people are not asking: How did Fable find this? Was it like guessing a bunch of families and then solving for possible solutions in those families? Was it clever search? something else?
Some speculation in this Claude chat: https://claude.ai/share/22abed98-d9af-43c5-9881-b19e009a07b0
linked from here: https://x.com/b_shrir/status/2079094004885668003?s=20
Very short version: there’s an existing false counterexample in the literature which holds almost everywhere except at a pole. It looks like Fable used this polynomial as a base & extended it in a way that eliminated the pole whilst preserving the structure.
I'm going to paraphrase what GPT told me: Consider the canonical degree 3 (subvariety of the trivial P1 bundle consisting of zeros) cover of the projectivization of homogenous polynomials of degree 3 in 2 variables (so it's a 3fold cover of P^3). The top space is P1 x P2 and if you take a standard affine open of the base and look at the cover over that restricted to a subset where the zero of the cubic is simple you get the map for some choice of coordinates...
I honestly have no idea if it's correct lol I didn't check it (I should given I actually work in AG) but it doesn't look impossible at first sight
7 replies →
I’m doing this by working all logical steps into lean (formal verification) the quick feedback loop between the AI prose and the Lean verification errors and warnings ensures that its logically consistent.
The issue that remains are two things, ensuring the idea of the proof is actually the thing you want to prove and the interpretation of the results you get. But besides that, everything inside of the kernel checked code is logically consistent
In my experience, lean will show that it's correct, but does it not lose the mathematical intuition that led to the result? As far as my experience goes, that's really hard to encode in lean itself.
Could we maybe get more information about the problem from the LLM trace itself here?
6 replies →
Since I actually don't know math, maybe my ELI5 understanding can be helpful (or corrected).
The conjecture says that you can always reverse (a process) to determine the original inputs.
But this proof shows multiple inputs creating the same output - which obviously cannot be reversed to determine the input - thus falsifying the conjecture.
[...] you can always reverse (a process) to determine [...]
There is a precondition - constant non-zero Jacobian - to the inverse existing and the inverse is claimed to be of a specific kind - polynomial. The counter example satisfies the precondition and by mapping two different inputs to the same output makes any inverse impossible, including polynomial ones. But maybe that is already ELI7.
This helps because the ELI5 was too general. To me at least.
It was known to be false in general. But there are many questions that are false in general, but true when restricted to polynomials.
Exactly, yes!
And the conjecture was for a specific class of processes.
And Fable found an example of one concrete* process in that class and three concrete inputs (two were enough of course) giving the same output.
*Concrete here means given by a finite string of characters
Maybe not?
https://en.wikipedia.org/w/index.php?title=Jacobian_conjectu...
I think the editor themselves misunderstood the conjecture.
UPD: The edit got reverted and there's this on the talk page now: https://en.wikipedia.org/wiki/Talk:Jacobian_conjecture#c-DaR...
UPD2: There are edit wars happening now: https://en.wikipedia.org/w/index.php?title=Jacobian_conjectu... https://en.wikipedia.org/wiki/Talk:Jacobian_conjecture#c-Sea...
Wikipedia's gonna Wikipedia. Unless there's a material debate over the Jacobian Conjecture itself, there's really no open question here. This isn't a complicated proof; it's a straightforwardly checkable certificate of a solution.
6 replies →
I agree, the conjecture never says anything about a proper map.
[flagged]
2 replies →
This is nonsensical: Properness of the map is equivalent to its being an isomorphism (quick proof: Jacobian invertible implies that the map is etale, and properness would imply that it is finite etale, but affine space doesn't admit non-trivial finite etale covers), so the lack of properness is just another way of verifying that this is indeed a counterexample.
sure it does? two copies of the affine line? (I guess there's no galois group & no connected finite etale things tho)
1 reply →
[flagged]
Anti-AI mentality claims another victim
The author has a PhD in math from Cambridge. If it turns out to be a false claim it is an interesting case study on AI's sycophancy causing even experts to drop their guard and make mistakes.
Who do you mean? The author of the tweet is Levent Alpöge, who does not have a PhD in math from Cambridge... but does have a PhD in math from Princeton. And his advisor was Fields Medalist Manjul Bhargava.
Also, there is no way this counterexample is wrong. You can very easily check it for yourself. (I did, I don't know why, obviously Levent wouldn't be wrong about this, but I guess I was in shock.)
1 reply →
[flagged]
21 replies →
I'd rather wait for independent seasoned mathematicians to verify such claims first before someone at said AI lab posting a claim about solving a proof online.
Let this be a lesson to those who fell for such AI psychosis and to not believe everything you see on the internet as real.
It's a Jacobian determinant and three points. You can check this yourself in Sage.
The author is a Princeton math doctorate.
10 replies →
It's a Princeton math PhD who posted. The verification is quite straightforward and was posted by the tweet author. Wolfram would have to also be producing incorrect outputs for the counterexample to be false. The counterexample works as claimed and conjecture has been proven wrong.
Maybe have a seasoned mathematician check if the counterexample is really bogus before saying people fell for ai psychosis.
4 replies →
Looks like this has already been formalized: https://github.com/deancureton/jacobian
2 replies →
This topic really is a testament to people's willingness to opine on things they have absolutely no clue about.
A first year undergraduate can completely check this counterexample in ten minutes. The original post even linked Wolfram alpha for the calculations.
And if you genuinely try you can very quickly understand using only high school math and a bit of Wikipedia that this counterexample is vanishingly unlikely to be wrong, even if you don't do the calculations yourself.
It's a counterexample, not a proof. A schoolchild can confirm it.
1 reply →
This is so unreasonable! As @__alpoge__ himself notes this is classic crank graveyard territory and yet the counter example is something a grad student in 1997 could have found w a ~3 day computer search. Wild!
The search space for a naive brute force of three polynomials of degree <= 7 with integer coefficients <= 12 is roughly 10^500. I think it would take a little longer than that.
> and yet the counter example is something a grad student in 1997 could have found w a ~3 day computer search
Is that true? Even restricting this to f(x,y,z) and coefficients and powers to 1 ≤ x ≤ 10, there are a lot of polynomials to check, and checking requires checking the Jacobian determinant and, if it’s a non zero constant, finding two points for which the polynomial produces the same value.
Or is there a way to generate all polynomials with a non-zero Jacobian determinant, and does that speed up things? (My intuition say it wouldn’t, because I guess those with zero determinants are rare)
I mean you’d probably just generate random low descriptive length f(x,y,z), check for a const determinant then poke for invertibility.
Be fun to ask Fable to write a search program to find more counter examples using only early grad theory to guide the search.
1 reply →
So IIUC no one really worked on that in „professional” or rather university world since then?
I suspect the LLM was able to synthesize a counterexample because of the availability of a lot of prior work:
> The Jacobian conjecture is notorious for the large number of published and unpublished proofs that turned out to contain subtle errors.
https://en.wikipedia.org/wiki/Jacobian_conjecture#cite_note-...
Maybe, maybe not. The "proofs" may not have helped at all with finding a counterexample. Either way, it doesn't matter. A counterexample was found, no one found one before even though clearly a lot of people have tried who also had access to the prior "proofs".
I think it's becoming harder and harder to argue that LLMs don't really reason and just mimicry human speech. This counterexample is clearly the result of a sequence of steps that build on previous knowledge in context and logically combine it to reach other true statements - to a degree and complexity that rivals the best human minds.
For someone that use Claude Code every day, this is obvious, but for some reason many scientists refuse to accept that it's truly reasoning; perhaps not in the human sense, but in a very profound and real sense. These powerful results are devastating to their point of view.
I can sympathize, because I too called LLMs "fancy Markov chains" in the GPT 3 era. But there comes a time where you have to update your world view to match reality, or be stranded in fantasy land.
10 replies →
This makes me wonder, what if anyone uses Fable-class LLM and passes of its novel results as their own work?
There's no shortage of folks doing that in software, right now.
3 replies →
How does that work in this case? What do those proofs do to help find this counterexample?
In my view, it's theoretically possible for a combination of the author's iterative prompts + evaluation with Wolfram Alpha to activate the weights that encode the language that describes the constraints on these polynomials (from the faulty proofs) in such a way that the author eventually arrives at this:
> ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3
8 replies →
I mean yeah, maybe, but this is how everything works
Would be really interesting to see the full conversation and understand how Fable came up with the counterexample in the first place.
I've been asking Fable very complicated questions in heterodox economic theory, which is something I know a lot about. The stuff it comes back with is incredibly deep.
To give a metaphor that everyone here on HN would understand, reading it's responses gives me the same level of wonder as one gets learning how quicksort works for the first time. It even stretches my brain to grasp what it's even come up with. I find myself getting mentally exhausted just digesting it's brilliance.
I think the singularity will have this point where AI comes up with ideas so profound, like a Ramanujen equation, that the most brilliant among us can't even decipher the answer to our questions. The internal reasoning of the machine is at a level of complexity that's beyond human comprehension to even keep track of everything enough to integrate the understanding of what it's come up with. This will happen with any even mildly complex question about any topic.
I don't consider myself an open weights supporter, but it's a bit of a bummer that if there's any novel search technique discovered by the model throughout this finding, it's possibly locked behind ant's reasoning summarization.
One wonders if they could turn their mechinterp work into analyzing the "thought processes" of these very special cases that turn into novel research and finally crack the creative thinking barrier.
I've come to understand that while an LLM is a parrot, it's a parrot that's smarter than I am.
At least you recognize this. The average Redditor is not capable of such self-examination.
Highly relevant comment <https://np.reddit.com/r/singularity/comments/1jh9c90/why_do_...>:
>Imagine you had a frozen [large language] model that is a 1:1 copy of the average person, let’s say, an average Redditor. Literally nobody would use that model because it can’t do anything. It can’t code, can’t do math, isn’t particularly creative at writing stories. It generalizes when it’s wrong and has biases that not even fine-tuning with facts can eliminate. And it hallucinates like crazy often stating opinions as facts, or thinking it is correct when it isn't.
>The only things it can do are basic tasks nobody needs a model for, because everyone can already do them. If you are lucky you get one that is pretty good in a singular narrow task. But that's the best it can get.
>and somehow this model won't shut up and tell everyone how smart and special it is also it claims consciousness. ridiculous.
If you have to be a parrot of another man's thoughts, at least let that man be a synthesis of all human endeavor to date.
"If I have seen further it is by standing on the shoulders of Giants." -Isaac Newton
Smarter in some ways at least. Probably still not quite as smart at understanding human emotions (and things that aren't well enough catalogued on the internet or amenable to Reinforcement Learning on virtual environments).
Probably not as smart as highest EQ humans but smarter than me, I use LLMs to help pin down what emotion I'm feeling.
1 reply →
Perhaps we are all parrots.
Hot take: even GPT-3 was not a parrot. Skeptics have never properly internalized that the fundamental operation is basically irrelevant to the gestalt. Humans are not parrots yet neurons likely also largely operate via predictive processing.
I agree, and was influenced by this, which discusses it more: https://www.astralcodexten.com/p/next-token-predictor-is-an-...
Context: https://en.wikipedia.org/wiki/Jacobian_conjecture
> Jacobian conjecture [...] states that if a polynomial function from an n-dimensional space to itself has a Jacobian determinant which is a non-zero constant, then the function has a polynomial inverse.
> ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
But 1 != -1 and -3/2 != 3/2 . So it's not its own inverse. Is the conjecture that it is its own inverse or that is has an inverse?
Edit: it was worded a bit strangely, but it is saying that [ (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) ].map(F) all produce (-1/4, 0, 0). Thus it has no inverse and indeed disproves the Jacobian conjecture.
I think the map sends (1, -3/2, 13/2), -> (-1/4, 0, 0) and also (-1, 3/2, 13/2) -> (-1/4, 0, 0) so it's not invertable which disprove the jacobian conjecture that polynomial maps over complex numbers with a jacobian that's non-zero are globally invertible.
(Just as a note for myself, I had to think of why the fact that such jacobians are constant is a byproduct, I guess it's because of lioville's theorem implying that any polynomial over C that never hits 0 must be a constant [because the reciprocal is bounded and thus must also be a constant])
That it has an inverse.
If f(a) = f(b) for a≠b then f can't have an inverse.
Suppose f has inverse g; then g(f(x)) must = x for all x.
But then g(f(a)) would have to equal a, and g(f(b)) would have to equal b. But they can't, because f(a)=f(b)
For all thehubbub, as far as I know, all the math breakthroughs via AI that I've heard about have come from Anthropic and OpenAI, not Chinese models. I could have missed those announcements, but one might think that between close to frontier performance plus cheap tokens, that they'd be leading the way on these things.
All (unless I've missed one?) the math breakthroughs are also coming from the relatively small number of mathematicians working at these companies, as opposed to the many orders of magnitude more mathematicians using LLMs for mathematics outside of the companies. I assume the missing link everywhere is being able to casually burn a few rainforests worth of tokens in pursuit of something publishable.
Depends on exactly where you draw the line at "breakthrough", but there's been at least a few novel and interesting results coming from outside mathematicians. Recently for example there were https://old.reddit.com/r/math/comments/1uxj3cy/after_openais... from Phillip Kerger at Berkeley, and https://www.erdosproblems.com/forum/thread/119/proof-claims from Samuel Korsky at Two Sigma (the latter of which was more of a collaboration between the human and machine, not a one-shot like several of the other results we've discussed).
No. Most of the results are from 3rd parties using publicly available models. The most impressive results have come from direct announcements, but most in total have not. Open AI have only announced 2 results and the total is up to a dozen or so now.
"Anthropic and OpenAI" it's basically been all ChatGPT outside of this one.
As with the case with AI and in fact a lot of things in life, sometimes you just need to wait. It will happen, rather soon.
I mean of course an employee has cheap tokens
> as far as I know, all the math breakthroughs via AI that I've heard about have come from Anthropic and OpenAI, not Chinese models
Maybe because dishonesty is more normalized in SV business culture
Post hoc ergo propter hoc.
Are you sure that this applies?
3 replies →
Because they've been proven equivalent, so too fall the Poisson Conjecture and the Dixmier Conjecture
That gave me something to play with with these models and this came out as a claimed disproof of the Dixmier Conjecture:
https://gist.githubusercontent.com/pedrocr/51157b9b2eed8152e...
It seems mathematics has at least gotten a powerful new tool to automate the work to cascade results after breakthroughs are made.
Somebody said this on the xcancel thread but this also means that the Dixmier conjecture for the third Weyl algebra is disproven
Is there a site or github repository that aggregates all these non-trivial AI-solved results?
I am a number theorist and a graduate math student at Bonn. It is eery to see that all of a sudden everyone cares about pure math. Anyhow, I think such incidents like this one are only to happen more often in the coming days, and while many mathematicians are concerned about their role in the community; I believe since we are still very early stages of AI-driven discovery, we still need the appropriate infrastructure for human-AI research, such as with versioning for proofs, a bigger library than mathlib and a central reasoning space so traces aren't lost- which happens to be the most valuable training data for models, and without which, these discoveries do not help much at advancing the field, and remain as blackbox.
> without which, these discoveries do not help much at advancing the field,
At least the fact that frontier models are not optimized primarily on formal math reasoning is in itself is good for not putting mathematicians out of their jobs, isn't it?
How do these results (and the future painted by them) affect your profession; do you expect a similar route as in software, where junior developers are unemployable, LLM-assisted development is the norm, and that great developers stand out partly through better communication with their managers?
I think in short term, it will get better, as more funding flows from industry to academia, you will see a spur in well-compensated PhD positions. However, in the long run, while I doubt that mathematicians will go extinct, they might have to move on to industries like in trading, chip-making, where they oversee AI models write code and also proof that the code works and is consistent with other parts of the software. However, until we have infrastructure for human-AI math research, as Tao himself said that the "roads" for human-AI research is yet to be built, humans would simply be working FOR AI models, and not working WITH AI, which will truly scale discoveries at a massive scale.
Looking at this poor tweet and only thinking that Telegram has proper LaTeX rendering.
You can have that also on Mastodon:
https://mathstodon.xyz/about
Thanks for the mirror link OP!
https://xcancel.com/i/article/2079135211196121363
can any serious mathematicians verify https://xcancel.com/i/article/2079135211196121363
he just...he tweeted it out
I almost felt bad adding {{Cite tweet ...}} to the wikipedia article
At least it's not {{Cite bridge graffiti}}
https://en.wikipedia.org/wiki/History_of_quaternions
The "thanx" is one for the ages.
Related & of interest: https://xenaproject.wordpress.com/2026/07/20/human-mathemati...
Overhead, without any fuss, the stars were going out.
The funniest thing about LLMs is the cognitive dissonance they cause people. People clearly recognize (and bemoan) the fact that LLMs produce derivative breathless prose ie they fundamentally fail at "unstructured creativity" (something the might accurately labeled intelligence) but are then shocked that the same LLMs can do math.
It's reasoning from a flawed premise that math universally requires intelligence and creativity. It does not. Anyone that's proved things via "diagram chasing" can affirm that. The conclusion you should draw is that math (at least the kind they excel at) isn't actually a creative endeavor.
I would think it's more that some aspects of math (like anything really, including programming) don't require creativity and can be solved through "brute force" or whatever you want to call what LLMs do. But it's pretty obvious that LLMs are not capable of solving the vast majority of problems in math (or programming) at this time. Eg. Google went 9/353 on Erdos problems. If a more powerful LLM is capable of solving those or if they require a certain je ne sais quoi of the human variety is still up in the air at this point. In either case it seems like they require a long and detailed prompt from a domain expert (ie. human) regardless.
Didn't humanity also score very low on 353, namely 0 since they were open? Probably collectively we could have gotten a slightly better score if all of Math started trying to solve those, but not by that much, I think, since they are precisely still open.
I dunno, this screams of goal post moving. Even if LLMs lack whatever nebulous definition of "creativity" that someone favours, there's no inherent reason for "creativity" to be required to solve any problems at all, "creativity" could just be a human method for solving problems that evolved because of it's broad applicability but is suboptimal at any given task.
> could just be a human method for solving problems that evolved because of it's broad applicability but is suboptimal at any given task.
Ya sure let's just posit another random hypothesis about evolutionary biology in order to substantiate the claim that LLMs are intelligent.
Or (bear with me) you can recall your (likely) experience proving stuff like SAS triangle identities and reflect on whether that required intelligence or just computation.
What's more reasonable?
5 replies →
Hilariously absurd cope.
They're still not very good at writing, but you've flipped the takeaway. The correct conclusion is writing style isn't very relevant to intelligence.
I think the real tragedy is that they are good at writing, just by default are tuned to have a kind of bland corporate tone. If you give the LLM a few pages of writing you like and tell it "Continue, but using this style" it will do a pretty good job of it. Most people just .... don't bother to do that.
1 reply →
It's just Moravec's paradox all over again
Lots of other point collisions too: {{{-1, 3/2, 13/2}, {1, -3/2, 13/2}, {0, 0, -1/4}}, {{0, 7/8, -(3479/1152)}, {48/7, 0, 49/1152}, {-(48/7), 7/16, 539/2304}}, {{-1, 7/8, -(40523/648)}, {160/7, 0, 49/5184}}, {{-2, 7/8, 2989/1296}, {32/21, 0, -(49/192)}}, {{1/63 (-120 - 16 Sqrt[3]), 7/8, 0}, {4/7 (-1 + 3 Sqrt[3]), 0, 49/288 (5 + 3 Sqrt[3])}} }; Next steps would be a simpler counterexample or family of polynomials where the conjecture breaks.
If there really was a "simple" solution to Fermat's Last Theorem, Andrew Wiles wouldn't have achieved the important result he did, ending up making connections across disparate fields of math.
The LLMs "sweeping up" easy, or previously missed, results seems like a net negative. It's probably better for humans to struggle and come up with new tools than to just "clean up" low hanging fruit that doesn't add much value to the field.
Can someone validate this counterexample from gpt-5.6 sol
(-(1+xy)^2 z - y^3(1+xy), 2x(1+xy)z + (1+xy)^3 w + y^2(7+12xy+4x^2y^2), 2x^2z + 3x(1+xy)^2w + 2y(1+10xy+6x^2y^2), 2x - 4x^2y - x^3w): C^4 → C^4 has Jacobian determinant 4, and sends (-2,0,1,0) and (-1,0,1,-2), (1,−2,−7,14), (2,−1,0,3) to (-1,-4,8,-4)
https://xcancel.com/__alpoge__/status/2079091571316912542
https://www.wolframalpha.com/input?i=Det%5BD%5B%7B-%281%2Bxy...
Anybody can ELI5?
Are you at least a little familiar with linear algebra?
If so, you've probably heard of the determinant. It's a certain way of "summarizing" a matrix with one value.
The determinant in this case is of the Jacobian, which is a matrix you can construct from a multi-variable function. Each term is the partial derivative with respect to each variable (x, y, z, etc.), with one line per output variable (vector element).
The Jacobian of a polynomial function is, in general, going to be a matrix where every term is some polynomial expression. And the determinant of that will also be a complicated expression. But in some cases all the variable terms cancel out and you're left with a single constant (0 or some other value).
The conjecture says that if the Jacobian determinant is constant (i.e., all the terms cancel out), then there must be a polynomial inverse. And the key condition for an inverse is that there must not be two input points that evaluate to the same output. It's just like y=x^2. It's not invertible, because both +2 and -2 square to +4.
So if you can find a function where the Jacobian determinant is constant and also find two or more points that evaluate to the same output, then you've found a counterexample to the conjecture. And that's what's been done. And remarkably, the counterexample is pretty simple. It would be tedious but a bright high school student could verify it.
I think another way to understand it is the generalization of the inverse function theorem. The inverse function theorem gives you local invertibility, but even being "locally invertible" everywhere does not imply global invertibility (you don't need too pathological an example to see this, a periodic function serves iirc).
The Jacobian conjecture roughly asks what whether local invertibility gives you global invertibility when you restrict only to polynomials (which we might hope "behave nicely"). Apparently for polynomials over reals this was disproved a while back, but up until now the general case of polynomials over complex numbers was open.
Nice!
(It's gotta be a nonzero constant, right, a nonsingular matrix).
1 reply →
Thanks that's a very nice summary.
1 reply →
Related:
Open Problems Solved by LLMs? A Survey of Verifiable Mathematical Discovery [pdf] -https://news.ycombinator.com/item?id=48914646 - July 2026 (110 comments)
I find it interesting that the counterexample uses C as a field. C is twisted and weird. Maybe the Jacobian Conjecture still holds for reals?
All the coefficients and evaluation points are rational, so it's a counterexample in all fields where 2 ≠ 0 and 3 ≠ 0, doesn't matter whether that field is the complex numbers, real numbers, rational numbers or even a finite field.
Ah you're right, thanks.
>C is twisted and weird.
Why do you say this? I've admittedly never done a proper complex analysis course but I got the impression that that complex differentiability was a very strong condition that results in holomprhic functions behaving "nicely" in ways that real functions do not
Should have used quotes.. I didn't mean it in any formal sense. What I am saying the nature of unit in complex plane makes it difficult to intuitively imagine invertibility and determinants.
Discussion on the talk page for the Wikipedia entry: https://en.wikipedia.org/wiki/Talk:Jacobian_conjecture#Appar...
Speaking as a mathematician, it does seem like we're a bit fucked as a community. Anything that is at all accessible to currently existing methods and mathematical infrastructure is probably going to fall to the frontier models of today, and at this rate of progress it's likely that, already by next year, we'll see new infrastructure being put into place by AI, giving us a world in which a few designated interpreters of the oracle get to 'do' mathematics, while it withers on the vine as an avenue for the exploration of human meaning.
Proving theorems will have lower payoff, but posing new questions (for AI to chew on) will have higher payoff. Math will go from theorem proving to conjecture farming/exploration. In a way this could be even more fun.
Of course AI can also farm conjectures, but they have to develop taste, which might be harder than just proving theorems.
Yes, as someone who prefers developing the 'correct' structure over 'merely' proving theorems, this is good for me in the short term. However, the writing I fear is on the wall for my medium and long term utility.
> Of course AI can also farm conjectures, but they have to develop taste, which might be harder than just proving theorems.
Do you have any argument why you might think this would be true?
2 replies →
I think in the long run mathematicians are probably fucked, but in the short run it's not that bad. All three of the big conjectures solved the answers were at the level where if you had given a grad student the questions and the right background reading there's a good chance they would have solved it. (This example, you could have given an undergraduate good at programming and computer algebra and told them to come up with a counterexample.)
At this point the advantage of AI is that it's read the entire mathematical literature, and it doesn't have to worry about wasting its time. The solved problems have all turned out to be surprisingly easy, so the real lesson is that we're bad at judging how hard problems are.
Assuming this state of affairs lasts, the medium-term problem is that you learn something when struggling with a problem, even if you don't solve it, and if mathematicians become too reliant on AI the skills they develop through struggle will erode.
The long-term problem, of course, is that it seems much more probable that a future model will make mathematicians all obsolete. But so far Fable hasn't. (Anthropic has probably burned a billion tokens on the Riemann hypothesis already, without telling anyone.)
> This example, you could have given an undergraduate good at programming and computer algebra and told them to come up with a counterexample
please try go try it. There's no way someone didn't do massive computer algebra searches before today.
> All three of the big conjectures solved the answers were at the level where if you had given a grad student the questions and the right background reading there's a good chance they would have solved it.
You cannot be serious... why didn't they solve it before then? Do you think no one tried it? What background do you give the double cycle conjecture student after the flow reduction? a linear algebra textbook???
9 replies →
Goldbach's Conjecture is very accessible.
Anybody that has to work for a living is fucked and not on the "can't do mathematics which they would find fulfilling"-level but on the "can't afford food, because human intelligence is simply not required anymore"-level.
next year is a long time away friend
It's so soon that it may be rational to delay certain math and software projects until smarter models arrive
no, he's right, we're fucked
1 reply →
The most interesting point is that he chose to simply post it on X, instead of pretending it was his own result and posting it on arXiv.
Respect for this spirit. Is this not also a sign that the old academic journal system is already outdated?
Interesting announcement on social media rather than publishing somewhere like arXiv.
You can't please everyone on the internet, if he had posted this on Arxiv then someone else will be complaining that Arxiv is for humans to publish and that llm output should be social media post instead. Also, the proof fits in a tweet. So why blow it up.
a short counterexample makes it likely the best medium for such an announcement. would have been inappropriate for a positive proof.
Apparently ChatGPT-verified: https://xcancel.com/__alpoge__/status/2079045382940573896#m
boo on the downvote. The response is pretty funny IMO
"Any idiot could have done this, it's just high school calculus and just a counterexample anyway. Stochastic parrot, spicy autocomplete, AI psychosis. Wake me up when an AI does something real."
The goal posts have moved. People generally stopped saying this stuff now.
Even if you go to the ultimate anti-AI subreddit r/betteroffline, they've changed from "AI is useless" to "AI is good but the AI bubble will collapse soon" over the last 6 months.
I think we’re not adapted well for this rapidly changing world. Here you have some people who were rightly skeptical about a new technology being shoved down their throats by giant tech corporations, and a technology that really was, and probably still is overhyped.
Yet it’s a technology which has rapidly grown in its capabilities.
So yeah now many of the people who thought it was useless before probably don’t think it’s useless anymore, but you’re holding them to their original words even though those words were about something completely different at this point.
If people aren’t saying it anymore it might be because they don’t think that anymore, and the people who have new goal posts might be entirely different people.
It’s like you’re looking at a different set of goal posts on a different field and saying, no! The goalposts have moved!
People have definitely not stopped saying this stuff.
3 replies →
Some quotes from a day ago, https://news.ycombinator.com/item?id=48957779:
> I hold my stance that LLMs are stochastic parrots... Making the parrots ever more complex and training
> Except solving problem is probably the least (even though it's important) interesting thing in research.
> Can we use AI to get a cure for cancer yet? Or is math-turbation the only thing these things are good for?
> Train on enough examples and statistical autocomplete gets you places. I'm surprised how anyone would even consider this intelligence?
And, as much as HN has declined in the grips of an anti-AI psychosis, Reddit is worse. I would love if social fora would switch to the reasonable claim that we're in a bubble; that's something that can be debated. That's not the dominant critique of AI, though.
3 replies →
I don't think the (fairly factual) description of these systems as stochastic parrots means that they will never do useful work, just that they are not intelligent in the way we believe animals to be (to "push back" on your anecdata, I've also heard fewer people claiming that LLMs are actually conscious in the past year -- maybe we're reaching the happy medium?). That was the point the stochastic parrots paper and Chinese room thought experiments were making -- nobody claimed that the man in the Chinese room would be unable to accurately translate Chinese text.
Fuzzers are another kind of stochastic generator but nobody would claim they don't do useful work in a way that is hard to replicate through deterministic methods. (I still find the code these models produce kind of awful, but advancements in harnesses do mean that they can finally produce code that works most of the time.)
6 replies →
the ROI just isn't there. they aren't making these capital investments back in the next decade even.
not to mention the trillions yet to be spent.
it kinda all works but it's an earth-scale blood from a stone. the resources needed to even do these parlor tricks is nutso.
2 replies →
I asked Fable to verify and it absolutely freaked out!
I have no idea what any of this stuff even means, but my AI thinks I’m a legend level mathematician!
https://en.wikipedia.org/wiki/Sycophancy_(artificial_intelli...
To be fair to the sycophancy tendencies, this was an open conjecture that held for 85 years, and not for lack of trying. So, maybe a bit warranted here? :)
5 replies →
The interesting thing about using Claude Fable 5 is it's nearly as irritatingly sycophantic as past Claudes while genuinely being smarter than the previous models. So you get a kind of yo-yoing of it glazing you as a creative genius and disappointedly revealing to you that your ideas are bad and dumb.
3 replies →
I mean if someone came to you with a value of n that disproved collatz, wouldn't you go crazy as well?
The results are great. They will still need to post the detailed chat session to show how AI is helping though.
Maybe instead of proofs, we should encourage the publishing of prompts and reasoning traces?
Maybe instead of proofs, we should encourage the publishing of prompts and reasoning traces?
The tweeter, Levent Alpöge, is a professional mathematician who works at Anthropic.
But I'm curious—can Fable handle cases where n=2 as well?
This seems like the next step, but the smallest counterexample in n=2 is degree greater than 100. (This is a paper of Moh. Wikipedia has details.)
Wasn't it proven true in general for n=2
(NGL I wanted to suggest someone to go for the JC using 5.6 after the CDC proof came out, but then on reflection felt I should neither waste people's time NOR contribute to the myth of AI :)
My prediction is that the bubble will burst in 2031 Q4, one year after the Riemann Hypothesis is expected to fall (according to Demis)
After 2031, I will suggest going for the JC for N=4 because they would (dis)prove the Dixmier conjecture for N=2 :)
https://xcancel.com/BrunsJulian1541/status/20790734625601334...
https://arxiv.org/abs/2410.06959
>Wasn't it proven true in general for n=2
Assuming you mean C^2 -> C^2, Do you have a link? If so it would be good to add to the wikipedia page. Also I'm not sure, but does the fact that there's a disproof for n=3 imply that it's false in all n>=3, or could there be higher dimensions where it still holds (I'd guess not since you could probably trivially "embed" this in higher dimensions in some way)
1 reply →
And yet it can't even replace a call center employee.
I guess you can't make a career out of being a frog anymore in math.
I want to see the Collatz Conjecture next!
See also https://en.wikipedia.org/wiki/Jacobian_conjecture
Now that is some spicy autocomplete.
The day is coming when some *BIG* problem is solved by AI just because someone jokingly asks about it.
"the Riemann Hypothesis is " <tab>
...not going anywhere.
I just fed this to GPT 5.6 Sol:
GPT wrote some SymPy code to check it. The response?
"As written, this is an explicit counterexample to the Jacobian conjecture. I checked it using exact symbolic algebra.
I do not see an algebraic catch in what you typed. Unless a term or exponent differs from the intended expression, it appears to disprove the conjecture. This deserves serious independent checking rather than casual dismissal."
Waiting for someone to write the Lean proof.
Here's your Lean proof https://github.com/google-deepmind/formal-conjectures/pull/4...
...no need for any lean here
[dead]
[flagged]
The surprising thing is that the counterexample seems relatively "simple" in that it's low degree, with coefficients that aren't too large.
Does anyone more familiar with this know why this _wasn't_ found earlier, when it seems like you could brute-force through some low-order polynomials?
Reading the chain of thought for Fable, it is incredulous as well. It keeps thinking that it must be missing something or that this is a trick, because it can't have just found a counterexample to a famous conjecture. It's just like us!
https://cdn.xcancel.com/pic/orig/EA1E99C3DAE99/media%2FHNpQX...
https://cdn.xcancel.com/pic/orig/3B5C0B867A7C3/media%2FHNpQU...
5 replies →
There are still many degree 5 polynomials, up to 35 terms for a single function, then you need to find three of them as well
[flagged]
Same awnser as much of the LLM Proofs - people cared about other things. There isn't a lot of money in academic math, and the ones that love it don't look for low value findings. Proofs like these are, funnily enough, usually the domain of hobbyists - but over the last few years, the "Monetize everything" mentality and struggling first world economy has pushed people away from interesting academic pursuits on their free time.
10 replies →
[flagged]
[dead]
[dead]
[dead]
[dead]
For the people saying goalposts have moved etc, whats your endgame?
What does this mean? Fallacious argumentation and deceptive rhetoric is acceptable, if the topic is sensitive enough / there is enough riding on a wrong answer being accepted?
Actual tweet:
>> hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final
So where does it say that Fable "produced" the counterexample? The tweet says it was a collaboration between two people, using Fable.
What’s more Impressive you got Fable to do anything useful.
The model is still useless for real data to day work, no matter how many parlor trucks it performs.
I’ve been staggeringly productive with Fable. Opus 4.8 fails a lot more for me.
Fable often just “knows” what I want with vague instructions. It also is able to autonomously perform work that lasts an hour long from my experience. I haven’t tested further.
Without Fable included in subscriptions, I would have moved my entire team over to Codex 5.6.
A counterexample to the Jacobian Conjecture - and also a counterexample to “AI will never be smarter than humans.” Even the most dyed-in-the-wool AI hater at this point must acknowledge that it is more intelligent than any human. The other day I found out that Fable could read seal script! The small seal, standardized stuff no prob, but I found it even did OK at the hardest you can get, Warring States regional scripts. That’s something maybe 2,000 academics worldwide can do and nobody’s even talking about it because it’s just one more item in a very long list.
It’s a strange feeling, to be overtaken by our own creation. Top dog for millions of years and then in the blink of an eye we go from “how many Rs in strawberry” to this.