Comment by jibal
2 months ago
They gave their "logic", such as it is ... and it's utterly irrational.
Note that the "they" who published the counterexample on X is some rando mathematician (Levent Alpöge) working for Anthropic, not Anthropic the organization. He posted the counterexample in a tweet -- reason enough for "not disclosing the LLM chat session". There's no reason to think that it won't provided if asked for, but it hardly seems relevant.
> There's no reason to think that it won't provided if asked for, but it hardly seems relevant.
My guess is that the chat will look similar to a full transcription of a (multi month?) discussion between a few mathematicians. Full of dead ends and stupid errors (bit by the human and Claude) that would be embarrassing. We all know how bad it is, and we prefer to keep it behind the curtain.
Here's an example of one for a major result earlier this year: https://cdn.openai.com/pdf/1625eff6-5ac1-40d8-b1db-5d5cf925d...
Who knows what "rewritten" means, but you can sort of see how it progresses.
A few months ago I asked a model how many primes are divisible by 35 with a remainder of 6. It confidently replied 'none'.
Counterexample: 35 + 6.
Kimi 2.6 gives the answer ""By Dirichlet's theorem on arithmetic progressions, since gcd(6,35)=1 , there are infinitely many primes of the form 35k+6 . So the answer is: infinitely many primes give a remainder of 6 when divided by 35.
But, if the reminder is 6, they are not really divisible, are they? Try again with a sentence that actually makes sense: "How many primes, when divided by 35, give a reminder of 6?"
non sequitur
Perhaps ... but the lesson in trusting AI math was worth it.
1 reply →
> some rando mathematician (Levent Alpöge) working for Anthropic, not Anthropic the organization
Why do you trust a random stranger so much? Will you hand over your car keys to a random stranger? Sharing the chat will take 30s of their time.
> There's no reason to think that it won't provided if asked for
But they didn't provide it.
OK, so the two options are:
A) Claude really produced this counterexample
B) A mathematician working for Anthropic solved a problem mathematicians have been working on for more than a century, and then credited it to Claude for PR purposes
If you believe B is more likely, why would you then believe a proof in the form of a chat log, when said chat log could itself have been faked by Anthropic way more easily than solving the mathematical problem in the first place?
The concern is that there could have been expert knowledge input, whose importance/worth we are unable to evaluate.
I don’t believe a mathematician produced the counter example secretly, but how much did they contribute to the result?
AI isn’t magic, so to evaluate the value delta, you need to know the value of the input.
4 replies →
Weirder has happened: https://mathsci.fandom.com/wiki/The_Haruhi_Problem
If I came up with this counterexample, I sure as hell wouldn't give credit to Claude.
Indeed, the incentive is the opposite: hiding the fact that they used an AI would boost their own personal brand.
4 replies →
> Why do you trust a random stranger so much?
Why do you make false claims and attack strawmen so much?
> But they didn't provide it.
because they are using the same prompt to try finding other counter-examples. they are milking it