← Back to context

Comment by SpicyLemonZest

9 hours ago

Their claim is only indirectly related to the motivations of the people using their models. What they're saying is that doing math in this way does not produce the same value as traditional mathematical research, and the people using these AI models aren't concerned about that because their marketing objectives don't depend on whether their results produce mathematical value. If people doing valuable work are made irrelevant by people doing a larger volume of non-valuable work, that's not a positive outcome.

But it does produce value. We now have an explicit solution and even a proof. Humans can then work on clarifying why it's true. Unsurprisingly, not all that different from software, where models can generate working code just fine. The details will all be there and all correct, but the architecture is currently not ideal, so a human guiding it can greatly improve the proofs.

I actually found this to be the case with some basic linear algebra notes I was recently doing in Lean (without using mathlib). The model could generate working proofs, but they obscure the basic ideas (actually I wonder somewhat if this is because the Lean code that's out there to train on doesn't make a huge effort to read like textbook proofs, which was my motivation in the first place). I give it a skeleton of a couple lines of `calc`, letting it fill in the reasoning for each line, and it does much better. Then ask it about making some macros to simplify "trivial" or "obvious" things, and it does even better. etc.

I suspect there's a good workflow where a big SOTA model makes an impenetrable proof (or code) and then a human works with a FIM model to simplify it (with the larger gnarly proof right there in context for FIM), but unfortunately everyone seems to only care about agents right now.

  • >> Humans can then work on clarifying why it's true.

    Presumably you're a human. Are you going to do that?

    • There are vanishingly few research mathematician positions and it's one of the most competitive fields, so no. But I'm not sure how that's relevant. As the OP says, usually the value of a proof is not the knowledge that something is true per se, but the reasoning techniques to understand why. How can it be anything other than helpful then to have a truth oracle as you try to figure out why things are true?

  • > But it does produce value. We now have an explicit solution and even a proof. Humans can then work on clarifying why it's true. Unsurprisingly, not all that different from software, where models can generate working code just fine. The details will all be there and all correct, but the architecture is currently not ideal, so a human guiding it can greatly improve the proofs.

    To me this analogy points in the complete opposite direction. Imagine somebody takes a half-completed project design you're trying to figure out, vibecodes a rough prototype of it, emails your manager to announce that the project just launched in alpha, and then dumps it back on your lap for approvals and testing and productionization. Would you say that they've added value to this process? Or did they just strip away all the hard parts of the problem so they could claim credit for the easy part?

    • One of the awesome things about LLMs is they make it quick and easy to make PoCs, so yes. Proving that an approach will work before spending a bunch of deep design effort is absolutely valuable. Your exact scenario is something I've literally done: give a half-completed design to a team member and asked them to vibecode a PoC to prove the approach will work and figure out some of the details, explore scaling and failure characteristics, etc. Or I do the PoC vibecoding myself too. LLMs have been a gamechanger here.

      4 replies →