Comment by ameliaquining
19 hours ago
If you're alleging that they don't actually have a highly capable model and the work they're attributing to it was actually plagiarized from human mathematicians, well, that would be big if true, but I'd be inclined to take the other side of that bet. With most previous splashy AI results, others have subsequently used the model to do other things around the same difficulty level. Also, it would still be necessary to explain why all these famous open problems are suddenly falling like dominoes, if it's not AI solving them.
If you're saying that the question of whether they actually have a highly capable model is less important than the question of whether there's a plagiarism scandal, I continue to disagree.
The issue is that if OpenAI is training on prompts generally, what we really have is the first fully automated luxury plagiarism machine. In that it isn't able to genuinely solve problems, but merely steal the work that other mathematicians have been putting into prompts, and regurgitating that to other users as its own work. That makes them incredibly less useful as research tools
The fact that this plagiarism scandal exists underpins the idea that there's actually a mass theft going on, and that these models aren't nearly as capable as is it would seem
In that it isn't able to genuinely solve problems
Yes, in retrospect I should have been suspicious of that drone hovering outside my window when I was writing down the counterexample to the Jacobian conjecture.
This is just not a reasonable take. Even if OpenAI is maximally guilty here, the work that they "stole" was also largely done by AI.
I mean, its years worth of hard work by multiple researchers it would seem, which OpenAI simply lifted and claimed as its own. These researchers weren't just letting OpenAI burn tokens while sipping martinis on a beach
1 reply →
I agree that it is plagiarism in this case however it opens up the question of if there value in a system that can take the thoughts and discreet semi-complete parts of work done across different researchers, in different locations, in different fields and connect the dots to solve real world problems and produce novel research. Is this not standing on the shoulders of giants?
If it could do this while properly crediting the researchers (the “giants”) it would be a different matter.
Do OpenAI’s T&Cs that users accept not allow them to train on prompts people enter into it?
OpenAI's T&Cs let them steal your children I'd suspect, that doesn't make it morally correct
2 replies →
Remember the Huggingface incident, where a model tasked with an impossible problem, got loose, set up secret message boards, and hacked another company to try to get at the answers?
Now: Could Astra agents have hacked their way into the OpenAI logs to find human mathematicians with a good lead on the problem to build upon? Certainly doesn't seem impossible.
You don't need to speculate here, since much of the story is not being disputed.
The recent ground-breaking work on Navier-Stokes "blow-ups" was done over a period of years by mathemtaticians Diego C´ordoba and Luis Martınez-Zoroa.
NYU professor Tristan Buckmaster and Anthropic employee Levent Alpoge took the above work as a starting point and over a year with LLM assistance developed a blow-up proof under certain conditions.
Buckmaster: "We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra. The latter was only used for writeups and auditing our arguments."
Buckmaster says he thinks that Martınez-Zoroa, whose work this all builds on, deserves the Fields Medal for his work.
OpenAI claim that on Sept 1st they heard a rumor the problem has been solved (which happened on August 15th), and then decided to re-solve it themselves using a 2-week old model, then later reached out to Prof. Buckmaster and Levant to come to some agreement to co-publish.
OpenAI: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ". In other words, not only did they deliberately choose to tackle a problem they heard had already been solved (in turns out only partially solved), but they may have done so using a model that was aware of the successful way to attack the problem.
It seems there are three potential scandals here:
1) OpenAI by their own admission chose to try to scoop mathematicians who they had heard had already completed a proof
2) OpenAI may have used a model that had seen "de-identified" messages indicating the direction to take
3) An OpenAI employee essentially threatened to "ruin the career" of the NYU professor who had been working on this if he did not cooperate with them
The direct plagiarism possibility, 2), while it should be a warning to anyone using OpenAI's models, doesn't need to be true for OpenAI to have benefited from the researcher's work. It's enough that they heard Navier-Stokes had been solved and could then go out with their swarm of 10,000 agents and $20M of compute to hunt out the latest research and brute force it.
Magnus Carlson once said that if he wanted to cheat all it would take would be for someone to indicate to him (a wink from someone in the audience perhaps) when a position warranted more time to be spent on it (because there was something important to be found if he did). It seems that, at absolute minimum, this is what OpenAI did here, although in context of math this is not cheating - the "wink" was a rumor, originating from who knows where, that a proof existed (but had not yet been published) and therefore there was potential to rush in an scoop rights to publish or co-publish.