← Back to context

Comment by famouswaffles

18 hours ago

1. I would agree if the rumours were that some mathematician(s) had solved them, but the rumors alleged it was Anthropic. I don't really see what the big deal was. They had a new model that was going along great and wanted to test its mettle.

2. Yes Brubeck's comments were weird at face value. That said, Open AI's proof isn't a duplication of anything. Not only is Tristan's work a sub problem but the methods are different. And what OpenAI didn't want was Levant on the paper OpenAI authored not whatever they were working on (Euler). It's petty sure but it's fair enough. Tristan and Levant didn't have anything to do with the Navier Stokes solution, so it's really their call if they didn't want to collaborate on their own paper with the Anthropic employee.

>OpenAI would have you believe this result shows how powerful their mystery better-than-Astra model is, but the reality here is that this model needed 10,000 agents, $20M of compute,

$20M in approximated API prices doesn't mean they spent $20M worth of compute. The real number would obviously be substantially less.

>and the assistance of a whole team of people at OpenAI

You can't eat your cake and have it. What sort of guidance do you think is happening in a 10k agent, 320b token, 88 hour run ? AI did this one.

>I'd say advantage humans this time....to work on problems that have not been solved yet, and that humans are NOT making nice progress on.

Interesting way to frame progress that didn't move along till an LLM generated proof.

> What sort of guidance do you think is happening in a 10k agent, 320b token, 88 hour run ? AI did this one

If you read the PDF release by Buckmaster, apparently the initial claim from Brubeck was that there as very little human input involved, then as the call progressed more and more people popped up that has been involved with it.

Does this aspect really matter? Not really, other than OpenAI wanting to present this as all the work of their model.

**

https://cims.nyu.edu/~tristanb/statement.pdf

I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used.

I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.

  • >If you read the PDF release by Buckmaster, apparently the initial claim from Brubeck was that there as very little human input involved, then as the call progressed more and more people popped up that has been involved with it.

    As it seems and as they tell it, they started the run modestly and diverted more resources towards it as it looked more and more promising. The run didn't start with 10k agents for instance. The point is there isn't anything humans are doing in this timeframe against all this text that would count more than "little human output". It's still a fair assessment I would say.

> Interesting way to frame progress that didn't move along till an LLM generated proof.

This part of your argument is totally wrong. The OpenAI approach begins with the B/L work. The belief / knowledge that their approach would pan out is worth a lot - it means essentially “depth-first” search in this direction will be more fruitful than a general search.

Unless you are counting the B/L work as LLM generated. Is that your argument? Even if you do consider it that way, to me racing in for a scoop isn’t a good look.

>Brubeck's comments were weird at face value

This is an odd way to gloss over threats.

  • I put it like that because of Brubeck's own words on the matter. You're acting like we've gotten email receipts here. I'm not really interested in going over a he-said she-said about strangers.

    • Brubeck has admitted what he said, but claims he immediately retracted it as a "poor choice of words".

      Given Buckmaster's telling, this seems beyond "poor choice of words"... It was a veiled threat, that he then doubled down on with his "If you don’t want me to be nice, then I don’t have to be nice." follow-up.

      **

      I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”

      **

      FWIW there are also other people on Twitter, such as this DeepMind researcher, saying this is a pattern for Brubeck.

      https://x.com/dheeraj_nagaraj/status/2097266146445774924?s=2...

      4 replies →