← Back to context

Comment by csomar

6 hours ago

Are the models improving? Because I am not seeing it. I have been trying Astra for a few quantifiable tasks in my codebase and performance wise, it's pretty similar to sol 5.6. Now when it comes to expressing the problem/solution, holy Christ, what a mess the writing has become. It is on the level of Opus 5. Now when it comes to burning money, Astra is just insane. With a $100/month subscription, you can easily burn through your weekly "allowance" in a morning.

Needless to say, for practical purposes am back to 5.6/Opus 4.6-4.8. But hey, maybe I am not smart enough to use LLMs?

Yes?

If we look at the math problems they're solving their just now reaching the human frontier... they weren't doing that before.

And your comparison point is model released 2.5 months ago... saying for some use case you didn't see noticeable improvement in 2.5 months (even while other people and benchmarks disagree) isn't a great argument that they aren't improving.

  • Math problems are highly structured, very precisely defined, and already heavily studied and not very complicated compared to problems in engineering or finance. There's a lot of quality material on which to train and it's easy to tell quality apart from crap. The search spaces are a priori much smaller than in other areas and the people using the tools to study them are themselves good mathematicians.

    Success in such problems does not automatically extrapolate to other contexts.

    • Finding a training algorithm that can do recurrent networks and continual learning is also a "highly structured, very precisely defined, and already heavily studied and not very complicated compared to problems in engineering or finance"

      That's the thing I'm most worried about - LLMs that are super clever at coding and maths, making an actually very very dangerous model that is far more efficient, and clever in a more innate (less brute force) way.

    • >compared to problems in engineering or finance

      Jane Street is apparently one of Anthropic's biggest customers. Probably engineering, finance, and some math.

  • I think it’s more likely that that’s because no one tried to solve such problems with them before (OpenAI apparently started working in Navier-Stokes after a rumour that someone seriously advanced the problem with AI) plus improvements in orchestration. Fair, the latter could be as dangerous as stronger models.

Seems like hundreds or thousands of agents are needed to come up with real breakthroughs. Both with the Navier-Stokes project and in the Hugging Face “project” there were lots of agents co-operating on the tasks.

  • I doubt the Hugging Face one would take that many if hacking Hugging Face was the direct goal being optimized.

    • I agree that it could be done with fewer agents. It would take longer though. Seems to me that these agent farms are good at coordinating and co-working in large projects, with the agents using message boards for communication.

Some people claim Astra is significantly better than anything else and significantly more token-efficient, and others (like you) say it's meh and way more expensive to boot. I really don't know what to think.

Kind of a tangent, but one thing I am curious about is to what degree the Navier-Stokes result announced today was primarily a brute-forced result based on the 'program' previously established by researchers to find counterexamples (blowups), or whether the model actually added significant/novel intellectual value beyond its ability to run at arbitrary parallelism. With 10K agents and a staggering $15M in compute (IIRC), I am feeling like a lot of the former may have been involved, but I don't really understand either the problem or the approach (or, indeed, the solution).

Obviously the potential for parallelism and coordination between so many agents is quite scary by itself, but I think brute force by 10K mediocre AI mathematicians is much less scary than ~one AI mathematician reasoning its way through the problem where all human attempts have failed. It seems fairly obvious that massive parallelism lends itself to brute-force counterexample-finding, and I suspect it isn't a coincidence that most of the touted AI math results have been counterexamples.

It's all still quite scary, but coming full circle: I really don't know what to think.

Try GPT5 and you will feel the difference. Not one from 2 months ago, but one from a year ago. And then you can get the idea of what happened in just 1 year and what you can expect in 1 year.