← Back to context

Comment by eigenspace

10 hours ago

Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.

I certainly wouldnt have predicted that 10 years ago.

Very glad to see Mistral still in the game even after some big stumbles with Large 3. I deeply hope that this model is 'good enough' that it becomes the European go-to, giving them the resources to keep the pace up.

I'm excited to try this out today.

I think a big part of that is the Chinese publishing the solution for everywhere hurdle in the road they've encountered in the form of a paper.

Deepseek essentially releases instruction manuals in paper form.

  • I think it might have accelerated things but on a much more basic level, there seems to be no real moat in synthesizing the world’s knowledge into LLMs.

    • There's no question that training leading LLMs requires some serious expertise and know-how, but surely already having advanced LLMs/agents must be helping tremendously not only for software engineers but also for those working on LLMs themselves.

      1 reply →

    • I think the moat is going to be compute. So far compute needed to push the frontier is still extremely cheap so the capital can afford to spread its bets. But when further improvement is going to cost in trillions, capital will have to pick a winner and bet only on him. It won't be a matter of finding the best bet, it will be a matter of survival.

      This will cause the picked winner to get massively ahead with sheer compute alone used both for training and inference dedicated to recursive self improvement.

      30 replies →

  • My spend on DeepSeek is not much and I regularly top up my balance every month as my support for all the good work DeepSeek is doing for the open science.

    • DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government. The reason they release the AI models is economic warfare against US, not because of charity or kindness. It's great for us consumers, but the goal is not to help humanity or open-source.

      152 replies →

    • They're trying to pull digitally what they already pulled physically. The reason we can't manufacture a grill brush for a reasonable price is the result of years of Americans choosing the cheapest price. We gave up our manufacture base. They want us to give up our labs.

      4 replies →

  • Architectural/algorithmic tweaks do advance the efficiency frontier nicely. But raw intelligence mostly comes from data (not just its sheer quantity, but also how it's curated & cleansed) and the scaling law. The know-how about data curation doesn't seem to get published much, even among the open-weight labs, though.

    • This. Even in the efficiency frontier, it is a lot of data curation that actually makes many of those tweaks actually work at scale in practice.

  • I'm not sure I like this framing - so much of AI research has been academic, in the open, building on others people's work. Much less comp sci generally, math & philosophy, etc. The idea that rich companies can just build stuff in secret because they have resources is a fantasy.

  • So boring to see conversations moved over to Chinese models when that’s not even what we’re talking about here. This is about Mistral.

    • Europeans are irrelevant these days, surpassed by China, S Korea, Japan, Hong Kong, Singapore, etc. Europe is coasting on former glory and now has regulated itself to death and vacationed its advantages away.

      2 replies →

  • Also, the field moves fast, but slower than people do. Researchers and engineers switch companies every year or two, and the know-how walks out the door with them.

  • How it should be. Knowledge should not be copyrighted. The world will be a better place with such information democratized

  • There would be a lot of competition even without DeepSeek. Workers can freely exfiltrate trade secrets without noncompetes in California.

> have not been a winner-take-all runaway acceleration game where catchup is impossible

From the Mistral site:

> ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe.

It is pretty capital intensive!

  • I’m pretty impressed that they managed to get that close to the frontier with such a small cluster!

    • > I’m pretty impressed that they managed to get that close to the frontier with such a small cluster

      Chinese companies also managed to put together their models with relatively small clusters.

      Perhaps US companies are desperately trying to brute force their way into workable models?

  • According to Grok thats 7-10 MW. Tiny numbers.

    To put that into context, the last wave of capacity SpaceXAI added 400-450 MW.

    • But how much of that are they using for training versus inference? They're serving quite a large user base.

For sure people who don't grasp the difference between models, might be stuck in 'good enough' models.

But Opus 5.5/GPT is such a game changer in comparison to sooo many others, its still a moat for now.

  • You've got to think the comments about "good enough" are people who have not yet tried Opus 5.5. I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived.

    As for Mistral - I got really excited when they said Large 4 was focusing on being #1 in cybersecurity, because that's somewhere that they genuinely could edge out Anthropic & OpenAI. Have it actually solve problems, instead of Anthropic flagging "you tried to find a null pointer exception bug in your own code, we're now reporting you to the US government". But on the Mistral benchmarks I'm seeing, this looks very disappointing, but at least they haven't entirely given up. I genuinely thought Mistral had given up on new general models. They need to learn the bitter lesson all over again.

    • I feel like the "good enough" argument isn't about how big the gap between models is but about how good they are at solving the tasks at hand.

      The capabilities of all models increasing so much all the time means there are simply less and less tasks you need a frontier model for.

      Even if Opus 5.5 is 500x better than Deepseek, if deepseek can solve all my problems, why do I need to pay for more?

      10 replies →

    • I use Opus 5.5 daily for my job. I am aware (and in awe of) it's capabilities.

      Look at the context in which I used that term 'good enough'.

      What i was saying is that there are tasks for which a dumber model can be good enough, and for organizations with sovereignty/ privacy concerns, those concerns can be strong enough to incentivize the use of a dumber model.

    • > I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived.

      I had the exact same experience. And unlike Fable, it doesn't gobble up your entire usage limit in a few hours.

      I always wonder what the "good enough" people are actually using it for.

      4 replies →

  • I agree that Opus and GPT are almsot surely better, but so many real users are nervous enough about giving Anthropic and OpenAI access to all of their internal information that they may be willing to stomach worse models if it gives them more security.

    The real question is if this model is good enough that it can still accelerate work, and not be a hindrance to real work like older Mistral models often were.

    If they can do that, they'll have customers.

    • The Navier-Stokes fiasco made me push for local/controlled models very hard. "Can't rule out" that they stole data (backed up by their backdoor offers of sharing credit).

      If these companies will steal from deep pockets like Disney or Sony (some of the most infamously litigious copyright trolls to ever exist), they won't think twice of stealing every bit of code you upload to them.

      If your code passes through an AI company's servers, you can assume you just gave it to them. In turn, when your competitor tries to copy that new feature you just added, the AI is now trained in exactly how to copy you and eliminate your competitive edge. Unlike your employees, the AI isn't bound by the same rules and even if it were and violated them, your company probably doesn't have enough money to prove it in court (and that's if we somehow reverse some of the stupid "AI is the most transformative use of copyright I've ever seen" judges who have drunk the coolaid).

      Most companies could build the compute to run GLM or Kimi models for way less than the potential loss due to IP theft from using third-party systems.

      7 replies →

  • No it is not. Only maybe for the noobs or vibe coders.

    People who aren't afraid of rolling their sleeves into any code base? The difference is practically zero.

    • I agree; yes, I can see that they need a bit less hand holding each cycle, but I also see these "frontier" agents do some absolutely dumb shit that I have to correct and then I'm wondering if I'm the looney one here.

      Maybe it's because people stopped watching what their agents are doing and stopped looking at the quality of the output. But I still see agents being absolutely mindless like a junior dev.

      Recent example: it updated an an API to add newly released models to the backend. There's a list of models that require specific configuration for the reasoning effort and temperature or the API call fails. GPT 6.1 Sol misses this and code fails at runtime because the newer models need to be added to the list for special handling of temp and reasoning. Fixes it for one model and tests it for that model using an E2E test. But doesn't test the other models that were added for the same error condition...I had to explicitly ask it to do so and it finds them and adds them to the list and says "that's on me."

      Yeah, not that smart.

      1 reply →

    • Try a bigger code base or more complex stuff and you will easily see that the solution, speed and amount of problems Opus5.5 solves vs older models is relevant.

      1 reply →

    • I’ve been saying that. When you have no idea what you’re doing, you *need* the latest greatest model because it’s the only way to reduce errors.

      For people who have some expertise, the models accelerate the grunt work, but you’re the one validating it.

  • Please, give it another 6 months and they catch up. The American labs are currently trying everything they can to block others instead of advancing their models, trying to build an artificial moat. The American models are not that great, they are good, and they have a lot of agentic workflows in the back, but its basically a hardware limitation at this point. Once the HW makers catch up, and we can move away from the Nvidia monopoly, things will speed up quite a lot IMO.

  • It's an improvement, but game changer might be a bit of a stretch. If I lost access to Anthropic or OpenAI models tomorrow, I would be annoyed, but would reach for a slightly inferior model. Last year I wouldnt be able to say the same, and rhe challenge is that the moat is drying up fast. Whether its general improvements in model training by other competitors, or straight up distillation of SOTA models, the moat is shrinking and the available capital and spend for American model providers is going to dry up quickly as competing good enough models are adopted by more consumers.

    It's especially the case as more non-Americans look to self hosted models and domestic cloud inference providers using open models that the US providers who are still leading the charge need to drastically drop their prices and find a path to profitability in order to maintain their lead and retain the advantage they had as AI turns into a commodity (which is happening faster than I think even the frontier labs initially predicted).

  • I use Opus 5.5 at work.

    I use MiMov2.6Pro, DeepSeekv4.1Flash, GLM5.3, Hy4, Qwen3.8 and KimiK3 at home. Opus5.5 is not a game changer.

    • I do a broad amount of diverse experiments/projects I always wanted to do and throwing Opus5.5 against it just works

      I have to admit, Sonnet got really good too.

      But Opus just uses tools, a broad spectrum of it, etc. it feels like sure if you add some router behind it you could split it up if you need to but if you give me the choice, its opus allll day long.

  • > X is such a game changer

    I hear this literally every other week about whatever the newest FoTM model is.

    Unless you can provide concrete examples of things you can do with them that you simply couldn't do with last week's model, it's absolutely meaningless.

  • Yeah, I agree with this. I think the "the models are good enough" narrative is a myth. I've heard it so many times over the last year, but the model number keeps changing...

    There is no ceiling on what you can accomplish with more intelligence, so there will always be a market for the best models, and that market is likely to just keep growing. If Opus 13.5 can one-shot a profitable company or discover a new disease treatment or whatever you can think of that a swarm of relentless super-geniuses could accomplish, companies (and governments) will throw money at it.

    I also think there will always be a market for many sub-frontier models that will continue to grow rapidly as well, because "good enough" is definitely a thing for a given task.

    • One could still argue that models are good enough for a given task. I primarily use Opus at work for writing code and I realized that for my usage the intelligence of Opus 4.8 is more than enough. Sure the newer models are better but I can still do my work with having access to newer models

  • Despite Opus 5.5 got really bad the last days for me. Looks like they nerfed it again. This is extremely unreliable.

    • Or maybe they secretly believe you are trying to distill their models and are deliberately degrading your experience. Who knows with them?

      1 reply →

It's shaping up to be much more like a game of 'chicken' where each company tries to raise more cash without going bust... Ultimately the game of musical chairs is going to have to stop. In the US it looks like they are trying to get a government sanctioned truce in the form of regulation. That's what 'Pacing the frontier' means...

Even "runaway acceleration" isn't instantaneous. People imagine the singularity as something that happens almost instantaneously. But obviously it happens over time, and that time might be decades. It might still end up looking like a vertical line on a long-term graph.

If the singularity is defined as an AI sufficiently intelligent to improve itself independently, that AI is still limited by the resources required to do this improvement.

Am I reading this correctly?

This appears to be roughly as good as Sol 6.1 (which is quite good), considerably faster in terms of wall clock for complete tasks, and considerably cheaper (where Sol 6.1 is already good value - just really slow).

That seems too good to be true...

But I really hope it is true...

Hear! Hear! I really want European models / AI labs to succeed.

I trust them and their populations to provide a more societal-friendly version of AI, putting pressure on the US tech oligarchy, while also providing democracy-friendly open models that I don't trust to happen with the Chinese labs.

It will become winner take all when AI companies manage to really get value from user logs.

Right now they don't even get good feedback from local sessions - I can see it make the same mistake two days running, and then months later when a new model comes out, presumably trained on my data, it still makes the same mistake.

  • law of diminishing returns? i.e. any reasonable frontier lab will have enough user logs...

    • My theory is there is actual knowledge in the user logs.

      When a user says "I'm struggling to undo a bolt on my 1952 Mustang" and the AI responds "try hitting it with a hammer" and the user replies with "that worked thanks" - that is a tiny piece of knowledge which exists nowhere else.

      Future AI's can say with more confidence that hitting it with a hammer will probably work.

      Across billions of conversations, that can add up to more knowledge than all books.

I strongly disagree with this "early days" framing.

AI is an idea 60 years old. We are on the 3rd or 4th generation of AI development. Three years into the current iteration of products.

This is not early days by any measure. LLMs are a result of a very, very mature research field.

What do you mean "good enough"? Did you mean "large enough"? ;)

Disclaimer: I'm not sure how much of an IYKYK factor applies to this joke.

We haven't reached RSI yet. Once any entity reaches RSI, the runway scenario will happen.

  • Truly, this is what the Lord's prophets have revealed to us! (Eliezer 11:52) Keep strong in your P(singularity), for when the Kingdom arrives, He shall judge us in His righteous glory, whether to eternal annihilation, or rebirth and life in His Memory Eternal!

  • Assuming RSI is something that is possible as you envision it in the near term. I think that it will happen at some point, but I think we could still be a long way off. I don't think anyone can truthfully say that it is right around the corner.

Especially with Mistral taking a fraction of the investment of the big guys. They can maintain the position pretty comfortably just by staying within a standard deviation of the leaders.

> Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.

Mistral is also an European company. As we live in a time where the US regime is engaged in pyrrhic geopolitical tactics, it's good to know that it can't threaten to cut access to models during s period where everyone is rushing to incorporate them more and more in our life.

I think it's a mistaken belief that AI as we found it is the exponential runaway train.

So it makes sense, since all you need is compute, that there's a ceiling and specialization is going to be more valuable then some super AGI.

Especially since the worst people seem to be the ones who think they'll all run away with the bag.

I mean, Mistral is about 9-12 months behind here when you look at its overall benchmarks versus the models released around a year ago.

  • Sounds ok to me. Claude was fine at the start of the year, and now with Mistral you also get EU sovereignty? I'll take that.

I don't think it is, and I think that is what will pop the bubble. All these companies have winner take all valuations, and that won't happen.

... unless they can legislate it, which is why they are flattering heads of state and scare mongering about dangerous AI.

On a purely technical level, maybe? But in terms of actual revenue, is there really any chance of anyone catching the big labs?

Obviously, this is only a valid question if you don't believe that open weights are about to eat their lunch and their revenue is about to collapse, or they're running a super unprofitable ponzi scheme propped up by investor money that's about to collapse like a house of cards. I don't find those positions credible at all though.

If you do, then this question isn't really for you, as I'm more interested in thoughts from those who think that OpenAI and Anthropic in particular are about to be the largest companies on earth in a couple years. Could anyone catch them at that point?

  • > But in terms of actual revenue, is there really any chance of anyone catching the big labs?

    I don't know about revenue, but I suspect multiple other labs are already beating OpenAI/Anthropic on profitability. Staying on the frontier is expensive, and it's hard to recoup those R&D costs when you have a bunch of other labs nipping at your heels.

    If you concede the previous point, then the only way for OpenAI/Anthropic to keep growing long term is to swallow the whole economy (i.e. mass job replacemnt), and that's a bet I wouldn't take.

    • I think the actual plan is to swallow a good portion of the job market. It’s the only thing that makes sense and I hear VC podcast debates on which percentage of jobs justifies the market cap.

    • Maybe. I can't freaking wait for the IPO filings so we can finally put all this to rest. (haha, like that'll actually put it to rest on HN, but at least we'll have better data)

  • The big lab revenue may not be catchable, but im not sure it needs to be.

    If they can carve out a niche of industrial and governmental partners who rely on them for sovereignty reasons, it may be enough.

  • They are very unprofitable…? I don’t think that’s really in dispute. We haven’t yet seen a profitable frontier lab and model pricing remains fairly subsidized

  • No, because compute, not model ability, is the moat.

    The second moat is convenience, which all the big labs make it (comparatively) easy to glide into their models.

Seems silly not to have predicted that 10 years ago. I feel like it's long been obvious that smarter models being available will mean way easier cheap synthetic data and access to tools that will speed up competitors as well as consumers.