Inkling: Our Open-Weights Model

2 months ago (thinkingmachines.ai)

Very nice, multi modal, largest open weight model that supports audio. Would be interesting to see how good the audio capability is.

If you want to run locally, checkout https://github.com/danielhanchen/llama.cpp/tree/add-inkling https://unsloth.ai/docs/models/inkling https://huggingface.co/unsloth/inkling-GGUF https://huggingface.co/unsloth/inkling-NVFP4

This supposedly is better than KimiK2.7, as much hype as GLM5.2 gets, I find myself using KimiK2.7 half of the time, so if the benchmark is true, then this can definitely go in the mix. My hope is that it might have strengths in some areas to beat all other open weight models.

  • Not to mention - it is American. This is the first competitive non-Chinese open weights model since what, Llama 3?

    • North Mini Code by Cohere (HQd in Toronto) has honestly been very competitive in my personal assessment with many of the models coming out of the PRC. I'd position it below Moonshot AIs and Z.ais recent releases, but above the varieties of Qwen, Deepseek, MiMo, etc.

      Depends whether America the continent or just the United States counts of course.

      3 replies →

    • +1 I enthusiastically use Chinese open weight models for a wide range of tasks (I also love Opus and Gemini) but I am so happy to see another high quality American open model (I consider gemma to be high quality, like qwen).

      I enjoyed turning off web search for Inkling to experiment with what innate knowledge is encoded in the model weights. A fun thing I do is check what innate knowledge very large models contain about me, as an individual. Inkling has an interesting concise shadow of what I do. (I have written a lot of books, so I am in training data.)

    • There is also NVIDIA Nemotron 3 Ultra, with 561B parameters, which was released a month ago.

      I do not know yet how smart it is, but the NVIDIA LLMs are very well optimized for fast inference (on their GPUs of course).

      Previously that was the biggest American open-weights LLM.

    • What does it mean, if it is American?

      Is it censored or will it eventually stop working in Middle Eastern countries?

      Or is it biased towards powerful political lobby group interests?

      When weights are open I usually don't care where is it from, as long as it is working for my use cases well

      8 replies →

    • "Competitive" is doing a lot of work there — it debuts at #41 on the AA index and the post itself calls it not the strongest overall. Mistral and Cohere's North are also non-Chinese open weights, so Llama 3 wasn't the last.

    • Llama 4 was unfairly hated on. Still the longest context window on any LLM ever (so what if you can't use it properly?) and unironically had decent image capabilities compared to llama3 which had none.

      Benchmark cheating aside, it wasn't that bad.

      2 replies →

    • GPT OSS was post Llama 3 and pretty strong at the time. But yeah this is the first seriously competitive non-Chinese open model in a good bit now.

    • > it is American

      It's also–hopefully–run by a cooler head. Altman launching a nuke into his own backside with his "stop me before I shoot grandma" routine was at least novel. Dario repeating the same playbook to the same effect years later still genuinely confounds me.

      If Thinking Machines pans out I could see it finding a welcome home at Apple.

      1 reply →

  • > This supposedly is better than KimiK2.7

    How can you tell?

    I just looked at the benchmarks and was kinda disappointed that it seems to be between KimiK2.6 and KimiK2.7 on most of the benchmarks.

    Do you refer to what it feels like to use the model? Or are there other benchmarks I haven't seen?

    • Most of the random comments you read on HN and reddit about how nice/bad various LLMs are, is basically based on the commentator's "vibe" about it, and almost nothing is grounded in evidence or actual usage. Don't read too much into it, want to know how good a model is? Run it with your own non-public benchmark, basically the only way to get proper answers you can somewhat rely on, everything else is manipulated, misunderstood or over-relied on.

      1 reply →

  • Oh thanks for sharing! The llama.cpp PRs should generally be fine for now - I'm fixing a few small edge cases as well!

  • I'm sure it's better than KimiK2.7 and GLM5.2. Benchmarks aren't the full picture. Despite GLM5.2 performing well on benches and supposedly near frontier, in reality it was nothing close to frontier in actual usage.

    • Have you used Inkling enough to be able to tell? Or how can you be sure? Please add some substance before making such claims

America needs its own DeepSeek or Z.ai, a lot of people (myself included) root for open chinese models to win because they have no other choice.

Thinking Machines might be it.

  • I don't hear about them a lot but it looks like arcee.ai is aiming to be just that.

    Here are some of their current open weight offerings: https://www.arcee.ai/open-source-catalog

    • You don't hear about them much because their models aren't really competitive. I really wanted to try Trinity Large as a daily-driver in the MiniMax M2 sort of niche but I couldn't make it through a single day. The models need another couple point releases worth of post-training to make useful agents and if memory serves they weren't any less slopped in writing style and those are really the only two things people look for in models.

  • It could be but there are a host of companies going after open weights models: Arcee, Reflection, Llama (TBD on Meta's focus on closed-source versus open-source), etc.

    That said, the fine-tuning API + open weight model at least is a semblance of a viable business that could work so I will be curious about it. I'm not sure the synergy is fully there (why is someone with an open weights model privelaged to fine-tune it better if it's just QLora or Lora) but let's see!

  • AllenAI is also one to keep your eye on. Founded by Paul Allen of Microsoft, they are one of the best teams working towards truly transparent / open AI (including training data)

  • What is the business model for an open weight model?

    • The same business model that Deepseek is using.

      Open-source models + services. This is more attractive because it doesn't lock in the vendors. If I grow larger, I can decide to deploy the open-source models.

      11 replies →

    • To compete against America. If your country has something like DeepSeek you really can't afford to let it fall as it's your best leverage if the US government decides to ban companies in your country from accessing American LLMs. And this is why there will never be a "DeepSeek of the US."

      2 replies →

    • Thinky has a potential answer in Tinker — give away the weights and charge for the SFT (and maybe RL down the line) to make the model more capable for specific tasks.

      2 replies →

    • In the US, there isn't one, which is why nobody in the US is currently doing it at frontier scale. And the people that were doing it stopped.

  • I’m trying to be charitable but your comment reads as “China bad” propaganda to me. Who cares that DeepSeek and Z.ai are Chinese companies?

    • It doesn't matter until it does. If the chinese government decides that open weight model releases are no longer allowed, that's a lot of companies that can't release new models. Same with the US government, etc. Having diversity is important.

      3 replies →

    • China’s got absolute control over its outputs. For America to have any guarantees around long-term availability of OW models, they need domestic production.

      FWIW this is the same logic for China’s need for their own OW models

    • LLMs aren't just for coding and math. Many people understand the world through LLMs, even when it comes to philosophy and politics.

      If you understand the world through a Chinese LLM, you are seeing it through a biased lens stemming from biased training data.

      (Also, in that way, having all major LLMs developed by the US carries a risk too. We need more diversity than just the viewpoints of the US or China.)

      1 reply →

    • I think practically every government will want to put restrictions on private companies building models.

      Frankly the EU and the US will practically be less involved and have more pushback from the public in this than China. I think that’s less “China bad” than recognizing that China is a more authoritarian state and has far more proclivity to interfere than western states.

      Maybe I’m wrong? What does deep seek say about Tiananmen square in 1989?

      7 replies →

    • Try selling SaaS for finance (think Private Equity/Wall St type customers) that is powered by a Chinese model. See how far you get.

  • Hopefully they'll release some smaller models (<100B) that we can run on home hardware at faster than 10tok/s.

  • Its not as good as GLM 5.2 for agentic workflows while also being bigger. Competition is going to be ruthless because the super low cost to switching.

    There is also AllenAi in the US, but they have yet to produce a model at this scale. Thankfully, new contenders can come out of nowhere and do well, as long as they can produce a competitive model.

    • > Its not as good as GLM 5.2 for agentic workflows while also being bigger

      GLM 5.2 underwent extensive post-training and iteration since its original release to reach its current state. This seems like an extremely strong model for a first release, with a lot of potential for improvement, just like DS4.

      Sometimes I wish Meta had stuck with Llama 4 a bit longer to see how much further it could be pushed.

      4 replies →

  • Also the fact that China is building solar power like crazy: that makes it fantastically more well spirited an endeavor to wish well.

  • Do you think American companies will secretly distill frontier models to build open weight ones?

  • unlikely I think, they're likely doing this to garner some interest in their company but they seem pretty interested in revenue (judging by the companies they're working with)

> Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning.

Open base models that can be fine tuned on Tinker is a great business model IMO. You (i.e. an enterprise) can own your own model & have it perform frontier-or-better at your task at potentially much lower cost and Thinking Machines gets to be your essential infra/service provider in this world.

Also,

> Inkling-Small matches or exceeds its larger sibling on many benchmarks — the result of improvements we made to the pre-training data and recipe for the smaller model.

Very cool! Excited to see the next generations of Thinky models.

  • > that can be fine tuned on Tinker

    Good source to understand why this is valuable?

    • Frontier models need to do everything for everyone. It's expected (though not often done) that smaller models fine-tuned on specific tasks can approach frontier performance on a specific area. [0]

      Post-training/fine-tuning is not trivial and having it as a service might make it more accessible.

      [0] https://surgehq.ai/blog/training-on-complexconstraints

    • If you want an LLM to have knowledge about something, the knowledge has to either exist in its weights or be provided to it in its context. Because context is expensive and limited, and models tend to get dumber the more their context is filled, there is usually more that you'd need to put into context than can reasonably fit in it in order for the model to answer questions about your data. So your options are basically

      1.) stuff it into context

      2.) figure out a way to determine what to put into context based off of what is being asked of the model (RAG)

      3.) change the weights of the model to have knowledge of your data baked into it (fine tuning)

      1 reply →

Here's a pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

What strikes me the most is just how many different tasks are involved in modern model design. It used to be the case that you come up with a new loss function, slight architecture changes, etc., run your train and eval loop, and publish the artifacts.

Now, there’s so much work to do just to keep up. It’s the ultimate red queen race. All of the 500 steps involved, each of which is its own little optimization loop, is sort of awe inspiring.

But obviously this inverts the previous rules that small teams run faster than big teams. AI requires a big team. It’s only once the team pushes past the 1000s that organizational inertia seems to become an issue. Because until then, there’s way too many pieces for even a dozen super stars.

Very preliminary testing so far, but there is something here, far beyond what the benchmarks suggest. Only ever saw such outperformance of public evals vs my private ones with Anthropic models and while it is far to early to make any judgement at this stage, this model will take up a lot of mine time in the coming weeks by the look of things. Only ever viewed Moonshot AIs models as something I'd be able to live with open-weight-wise (Z.AIs output simply does not perform as well in my task set), but this has the potential to be the second. If Mistral came out with something like this, I suspect every Europhile (me included) would never stop talking about it.

  • Quick and still very early update, the model has (with web search disabled which was verified via the reasoning traces) accurately answered a number of questions focused on very niche details (engine specific maintenance in certain newtimers, very niche bag construction and material details) that I have only ever seen Gemini 3 and 3.1 Pro get correct. Neither Fable 5, nor GPT-5.6 Sol or any other model by any other lab has ever provided accurate information without web access for these specific questions for which an objectively correct answer absolutely exists and is general knowledge if one is versed in the specifics.

    Being ahead of Fable 5 in any task, that is not included in public benchmarks and thus could be overfitted for, is impressive to say the least. Last time a model exceeded the expectations I had based on the release notes to such an extent was Haiku 4.5, which I still wish we got a solid replacement for.

For a first model, and given it's open, I am gaining some faith in American Open research labs again...

I couldn't test it since it's not on openrouter or something, but even if it's only as good as GLM5.1 that's more than good enough first attempt, I think.

Perhaps a lot more labs will catch up to ballpark frontier esque level soon, I am all for more competition in any field.

It's nice to see a strong long context open weights model that is multi-modal.

There are many applications that will benefit from the strength in audio here and until z.ai and co work in visual this could be very strong for general agentic applications, though I see there's a bit of weakness in the benches for areas that might make that less true.

Like all models need to slap it in your harness and do proper evals on the tasks you care about.

  • MiniMax M3 and DeepSeek v4-Pro are highly capable long context open weight multi-modal models. But long-context is a trap, because performance still falls dramatically after 150k-200k context.

    • > But long-context is a trap, because performance still falls dramatically after 150k-200k context.

      I often see this repeated, and it is not true task to task. I work on this daily and we have several tasks where long context is advantageous and our evals against a whole battery of models with different windows show it as being so.

      This is why having good evals for the tasks you're working on is so important.

      I do grant it's a good rule of thumb.

      1 reply →

    • > But long-context is a trap, because performance still falls dramatically after 150k-200k context.

      I'm not sure exactly what causes the difference, but this heavily depends on the model. In my experience with Opus 4.8, I can go well over 500k and still get extremely good results. A drastically different example was GLM-5.1, which worked great until about 100k and then turned insane almost immediately. They did fix that with 5.2, though.

      2 replies →

Seems like this is particularly good at instruction following, but not as strong at coding as others. It's always great to get more diversity of open weight models though! I'll need to test this out to see what its "personality" is like.

For the most part it’s better than Nemotron, worse than GLM. This makes it the best American open weights model from what I can tell?

  • It's nearly double the size of Nemotron 3 Ultra, so I'd expect it to be considerably better, although the active parameter count seems to be a touch lower at 41B vs 55B

  • I'm surprised that Nemotron gets mentioned at all. In my experiments with it for coding tasks it performed extremely poorly, essentially unusable.

    • I focus on realtime voice AI uses cases and nemotron's time to first token is INSANELY fast. It's become a legit option for voice use cases

“Alongside Inkling we are sharing a preview of Inkling-Small, a 276B-parameter Mixture-of-Experts model (12B active, vs. 41B for Inkling) with a different performance/latency trade-off.”

Buried at the end there is the details I was most interested in - a possible competitor for DeepSeek V4 Flash? Excitedly awaiting the release of the weights for this one.

I never thought i'd see the day they released a model, rather than a blog post. The Figure 3 demo being a screencap of chrome in localhost made me feel better about myself. Jokes aside, best western open weights model- very cool.

Interestingly, when opening this page, the first thought I had was not that the benchmarks should be high, but 'I really hope they did not benchmaxx'. I think a model with modest benchmark scores can have much better real world utility as opposed to the current frontiers that are RL'd into being robotic and rigid.

What are the different business models for open-weight AI companies?

  • For thinking machines, they provide super simple finetuning APIs.

    if it is their model, they can have more lower level integrations for that. Thinking machines might be the only large lab in the US to have business interest aligned with open sourcing strong models that are customizable.

  • Just serving the model over API seems like a natural fit and is what many of them are doing. So simply being the cloud provider for your own open weight model can be a source of revenue

    • But so can everyone else. What’s the moat for spending all those billions. I understand the Chinese angle, they need to undermine American models as a matter of statecraft, but what is the business model here? It just seems like VC charity.

      5 replies →

    • What is the moat? The time it takes for AI to rewrite an efficient inference stack for a new model? Considering most LLMs follow a similar architecture, adapting to a new model shouldn't take that much time.

      5 replies →

  • Similar to companies working on FOSS codebases, hosting (sometimes with the license restricting third-parties in some way), providing tailored models and services to customer's and getting bought for your team if your model happens to be competitive enough.

  • Maybe the thesis is that

    Open source low cost models will dominate most enterprise tasks as cost curves will dictate usage. TM is trying to replicate that especially as the US and China gets more defensive with their tech

  • - inference

    - RLaaS (Tinker, or the more involved FDE motion a la Reflection / Applied Compute)

Not compared against Gemma 4? That is a big omission.

  • Gemma 4 wouldn't really be a competitor to this. Gemma has the dense 31b model, but this has like 25x the total number of params.

The actual part on fine-tuning seems very short in the article. Did I miss a page where they have examples of fine-tuning it for different niche use cases?

Optimizing models to be fine-tuned is an amazing direction, but just makes me wonder how much better this actually is at being fine-tuned compared to other models. As none of the modern models are great at being fine-tuned afaik. Basically looking for some sort of benchmark showing that it's resistant to overfitting / catastrophic forgetting, etc.

Would be very interesting to see concrete demonstrations of different fine-tunes of the model. I'd imagine they've done hundreds of those internally.

  • They have good docs on finetuning in general here: https://thinkingmachines.ai/tinker/

    I used it last week for an application using a small model just as an experiment. It all went very well. The model did not turn out to be good though because my training data was of bad quality. I plan to work on it more this weekend.

The most important observation is that open source matches their business model, which is to provide fine tuning services for enterprises, etc.

That what makes this a (potentially) safer model to build on top of

Do you think the barbarians are at the gates of OpenAI and Anthropic? If cheaper, open weights models can seriously take revenue away from those two labs for (frontier - 1) model use cases (which are the models most enterprises will choose) then OpenAI and Anthropic are left only with users using their latest and greatest model AND who will keep upgrading to the newer ones?

  • Bull case: iteration becomes so quick that frontier-1 won't cut it. OpenAI and Anthropic are both betting on the singularity, I suppose.

    • The "singularity" as stated requires AI to either make a technological advancement strong enough to be deadly to humans (besides just intelligence), or spreading deep institutional support for itself among society.

      The idea that "we will get superintelligence first, then... ???" is kind of a weird notion. I mean, it's pretty arguable that we do have at least some form of superintelligence. The AI itself needs to actually do something with it though. Either that, or more likely, someone needs to do something bad with the superintelligence.

      That could be both re-assuring or not. Because under that view, given how AI is being integrated so quickly into society, it's not going to take this fantasy view of superintelligence to reach the singularity. If you have broad institutional support and crowd out the thing we call humanity over time (the two ways to 'solve' a problem: solve it, or declare it meaningless), that is another way to reach the singularity.

    • Without the singularity, I think Frontier labs will offer intelligent model blends. They'll have their own versions of "cheap" models and be expert at using the appropriate amount of compute for a task.

competition in this space is great, especially with open models/weights. I think the answer is not closed source models. Similar to the Unix versus Linux situation in the 1990's, open source wins out. Yesterdays story about how OpenAI has now began encrypting traffic between model and agent [0], this story brings a breath of fresh air. There is nothing "Open" about hiding the communication between model and agent, especially with software that is running within a trusted environment/network. It needs to be more transparent, not less.

[0] https://www.theregister.com/ai-and-ml/2026/07/15/openai-hide...

  • Open Source won out because the cost of compute fell through the floor. I'm not sure whether we're going to see a similar dynamic play out this time, although I would greatly prefer it to.

    • cost of compute, also the less likelyhood of underhanded tactics or vendor lock-in

Give me a good 180B param model that fits snuggly on an single DGX spark and I will sing your praises.

Smart that she says "not the strongest overall model, open or closed". This is a rare for an AI lab to say out loud. They basically decided to compete on customizability, and not on topping the temporary leaderboard. Also corroborates what we recently wrote: any Lab's capability lead cant hold for long anyway, it's a "red queen race" that never settles: https://news.ycombinator.com/item?id=48892559

This is a winner IMO. Lots of cost pressure on token spend atm within enterprises and tasks that don't require Opus / Codex class models.

These companies have hopefully captured all of their traces and now have enough to fine-tune an open model and host themselves.

Inkling feels like the right base - not obsessed with benchmaxxing on coding but rather being adaptable to the task required

For tasks like GTM, support, content writing etc. seeing 80%+ savings

  • Open weight models as a category might be a winner due to cost pressure, but I don't think Inkling is a top performer in that respect. Pricing on a token basis is 6-9x higher depending on the provider.

    Can chime in on the support use case specifically: GPT OSS performs really well here and has been somewhat of a benchmark with our customers, limited testing [0] against Inkling reveals basically identical performance, but with a significant cost increase at scale.

    I'd say that for real-world tasks that aren't coding most companies don't see value by being on the latest and greatest model.

    [0] https://valiopt.com/blog/inkling-model-customer-support-revi...

How much mortgage equity would I need to do that 27min fine tune demo on local :)

Self fine tuning like that though seems like a whole new set of possibilities unlocked.

Happy to see an open weight model ! This has all the right ingredients for success.

They also indicate they have a 276B A12B version, but it doesn't seem the weights are available. This might actually be able to fit in 128GB when quantized to 2 bits or so which makes it interesting.

Open weights are great, but open payment matters too. DeepSeek V4 is open-weight but closed-payment. api-hub.cc makes it accessible with just a credit card.

Interested in the implied strategy - that training a bespoke model for what you need will make economic sense over using a mass-trained model. I wonder if that's true?

  • Same. Gutsy bet to make in the face of Fable / Mythos, but the multimodal quality is at least a promising technical/ product story to tell. Everyone knows throwing Opus at everything is wasteful and domain expertise should live in the weights eventually; the question is whether foundation model scaling will slow down enough soon enough for that to matter.

    Or maybe this is just a warm-up / stopgap and Thinking Machines is betting on finding the next architectural breakthrough that lets it compete with the big foundation models?

is there something in the space of "taskifying" enterprises data for them? inkling on its own looks high-quality, but expecting companies to spend $ figuring it out before spending more $ on the actual fine-tuning job seems ... hard, especially if making the model especially customizable is the goal? or do the unit economics just work out with a small number of fine-tuners training a ~1T model on Tinker?

Giving the accuracy-token graph and not the accuracy-cost graph, thus we cannot easily compare costs with other models, is not a way to gain my trust

My personal bet is that this model should really shine in Autoresearch NanoGPT-style speedruns because its first-class integration with Tinker

Do they have an api to try the model in real envs?

  • They've got an openai + anthropic compatible endpoints. I got far enough to run some tests on the openai endpoint, albeit with some finagling (their /models list is empty, my tool auto-configures using that, was an initial stumble).

    • Thanks! I found the OpenAI-compatible endpoint and got it working. I ran Inkling on a couple of my own evals. It looks promising, but on my cases it still fell short of GPT-5.4 and GPT-5.6 Luna.

  The MoE design largely follows DeepSeek-V3

why is the model never compared to deepseek in their blog post ?

Lol slither.io is the new benchmark now? I guess my game slitherworld.com is now something that can be vibecoded too

I tried Hy3 today and liked it. It's a (small) step up from DSV4P.

Something on that level but multi-modal would be quite nice!

Very impressive model, exciting to see an American open-source lab with such competitive results.

If it's ~30% bigger and not as good as GLM 5.2, why would I tinker with this model?

Maybe for the multi modal?

  • > If it's ~30% bigger and not as good as GLM 5.2, why would I tinker with this model?

    The benchmarks never tell the full story. Some of the open weights models have been benchmaxxed for a while. Their utility on real work can be different than the benchmark number.

    The multimodal input is also a big deal. Having vision input is really helpful for a lot of tasks.

  • If they have a really seamless fine-tuning experience and maybe can help you extract the data you need to FT (which is one of the big challenges in actually getting fine-tuning democratized), maybe you would use it because "Tinker" defaults to it.

    The model could also be more flexible for non-coding use-cases (they show the results for reasoning being strong) so maybe the argument is to use it for non-coding use-cases to drive relatively deterministic conclusions for non-coding agents (they have also done some determinism work on kernels, which could be useful in pulling on that thread of deterministic models that are fine-tuned for everything that is not writing code.)

    That said, I'm not sure how much all the work they have done actually synergizes or if the market size (at least in the short to medium term) is big enough for a huge outcome from the company's current valuation with those bets as the enterprise agent estate is taking a while to evolve. Hence companies like Anthropic and OpenAI are throwing tons of consulting money at the problem.

  • There's also an Inkling-Small that is 276B, 12B active that is much smaller than GLM 5.2 and still multimodal. Not released yet, but in the announcement link they mention that they're testing Inkling-Small & will release as open weight after testing. That one may be interesting as a Deepseek V4 Flash replacement.

I think we’re going to start seeing more OSS models that perform especially well on certain tasks instead of trying to be generalists like the frontier models. That’s a winning formula because if you’re building an app on a model it often has a specific set of use cases

Excited to try out its capability, especially audio and video.

It's nice that it has a long context window, but in practice, I find I always have to clear context btw 150k-400k context even if the context window is 1M on paper.

everyone and their grandma shipping top tier models now. anthropic and openai trying to capture the app layer with their shitty 'super app'

  • We always have been, Big Tech has been extremely slow to catch up to the indies.

    Nobody is making money lmao.

    I would not bet on OpenAI creating any good products, they never have. They are like Meta in all this - never innovated anything themselves, can only acquire others to stay relevant. They'll never do an incredible consumer experience on the level of a PlayStation or Blizzard or even Google.

artificial analysis puts it quite low below previous gen open weights models from China. Why are people so ecstatic?

I really respect the epistemtics work here. It might become an accurate, inexpensive open-weight workhorse for high-level prioritization and decision-making work. (Finance bros will also love this)

[flagged]

why is this website ai slop

  • Is it really that bad? I always get the impression that their blog posts look especially beautiful with their font choices and overall design. They are typographically pleasing, and if I could, I would use this as the distraction-free reading mode for every web page.

    It feels like I’m reading a newspaper, but oddly, without them resorting to any skeuomorphic tricks.

  • do you want them to focus on the website or their model? do you buy a device because of its unboxing experience?

Raised 2 billion dollars at a 12 billion valuation and debuts at 41 on the Artificial Analysis Intelligence Index, while KIMI and DeepSeek will release Fable-class models this week. What a joke.

  • Moonshot (Kimi) has raised $3.77B and been around for >3 years, Thinking Machines raising $2B and releasing a decent open weights model in 16 months is actually quite comparable.

  • I think your comments might be overly negative. Would you expect the first model from an organization to top the chart? It's a process, I think they did pretty good. They have enough resources to continue improving on it

  • > ...while KIMI and DeepSeek will release Fable-class models this week.

    What new model is DeepSeek releasing? Their current V4 Pro at Max reasoning is consistently worse than GLM 5.2 at Max reasoning, though the latter is close to Opus 4.8 at Extra/Max reasoning, albeit a little bit worse in my experience (though if they gave comparable amounts of tokens to Anthropic 5x Max subscription I could see myself moving over, currently they give you less though even with their ZCode discount).

    In practical agentic development, none of those seem to be that close to Fable to me. Spent 181 million tokens with GLM 5.2 with ZCode in the past month, 142 million with DeepSeek V4 Pro with ZCode and OpenCode and about 3.45 billion across all Anthropic models with Claude Code, though understandably with my workload between 95-99% of them are cached (very docs/plan/tooling/read heavy work to limit slop, albeit with sub-agents and workflows).

    • DeepSeekV4 was a preview model, read the papers. It's not the final model. They released it to demonstrate architectural capabilities. They are still training and the model release is planned within the next month.

      5 replies →

How does this compare with Grok 4.5, Fable, GPT 5.6 etc I can use them for a few bucks a month, whats the benefit of using this model? is it more intelligent, faster, cheaper (and no I don't want to spin-up my own mini datacenter to run it 'in house') I want to install a harness/app/visit web page auth and start as 99% of people using AI want to do.