Comment by ls_stats

2 months ago

America needs its own DeepSeek or Z.ai, a lot of people (myself included) root for open chinese models to win because they have no other choice.

Thinking Machines might be it.

I don't hear about them a lot but it looks like arcee.ai is aiming to be just that.

Here are some of their current open weight offerings: https://www.arcee.ai/open-source-catalog

  • You don't hear about them much because their models aren't really competitive. I really wanted to try Trinity Large as a daily-driver in the MiniMax M2 sort of niche but I couldn't make it through a single day. The models need another couple point releases worth of post-training to make useful agents and if memory serves they weren't any less slopped in writing style and those are really the only two things people look for in models.

It could be but there are a host of companies going after open weights models: Arcee, Reflection, Llama (TBD on Meta's focus on closed-source versus open-source), etc.

That said, the fine-tuning API + open weight model at least is a semblance of a viable business that could work so I will be curious about it. I'm not sure the synergy is fully there (why is someone with an open weights model privelaged to fine-tune it better if it's just QLora or Lora) but let's see!

  • I don’t really get the business plan for open weights model companies, is the idea companies would pay them for serving?

  • Llama is dead. Meta is now releasing proprietary models (Muse Spark).

    • They’ve made some wishy-washy statements about their intention to release a future version of Muse Spark as open weights. We’ll see.

  • > It could be but there are a host of companies going after open weights models: Arcee, Reflection, Llama (TBD on Meta's focus on closed-source versus open-source), etc.

    my bet is that Chinese government fund Chinese models way more compared to what those companies receive (except llama, which is outdated but was strong foundation at its time)

    • The story of Reflection AI is supposedly that the company was faffing and failing at winning in the coding agent space, but was introduced to Jenson, who suggested they build an open-weight model and said he would fund it. That turned into a $2 billion financing with NVIDIA doing roughly $500 million and was a complete pivot.

      I think the bet would have to be that a US Open Weight company either: 1. Gets a lot of money from Jenson who views them as a counterbalance to the big labs in his ecosystem and a way to generate leverage (the same way he is positioning neoclouds-- it also could be synergistic with neoclouds who could offer the model serving endpoints) 2. Can fast follow the same way Mistral does (which, honestly, seems like just distilling the Chinese model, which distills the US lab but is pretty innovative on a whole lot of architecture both in training and serving land.) 3. AND figure out some (maybe not super lucrative but lucrative enough) sort of business model, as well. There are lots of possible business models, so I will be curious how this whole space evolves.

      6 replies →

AllenAI is also one to keep your eye on. Founded by Paul Allen of Microsoft, they are one of the best teams working towards truly transparent / open AI (including training data)

  • I love Allen AI.

    I find it wonderful that, as a non-profit, they are only one to two years behind SOTA models that cost billions of dollars to build, if not more.

  • AllenAI is great, but they don't have the budget or remit to build large models.

    • I wonder if the recent sale of the Seahawks will change that. IIRC, ~$10B and all is supposed to go to charity. Not sure how much of that will go to AllenAI, though. (If any.)

      1 reply →

What is the business model for an open weight model?

  • The same business model that Deepseek is using.

    Open-source models + services. This is more attractive because it doesn't lock in the vendors. If I grow larger, I can decide to deploy the open-source models.

    • So they're constantly hemorrhaging their most valuable clients?

      Tech history is littered with the corpses of "open source but we sell hosting" services. Models are so expensive to train, you can't be losing the big clients once they get super profitable.

      4 replies →

  • To compete against America. If your country has something like DeepSeek you really can't afford to let it fall as it's your best leverage if the US government decides to ban companies in your country from accessing American LLMs. And this is why there will never be a "DeepSeek of the US."

    • Considering how volatile things can get depending on who's president, I'd say even American companies need to "compete against America" if they don't want to get their rug pulled from under them (which, apparently, the legal system allows to easily happen in the US).

      1 reply →

  • In the US, there isn't one, which is why nobody in the US is currently doing it at frontier scale. And the people that were doing it stopped.

I’m trying to be charitable but your comment reads as “China bad” propaganda to me. Who cares that DeepSeek and Z.ai are Chinese companies?

  • It doesn't matter until it does. If the chinese government decides that open weight model releases are no longer allowed, that's a lot of companies that can't release new models. Same with the US government, etc. Having diversity is important.

    • However unlike the US models, China banning the release of new models would not break existing ones. Betting on US models only can get you locked out in just a few hours.

      1 reply →

    • It's a similar problem the human DNA solved by telling our teenage selves that our parents are dumb and we needed to move to a new tribe. Genetic diversity, but a digital equivalent.

  • China’s got absolute control over its outputs. For America to have any guarantees around long-term availability of OW models, they need domestic production.

    FWIW this is the same logic for China’s need for their own OW models

  • LLMs aren't just for coding and math. Many people understand the world through LLMs, even when it comes to philosophy and politics.

    If you understand the world through a Chinese LLM, you are seeing it through a biased lens stemming from biased training data.

    (Also, in that way, having all major LLMs developed by the US carries a risk too. We need more diversity than just the viewpoints of the US or China.)

    • All 6 UN languages should have their own dedicated LLM at the very least: American for English, Chinese for Chinese, LATAMgpt for Spanish, Russia has their own, Mistral for French, the Arabs don't have one yet I think, then there's sarvam for India and South Korea's sovereign models like SK telecom's

  • I think practically every government will want to put restrictions on private companies building models.

    Frankly the EU and the US will practically be less involved and have more pushback from the public in this than China. I think that’s less “China bad” than recognizing that China is a more authoritarian state and has far more proclivity to interfere than western states.

    Maybe I’m wrong? What does deep seek say about Tiananmen square in 1989?

  • Try selling SaaS for finance (think Private Equity/Wall St type customers) that is powered by a Chinese model. See how far you get.

Hopefully they'll release some smaller models (<100B) that we can run on home hardware at faster than 10tok/s.

Its not as good as GLM 5.2 for agentic workflows while also being bigger. Competition is going to be ruthless because the super low cost to switching.

There is also AllenAi in the US, but they have yet to produce a model at this scale. Thankfully, new contenders can come out of nowhere and do well, as long as they can produce a competitive model.

  • > Its not as good as GLM 5.2 for agentic workflows while also being bigger

    GLM 5.2 underwent extensive post-training and iteration since its original release to reach its current state. This seems like an extremely strong model for a first release, with a lot of potential for improvement, just like DS4.

    Sometimes I wish Meta had stuck with Llama 4 a bit longer to see how much further it could be pushed.

    • Llama 4 wasn't deemed a success, and Meta pivoted away as its now former head of AI couldn't demonstrate, nor even showed interest in, business profit.

      They overspent on llama 3 anyway so money ran dry, LeCun is good at running research, but budgets didn't stretch. Meta isn't investing in frontier big models anymore.

      1 reply →

    • Llama 4 was a bad architecture.

      Meta Spark is moderately promising but of course closed source.

Also the fact that China is building solar power like crazy: that makes it fantastically more well spirited an endeavor to wish well.

Do you think American companies will secretly distill frontier models to build open weight ones?

unlikely I think, they're likely doing this to garner some interest in their company but they seem pretty interested in revenue (judging by the companies they're working with)