Comment by flatline
12 hours ago
I knew when I wrote that it was a bare assertion, based partly on memory. This is an approximation based on a few sources, the principal of which was this article, which pulls from a bunch of other sources in turn.
12 hours ago
I knew when I wrote that it was a bare assertion, based partly on memory. This is an approximation based on a few sources, the principal of which was this article, which pulls from a bunch of other sources in turn.
Oh, it's the OpenRouter number: https://finance.yahoo.com/technology/ai/articles/china-ai-mo...
Those numbers aren't credible IMO because OpenRouter only see traffic for people who have chosen to route their traffic through OpenRouter. If you do that, you're much more likely to be experimenting with alternative models. They have no insight at all into people who point their applications directly at OpenAI or Anthropic without having OpenRouter in the middle.
I agree about OpenRouter. The AI Gateway number [0] is likely the figure that was actually coming to mind. Moreover, Qwen models alone have overtaken the previously-dominant Llama models in hf downloads by quite a margin.
Real question, and a refinement to my previous statement: would you find it more surprising if over 25% of worldwide inference was running on Chinese open-weight models, or not? I personally would not be shocked.
[0] https://vercel.com/blog/ai-gateway-production-index-july-202...
I wouldn't be too surprised by that, given both the size of the Chinese market and the enormous price discount you get compared to the US models.
Within China itself, inference is overwhelmingly on Bytedance models which, by the way, are just as closed as those of Anthropic and OpenAI. They are integrated into everything, not just through a dedicated app, the way Gemini is integrated into Chrome.