← Back to context

Comment by m-hodges

16 hours ago

> We discovered that Moonshot AI, the company that produces the Kimi family of models, silently forwarded customer requests to Claude, instead of processing them using Kimi. Moonshot then displayed Claude’s responses to users. These users thought they were using a Kimi model, but received responses from Claude instead.

> DeepSeek also silently relayed exchanges to Claude without informing DeepSeek customers.

> MiniMax built its own proxy network service through a shell company. This shell company has no obvious links to MiniMax and does not disclose its relationship to its parent company. This shell proxy network service only offers access to models developed by Anthropic and OpenAI. The service does not offer access to any Chinese models, including Minimax’s own.

DeepSeek, minimax and so on have razer thin margins but unlike openai and Anthropic they are actually making some profit. Doing this doesn't make any financial sense.

Maybe Anthropic is confusing Chinese AI providers with token resellers using the same alibaba infrastructure? Or maybe something like openrouter was switching between operators depending on price/demand/availability?

Also, how can Anthropic have such accurate information about state actors and cybercriminals? This is the same company that hacked itself and realised that first months later..

  • My understanding of what Anthropic are saying about this is that the labs in question aren't forwarding things to Claude to make money nor even to look better to the customers whose queries they forward to Claude but to get access to conversations between real users and Claude, which they can then use to help train their own models.

    (I do not guarantee that I'm understanding right, and still less do I guarantee that what Anthropic say is actually true.)

  • Maybe there is some truth in that reselling Claude subscriptions/trials/api bundles via third parties breaks Anthropic's ToS. The rest is putting a maximum spin on it in order to achieve the political goal of banning Chinese AI. Anthropic is a highly ideological company and they are convinced that they are just in what they pursuit.

  • I’ve seen the supposed Kimi thinking output yap about Anthropic’s guidelines and whatnot on many occasions - could also be the result of distillation, but also that straight up being Claude’s output.

    To be honest I've also gotten Kimi to do an okay proof of concept for SQLi though mostly in a more defensive role, like "Let's see how big of a problem this is", while Claude complained about CVP on the same task.

  • You do this to distill a model.

    You can submit your users' questions async too, but if you do it sync, then you can also RLHF on the users' behavior after the output.

    • Ah, that makes way more sense than Anthropic's (probably deliberately misleading) insinuation that Moonshot has been burning millions of dollars in Claude API credits by swapping in a slightly better but infinitely more expensive model just to trick their users.

      I get those A/B responses chatting in Gemini fairly often, and I really don't think I'd feel deceived if I later learned one of the choices was actually from a competitor's model.

      2 replies →

  • > Maybe Anthropic is confusing Chinese AI providers with token resellers using the same alibaba infrastructure? Or maybe something like openrouter was switching between operators depending on price/demand/availability?

    Or maybe Anthropic is scared shitless of those competitors and is trying anything to smear them.

    Don't forget their goal is to ban open source and foreign AI. Being the sole legal provider is their business plan.

Consider me incredibly skeptical of any of these claims.

  • Same, I don't even see how that would work since you see the full thinking traces in Kimi but are hidden with Claude.

    And the Deepseek one sounds even more dubious as Deepseek is one of the cheapest model around, why relay anything to a more expensive model? I'm sure even the gray market Claude prices are still higher than Deepseek.

It seems hard to believe they could expect to get away with this, given model-to-model differences in writing style.