Comment by remexre
2 days ago
my first prompt to any Kimi model was K3 via Pi, some version of "hi kimi!!" and the response was telling me "I'm actually Claude."
this is not hard to repro, just use a system prompt that doesn't mention the model name.
that said, if they bootstrapped with opus 4.6 convo sft data they had sitting around... so what?
The main story is what isn't being talked about. Chinese labs exfiltrated trillions of tokens of high-quality output from Anthropic and OpenAI, through proxies and heavily discounted token resellers, which they distilled and used for training data for their own models.
Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and OpenAI did), they distilled Anthropic's models to bypass the hardest parts of development. Chinese labs compressed 18 months of intensive research and development into just 6 months, and are now head-to-head with their American counterparts.
Anthropic tried to complain about this unauthorized "token theft", but they burned too much public goodwill with BS safety restrictions and users don't care. The US government is too busy fighting a war to help. Chinese labs are offering highly capable, cheap, open-weight models; exactly what users want. The community is happy to overlook any questionable methods Chinese labs used to build them.
The cope is incredible. There's people in this thread in denial that Moonshot AI is trained on exfiltrated Anthropic's model output, even when shown substantial evidence this has been happening since Kimi 2.X
Chinese labs were even paying an absurd $0.01 per Opus tool call trace, to get the quantity of training data needed.
Kimi K3 has reached the point of RSI, and no longer needs synthetic data generated by Anthropic/OpenAI models. K3 is now capable enough to generate, iterate, and improve its own training data recursively. The data exfiltration is complete.
We witnessed the most extensive industrial espionage campaign, probably ever, and nobody in the industry cares at all that it happened.
I could perhaps get myself to care just the tiniest bit if the information that was supposedly stolen wasn't generated by "stealing" from everybody else. Either it is fair use to train AI models on whatever information you can get your hands on for everyone or for no one.
Training data that’s still sitting in the pages of a book is not really all that useful.
1 reply →
Stack overflow is pretending to be Claude now. I wonder if one can get it to say your question had already been asked.
“How dare they steal what we rightfully stole first!”
Ask claude its name in Chinese and it says Qwen or Deepseek. Anthropic distilled Chinese tokens rather than create their own Chinese language training data.
Why should anyone care? I couldn't give a single fuck, in fact if what you assert is true (definitely not proven), I applaud Moonshot - seems like a very smart way to operate.
Shrug.
Hard to feel sorry for companies that created their empires by ignoring copyright themselves.
Also, 'most extensive industrial espionage campaign, probably ever' is absolute nonsense. They did not need to infiltrate the companies for this nor are you accusing them of stealing any trade secrets. This is only about whether they looked at their competitors' products from the outside (in the form of conversation tokens) and used it to improve their own product (by training). Hardly the crime of the century.
> 'most extensive industrial espionage campaign, probably ever' is absolute nonsense
You are completely underestimating the scale of what is happening here.
Chinese AI labs are actively facilitating an industrial-scale network of tens of thousands of bot accounts, that resell Claude tokens at 97% below official API prices. They buy subsidized Max 5x plans (sometimes with stolen credit cards), then split the subscription across dozens of clients and reselling the output. They are running a massive data-harvesting operation. Chinese labs and token resellers subsidize the cost of the tokens in exchange for the API metadata (detailed reasoning traces, model outputs, and tool calls) to use as high-quality training data for their own models.
They are buying Anthropic's own product, just to resell it below cost, just so they can capture the training data. Reportedly, they are paying as much as ~$0.01 per tool call.
https://news.ycombinator.com/item?id=48664814
32 replies →
>The community is happy to overlook any questionable methods by Chinese Labs.
Using the US-based models are arguably even more questionable. You have to be content with the OpenAI and Anthropic literally scraping the entire internet. They've all pirated content, scraped against ToS, ignored robots.txt, bypassed paywalls all to train their models. It's well known these AI labs have ingested the entirety of Annas-Archive into their models, the largest collection of books ever assembled.
They didn't credit or compensate literally any artist, author, scientist or publicist in the creation of their models.
They've lobbied against local Governments to shove big and loud datacentres in peoples backyard. They've polluted local water supplies, they've doubled energy costs for these regions. They've tormented locals with subsonic frequencies.
They've given access to the DoD to use their models to kill people, or assist in killing people. They've used these models to enable mass surveillance, allegedly not domesicially but since when can we trust any of the 3-letter agencies.
Using "US Models" is not the moral high ground you think it is. Kimi saying its Claude, Gemini or ChatGPT is not the "substantial evidence" you think either.
>We witnessed the most extensive industrial espionage campaign, probably ever, and nobody in the industry cares at all that it happened.
Because they stole for every one of us without permission. Thousands of my comments on this site and others (Stackoverflow, etc) are all used in their training data.
Not to mention, OpenAI has allegedly just stole tons of internal Apple documents... I guess we'll just ignore that too.
assuming the k3 model weights do indeed get published, if your model of the world is "achieving RSI is beneficial and K3 has done so," this feels structurally different from ordinary industrial espionage, because the knowledge has enriched the commons
more like silk than capacitors
if, again, your model is that RSI will be beneficial, why wouldn't making it available to all unlock more benefit globally than not doing that
Sorry, but I have a thought that's off topic: If Kimi is good enough to improve itself, what's preventing someone who owns a big datacenter and nothing else, to just run Kimi to do AI research, thereby rendering the frontier labs more or less unnecessary?
>through proxies and heavily discounted token resellers
Could you explain a little more about how this works? Are you saying that the Chinese run or have backdoored something like OpenRouter?
Have a look at https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens... and https://x.com/yan5xu/status/2029743983522631698
Chinese resellers acquire hundreds of Claude Max 5x accounts and set up a custom proxy server. Customers point their ANTHROPIC_API_KEY at that proxy, and requests are routed to Anthropic through one of those hundreds of accounts. Because one $200 Claude Max 5x account gets the equivalent of ~$2000 in of API credits, these resellers can resell Anthropic tokens at a massive discount, undercutting official API prices by more than 90%.
To cut costs even further, these accounts are funded using educational discounts, startup credits, or stolen credit cards.
The resellers log all data traveling through their proxy networks, which they then resell to Chinese labs as high-quality training data for significant profit. https://x.com/xkajon/status/2050445443889525235
The resellers also loan these proxy networks to Chinese labs, allowing them to can run distillation attacks on Anthropic, while blending in with regular user traffic. https://www.anthropic.com/news/detecting-and-preventing-dist...
This is a widespread tactic, there's hundreds of proxy resellers operating. Some even offer enterprise SLAs.
1 reply →
> nobody cares at all that it happened.
Who in their right mind would care? Why care? Misplaced patriotism?
"A thief who steals from a thief has 100 years of forgiveness". Spanish proverb.
In fact, I would be very concerned about the sanity of someone who cared about this sort of thing, unless they were Dario themselves.
> We witnessed the most extensive industrial espionage campaign, probably ever
This is the funniest way of saying “going to a company’s website” I have seen in my entire life
> nobody cares at all that it happened
Oh, no. I wouldn’t say that. If that happened, I definitely care: I’m positively delighted about it.
They stole from me first. And are spitting in my face and telling me they’ll take my job while they do it. I have negative sympathy for them.