← Back to context

Comment by tristanj

1 day ago

Including the three sources above, multiple others have reported that K3 self-identifies as Claude.

"I'm actually Claude - not Kimi". https://x.com/PimDeWitte/status/2077884701470040083

I regret to inform you that it is, in fact, real and from their own website - you don’t even need to try hard to reproduce it. https://x.com/PimDeWitte/status/2078105292965912690

lmao this is so funny, if you ask Kimi K3 for something with an empty system prompt it will consistently think of itself as Claude https://x.com/__alula/status/2078359305741275445

"I genuinely believe I'm Claude based on everything in my training" https://x.com/williawa/status/2077869021589033002

another "I'm actually Claude - not Kimi", including the system prompt https://x.com/jchudnov/status/2078661564803207406/photo/1

my first prompt to any Kimi model was K3 via Pi, some version of "hi kimi!!" and the response was telling me "I'm actually Claude."

this is not hard to repro, just use a system prompt that doesn't mention the model name.

that said, if they bootstrapped with opus 4.6 convo sft data they had sitting around... so what?

  • The main story is what isn't being talked about. Chinese labs exfiltrated trillions of tokens of high-quality output from Anthropic and OpenAI, through proxies and heavily discounted token resellers, which they distilled and used for training data for their own models.

    Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and OpenAI did), they distilled Anthropic's models to bypass the hardest parts of development. Chinese labs compressed 18 months of intensive research and development into just 6 months, and are now head-to-head with their American counterparts.

    Anthropic tried to complain about this unauthorized "token theft", but they burned too much public goodwill with BS safety restrictions and users don't care. The US government is too busy fighting a war to help. Chinese labs are offering highly capable, cheap, open-weight models; exactly what users want. The community is happy to overlook any questionable methods Chinese labs used to build them.

    The cope is incredible. There's people in this thread in denial that Moonshot AI is trained on exfiltrated Anthropic's model output, even when shown substantial evidence this has been happening since Kimi 2.X

    Chinese labs were even paying an absurd $0.01 per Opus tool call trace, to get the quantity of training data needed.

    Kimi K3 has reached the point of RSI, and no longer needs synthetic data generated by Anthropic/OpenAI models. K3 is now capable enough to generate, iterate, and improve its own training data recursively. The data exfiltration is complete.

    We witnessed the most extensive industrial espionage campaign, probably ever, and nobody in the industry cares at all that it happened.

    • I could perhaps get myself to care just the tiniest bit if the information that was supposedly stolen wasn't generated by "stealing" from everybody else. Either it is fair use to train AI models on whatever information you can get your hands on for everyone or for no one.

      4 replies →

    • Ask claude its name in Chinese and it says Qwen or Deepseek. Anthropic distilled Chinese tokens rather than create their own Chinese language training data.

    • >The community is happy to overlook any questionable methods by Chinese Labs.

      Using the US-based models are arguably even more questionable. You have to be content with the OpenAI and Anthropic literally scraping the entire internet. They've all pirated content, scraped against ToS, ignored robots.txt, bypassed paywalls all to train their models. It's well known these AI labs have ingested the entirety of Annas-Archive into their models, the largest collection of books ever assembled.

      They didn't credit or compensate literally any artist, author, scientist or publicist in the creation of their models.

      They've lobbied against local Governments to shove big and loud datacentres in peoples backyard. They've polluted local water supplies, they've doubled energy costs for these regions. They've tormented locals with subsonic frequencies.

      They've given access to the DoD to use their models to kill people, or assist in killing people. They've used these models to enable mass surveillance, allegedly not domesicially but since when can we trust any of the 3-letter agencies.

      Using "US Models" is not the moral high ground you think it is. Kimi saying its Claude, Gemini or ChatGPT is not the "substantial evidence" you think either.

      >We witnessed the most extensive industrial espionage campaign, probably ever, and nobody in the industry cares at all that it happened.

      Because they stole for every one of us without permission. Thousands of my comments on this site and others (Stackoverflow, etc) are all used in their training data.

      Not to mention, OpenAI has allegedly just stole tons of internal Apple documents... I guess we'll just ignore that too.

    • Shrug.

      Hard to feel sorry for companies that created their empires by ignoring copyright themselves.

      Also, 'most extensive industrial espionage campaign, probably ever' is absolute nonsense. They did not need to infiltrate the companies for this nor are you accusing them of stealing any trade secrets. This is only about whether they looked at their competitors' products from the outside (in the form of conversation tokens) and used it to improve their own product (by training). Hardly the crime of the century.

      31 replies →

    • assuming the k3 model weights do indeed get published, if your model of the world is "achieving RSI is beneficial and K3 has done so," this feels structurally different from ordinary industrial espionage, because the knowledge has enriched the commons

      more like silk than capacitors

      if, again, your model is that RSI will be beneficial, why wouldn't making it available to all unlock more benefit globally than not doing that

      1 reply →

    • >through proxies and heavily discounted token resellers

      Could you explain a little more about how this works? Are you saying that the Chinese run or have backdoored something like OpenRouter?

      2 replies →

    • Why should anyone care? I couldn't give a single fuck, in fact if what you assert is true (definitely not proven), I applaud Moonshot - seems like a very smart way to operate.

    • > We witnessed the most extensive industrial espionage campaign, probably ever

      This is the funniest way of saying “going to a company’s website” I have seen in my entire life

    • > nobody cares at all that it happened.

      Who in their right mind would care? Why care? Misplaced patriotism?

      "A thief who steals from a thief has 100 years of forgiveness". Spanish proverb.

      In fact, I would be very concerned about the sanity of someone who cared about this sort of thing, unless they were Dario themselves.

    • > nobody cares at all that it happened

      Oh, no. I wouldn’t say that. If that happened, I definitely care: I’m positively delighted about it.

    • They stole from me first. And are spitting in my face and telling me they’ll take my job while they do it. I have negative sympathy for them.

Moonshot AI should have made it identify as Mythos as a practical joke to make US go crazy trying to figure out how they got access to it.