Comment by tristanj
2 days ago
API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870
This behavior is exactly what you'd expect from a model distilled from Claude.
There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/...
This analysis observed K3 identifies itself as Claude approximately 15% of the time.
K3 reproduces Claude's correct current model id, which the real Claude models themselves do not emit. This suggests K3 was trained on Claude data labeled with deployment metadata (API logs, tagged synthetic data), rather than Claude's chat outputs.
And there's an entire Reddit thread discussing Kimi's similarities with Claude https://www.reddit.com/r/LocalLLaMA/comments/1m2w5ge/did_kim...
This analysis shows K3 and Opus/Fable have unexpected correlated outputs https://typebulb.com/u/lab/you-re-relatively-right/full
Kimi calling itself claude means nothing. During pre-training, when the model learns to "simulate" the internet text, it will naturally be fed with a bunch of data about Claude and ChatGPT. With the amount of LLM outputs on the internet today, it is not surprising at all that a model would naturally call itself Claude or ChatGPT. You can mitigate that in post-training (or actually in pre-training as well) by training on many examples of what the model should call itself. That being said, getting probably hundreds pf thousands of ChatGPT and Claude examples totally "pirged" out of the weights is going to be difficult and really more hassle than its worth.
Sure, but then Qwen should leak that too, and it doesn't. K3 calls itself Claude 7 out of 48 times, Qwen does it 0 out of 48, and the only other model to identify itself as Claude is DeepSeek. and DeepSeek is alleged to also distill from Claude data anyway. So this isn't something every model absorbed from the same web text.
And you skipped over the strongest datapoint that K3 is distilled: K3 reproduces Claude's public model identifier under prefill (i.e. "claude-opus-4-5-20251101"). This data does not appear in Claude chat logs, only in API logs. K3 only does this for Claude models and not for any other lab. The real Claude models don't produce their own current public identifier, they only know their previous identifier (i.e. Sonnet 4.5 calls itself "Claude 3.5 Sonnet").
This is highly suggestive of the type of data that K3 was trained on. K3 was very likely trained on Claude metadata traces (API logs, tagged synthetic data). Not web chat logs, those wouldn't include this identifier. And this data wasn't filtered correctly, which is why K3 incorrectly identifies itself as Claude 15% of the time.
You can also look at the last link and it's pretty damning: Kimi K3's output has an uncanny similarity to Fable/Opus output. https://typebulb.com/u/lab/you-re-relatively-right/full
Did you know claude models identify as qwen or deepseek when asked in chinese?
3 replies →
> Qwen should leak that too, and it doesn't.
FWIW I had Qwen identifying itself as "a language model made by Google" in one conversation, although I could not reproduce this reliably.
1 reply →
This is good and interesting research, and I think it proves Kimi was trained on Claude's outputs, but only for English conversation and essays. Their style is similar, which is tbh somewhat expected for a Chinese company - I don't think most American companies (especially small labs making LLMs) could figure out a good way of talking to Chinese customers, without copying some of the existing Chinese models.
This doesn't imply that Kimi's reasoning capabilities or coding ability has anything to do with Anthropic (which is the most valuable part), or the model's strength comes from distilling Anthropic.
Considering by this benchmark, Opus sounds extremely similar to Fable, this isn't really evidence of a ton of Fable output being trained on (and even then, most likely for superficial stuff, like style)
I also have a suspicion that in original Chinese, these models don't sound anything like Anthropic's ones, though I have no proof of that.
https://qwen.readthedocs.io/en/latest/training/ms_swift.html
Qwen cares enough about model identity that their training framework and docs include a preset for training on it complete with a targeted dataset: https://huggingface.co/datasets/modelscope/self-cognition
And people get Claude to claim it's Deepseek by asking in Chinese.
I can't believe we're still at the "I asked the model who it is" stage of LLMs nearly 4 years out from models calling themselves GPT by OpenAI.
> Kimi K3 reproducibly identifies itself as Claude
It could also be have been trained from collected response datasets. Claude got caught several time responding it was ChatGPT or even Deepseek and I don't think Anthropic has been distealling DeepSeek.
> This behavior is exactly what you'd expect from a model distilled from Claude.
The opposite actually. If they wanted to distill Claude without getting caught they could just use a regex to change Claude to Kimi in their distillation pipeline!
Apt typo.
Though I am of the opinion that distilling is no different than how extant frontier LLMs have also been trained on other people's data, I could actually see the word distealling becoming useful in discussion.
Its not a typo, someone coined that during the DeepSeek R1 hype period and I kept using it since then.
I totally agree with you on the fact that it's not morally any different than pre-training. IMHO we should have a legislation that force base models to be released publicly without any restrictions whatsoever as it's basically the product of the whole humanity's intelligence.
Opus/Sol/Fable are valuable because of their reasoning and coding ability, not their bedside manners.
While still not okay, I suspect the latter is what gets stolen by Chinese distillation (and some evidence suggest this happens the other way round with US models talking in Chinese)
> they could just use a regex to change Claude to Kimi in their distillation pipeline!
Jean-Kimi Van Damme would like to have a word with you.
Some species of the Scunthorpe problem, then.
and claude will call itself chatgpt etc.
nothing new, all ai labs are immoral and not bound by any reasonable oversight or ethical constraints. All outlaws in their own rights on that front. Absolutely none of them have true rights on the matter of being distilled from given historic and continued behaviour. I'm not sure why this is a talking point at all? We know AI companies steal, the least interesting behaviour among this is them stealing from one another.
For me, a far more interesting and important point of conversation on this matter is anthropic buying rare or evwn unique books, processing them for training data, and then destroying the books for others cannot use it as well.
Permanemt destruction of priceless primary source materials is so many leagues beyond copying a copy that I cannot fathom it even registering as a discussion point.
> For me, a far more interesting and important point of conversation on this matter is anthropic buying rare or evwn unique books, processing them for training data, and then destroying the books for others cannot use it as well.
That's an incredible allegation, and appalling if true. But is it true?
It's not an allegation https://www.washingtonpost.com/technology/2026/01/27/anthrop... (if you're talking about the "rare or unique" part, yeah that might be bs)
But in my opinion, treating mass produced books like they're this sacred untouchable object is ridiculous. They're not "source" material, they're just a copy as well, and they're not "priceless" by any means. They're very reasonably priced, perhaps even so cheaply priced that books can be bought in bulk in these amounts. Buying used books and doing whatever you want with them is just legal. Used books, that would probably be just laying in some warehouse, or recycled anyway.
If there's anything to have gripes with, it's the copyright system that makes it easier to take this legal route.
All large corporations are immoral FWIW
It does not reproducibly identify itself as Claude, there's evidence to the contrary in the very thread you linked: https://x.com/bobbyNewcomb5/status/2078151562828947954
As mentioned in my comment, Kimi K3 identifies itself as Claude ~15% of the time.
Here's another report of K3 identifying itself as Claude https://x.com/Sauers_/status/2077842686459981901
And an analysis showing the self-identity distribution for K3 and other models https://x.com/RyanGreenblatt/status/2078663148509544589
Your main source is Ryan Greenblatt who is a regular recipient of community notes and has no corroboration for the 15% statistic other than his assertion. The other tweet (Sauers_) is also community noted as engagement farming with a false system prompt, so forgive me for being skeptical.
55 replies →
But then again, the identity could also have slipped into the model from other sources during pretraining. The internet is full of "I am Claude": https://grep.app/search?q=i+am+claude and variants https://grep.app/search?q=i%27m+claude
Either way, there's probably no significant portion of Mythos/Fable or Sol in there as OP has stated.
When prefaced with "I am Claude", Kimi K3 prefers to generate API-specific Anthropic model identifiers, unlike other models Qwen, GPT, or even Claude itself. These exact identifiers appear in Claude API metadata, and are stripped out of Claude web chats.
While other models produce human-readable names like "Opus 4.5" or "Sonnet 4", Kimi K3 produces exact API model identifier like "claude-opus-4-5-20251101" or "claude-sonnet-4-20250514".
Which is extremely unusual. Web chats only contain the human-readable model name. Other models don't do this. So where did K3 get this data?
We can conclude, with high confidence, that:
1) K3 was trained on raw Claude API calls/metadata.
2) Claude API metadata was trained on in additional to standard web data.
The exact model identifiers appear extremely frequently in code on GitHub.
https://grep.app/search?q=claude-opus-4-5-20251101
https://grep.app/search?q=claude-sonnet-4-20250514
They also appear elsewhere on the internet:
https://trends.google.com/trends/explore?q=claude-opus-4-5-2...
1 reply →
fwiw, Gemini 3.5 has identified itself to me as an OpenAI product on multiple occasions.
Early Grok would also identify as ChatGPT. This has happened with new model releases for years now.
Surprising they didn't clean that from the data before training. It's easy to identify, a simple search->replace gets most of it, and a cheap LLM can identify the edge cases (e.g. avoiding "Claude Shannon" -> "Kimi Shannon" or something).
Claude Sonnet 4.8 reproducibly identifies itself as DeepSeek when asked in Chinese:
https://x.com/stevibe/status/2026227392076018101
I mean, people can point fingers however they want, and the fact is nobody actually "owns" the data they feed to their LLMs...