Comment by petesergeant
9 hours ago
I wish nothing but luck for an EU model, but:
> intellectual-property safety
My suspicion is that you simply can't build an even slightly competitive model without liberally stealing your training data, in 2026, as much as I'd like it to be otherwise. You can get to the point that I suspect most of the frontier labs are at, where you've laundered the initially stolen data through the creation of huge amounts of derivative synthetic data, but still. Anyone who isn't comfortable stealing their training data is bringing a knife to a gun fight, and is going to die a noble but inevitable death.
This doesn't seem to be true. There's a clear legal path via the first-sale doctrine to train models on copyrighted works. It's been years now, and publishers still don't seem to be offering anything for training (e.g. bulk licenses solely for training use), but adversarial interoperability via cutting up books and scanning them remains perfectly legal.
There's also the ability to distill other models, which is also not illegal (though I'm sure they like to come after whomever for TOS violations, but thats a civil matter).
And, of course, the obligatory copying-isn't-theft observation. A recent supreme court judgment put it well.
> Since the statutorily defined property rights of a copyright holder have a character distinct from the possessory interest of the owner of simple “goods, wares, [or] merchandise,” interference with copyright does not easily equate with theft, conversion, or fraud. The infringer of a copyright does not assume physical control over the copyright, nor wholly deprive its owner of its use. Infringement implicates a more complex set of property interests than does run-of-the-mill theft, conversion, or fraud.
Folks are pretty smart here, I think we can handle these nuances, even if we don't agree about whether they are good.
Edit: reading through the full text of their post, it looks like they are using common crawl, which is likely just as much of a copyright infringement as Anna's Archive -- it's not like published works have a unique claim to copyright. I think this strengthens your point, though: I was expecting to see scans as training data, but it doesn't appear to be the case.
"The Congress shall have Power To ... promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries." - The United States Constitution
Copyright is a government mandated monopoly that was only granted in order to advance the arts and science. Any interpretation that runs contrary to that is bollocks being used by the religiously or financially motivated to serve their own petty interests to the detriment of societies.
Disclaimer: I am part of the team that trained Kolibri, opinions are mine.
You're right that especially big models benefit from training on copyrighted material in terms of world knowledge (especially from books). However, in the small model space imho agentic capabilities where the model looks up knowledge on the fly are much more important. That's what we focused on quite a bit during training. Personally, I also don't think stealing stuff is okay.
I hope you're right, but I guess we'll see when we get independent benchmarks. My intuition is that even for small models and models mostly focused on tool calling, you _still_ need all that extra contextual stuff for the magic, but I am further from the coalface than you are.
Why should we concede this "stealing" framing? If I want to put texts into my computer program, why should that count as copyright violation? I think that idea is as ridiculous as saying that reading a document is a copyright violation.
Regulate large cloud services and proprietary software - yes! But not on the basis of "Intellectual Property".
The "legal" issues here are very very complex and we should not passively wait for or accept corrupt court rulings, international trade agreements, proposed laws, or worst of all propaganda that pushes a parochial and craven view on this.
This model isn't terrible, at least on the benchmarks. It's 78B A3B and performs about like Qwen3.6 35B A3B. You can probably run it comfortably in 96B of RAM with a decent quant that doesn't lose too much.
Unfortunately, Qwen3.6 35B A3B isn't really a useful coding model. You'd probably want Qwen3.8 27B at a minimum, which requires at least 32GB of VRAM (not system RAM) to run semi-comfortably.
So this isn't going to be a competitive model for hobbyists, and you'd have to be a bit desperate to use it for coding. But if you work in a regulated industry and don't mind paying for a bit of extra hardware, it isn't catastrophically bad, either. Probably would work fine for information extraction or as a "classifier" like Jev. (Almost any GGUF model can be turned into a classifier using llama-server. See pi.dev codemode for sample code.)
So they're not a real contender yet, but they look like they're probably at least minimally credible.
hasn't IP law passed the statute of limitations? As in most models are probably trained on output of other models, as creating enough data otherwise is not feasible. Additionally, they are trained on github repos made since the AI boom, which were generated by models with IP issues (who knows what and how).
Thus training on 'clean' data is like trying to unscramble an egg.
There is no word for copyright in Mandarin :)
Is this true? Do they possibly use a loanword or a descriptive term? Certainly you are not implying that the concept of copyright does not actually exist in Chinese society?
For what it's worth, in my language we don't have a word for copyright either. We have the concept, though, we just call it literally Creators Rights זכויות יוצרים and the borders of what is and what isn't covered broadly map to the familiar concepts of IP.
版权
I mean, there is one, they have copyright law. Forgive me for being slow is this a joke about the widespread theft of IP in China? Or was the acquisition of training data just much more 'accepted' in China compared to the west?
I feel I messed up your quip =/ I'm new here, go ez. Not looking for excuses to hate on China either.
Yes, the law exists, and anyone can go and ask their IP back in the court of Beijing.
Not really. The upside of competitive newer models is all in the proprietary data they are trained on. This is why data labeling, RL environments et al have been such a big industry, OpenAI and Antrhopic are paying literally billions to get the data they need. Do people think the ability to do research level math or advanced cybersec comes from just training on more public data?
A real sovereign effort could invest heavily in this, whatever people accuse China of “stealing” I’m sure they are also generating tons of their own data and are probably the primary sovereign doing so outside the US labs.
Have you tried the model itself and seen if it's "even slightly competitive" or not, and have specific complaints about it? Otherwise it feels like you're complaining about something that is easy to test but rather than taking the time to actually figuring that out first, you're arguing about some general and theoretical thing which the submission (may) directly disprove.
No, I haven’t, but I’ll donate $20 to the non-political charity of your choice if it doesn’t turn out to sit a significant difference from the frontier.
I think it’s a safe assumption that they’re leaning into “sovereign” because performance is bad.
I think you have it backwards. Sovereign is the goal, good can come later.
There is a proliferation of sovereign models under development specifically to address data sovereignty, and a loss of performance is absolutely acceptable over the risk that a once ally will turn adversarial, or a foreign business stops serving what has become critical infrastructure.
Is it noble? The entire notion that training data can be "stolen" at all is quite silly. If I "steal" content that someone created to use for training, what am I actually stealing? They didn't lose anything. They still have everything they had before. What was "stolen" was "unrealized profit", or put another way: money that wasn't theirs, that they had no entitlement to. The only actual crime that is committed is "unauthorized copying", not stealing. Support and enforcement of copyright feels wildly authoritarian. It's hard to see it as noble.
The same could be said of any digital product being sold. Nobody actually loses anything but the actual sale either when you download a cracked game or piece of software, a movie, music, etc.
That's fair, but then people like you complain when someone "steals" I mean distills openai or anthropic models.
It’s poor form to argue against someone by imagining something totally different that they might believe, which would then make them hypocritical.
I don't complain about that. Model distillation is excellent.