It is not feasible. They never made much of an inroad against torrents and that is a much easier target than abliterated models. As the linked website shows; the process to abliterate a model can be as simple as
Torrents are legal until someone, at great expense and difficulty, proves otherwise (even then, jurisdiction and content dependent). At which point everyone involved will ignore the fact and carry on. That is a situation with enormous will, lots of money and ongoing enforcement effort to suppress the things.
And compared to torrents abliterated models are more complicated to identify, harder to suppress and there is a lot less reason for anyone to care.
I think that was the point being made? Outlawing something does nothing if enforcement is not feasible. The music and movie industries didn't crush torrents, they switched business models to streaming with prices being determined mostly by how much hassle was avoided by skipping the torrents.
Maybe it would be easiest for everyone if you clarified what country you live in, because abliterated LLM torrents are not "outlawed" in any of the major Internet-using countries that I am aware of.
Just because something is easily available does not mean that it isn't easily banned.
That doesn't make it go away, but it gives a dystopian government a lot of excuses to go after people breaking the law.
At the moment is there any legislature that is seriously pressing regulation to as you say outlaw .. or ban outright AI models that do not have guard rails built in? i know there is a lot of moves about this for things used in critical infrastructure.. but i thought no one is really saying ban these things completely.
If you think closed source software/binaries only is bad, wait until you see how awful the state of the art is with a clear-as-mud bucket of matrix weights.
We know it's possible to train an LLM to secretly respond to certain trigger phrases, and last I checked these could only be detected with the assistance of whoever chose those phrases.
The trigger condition for such backdoors is not something anyone can do a systematic brute-force check for, for the same reason we had to invent LLMs in order to do natural language processing: combinatorial explosion.
Passing around open weight models from known sources is already asking you to trust those sources; because of how difficult this is to do correctly even without deliberately inserting such things, we still don't know if China has already put such trigger conditions into their models despite headlines such as these: https://venturebeat.com/security/deepseek-injects-50-more-se...
Regardless of if it was deliberate or not, we don't know if we caught all of these misbehaviours. We don't know how to.
And note, I'm not saying "and therefore you should trust the Big Name Models". If open weight models score 2/100 in this context, closed ones score 1/100.
You can actually discover those in open weight artifacts, reproduce them, study them and issue a security bulletin.
With proprietary hosted weights you can be specifically targeted and you would not be able to reproduce nor prove anything.
Poisoning open models would be of short-term benefit to China only if they could target US (and maybe EU + Commonwealth) specifically. Damaging anyone else would be a net loss and would erode the partnerships and alliances they are trying to build elsewhere. So it's a fire-once weapon with a huge risk of collateral damage.
Much more plausible is simply making the models ideologically biased, but as history teaches us, preferring ideology or religion over science is a well-known path to ruin. It would be weird to simultaneously warn public not to use their own open models, so.
I think the most plausible explanation for open models is simply that Huawei wants more customers and is willing to compete on the hardware front.
> You can actually discover those in open weight artifacts, reproduce them, study them and issue a security bulletin.
No, you actually cannot. Not in general and without already knowing what the whole trigger pattern is. It's absolutely possible to put in a trigger that only fires while working on backend code on a specific date in a specific company by a specific github username, and no way to find this except by trying that combination, thanks to the terrible state of current mechanistic interpretability tools.
Remember: an AI model is not code. Solving this problem is as hard as the entire alignment problem.
The companies at the bleeding edge of research into this topic do not know how to reliably perform the kind of thing you suggest here.
The only reason we can point at DeepSeek-R1 and say the following, is because we can guess the magic keywords:
we found that when DeepSeek-R1 receives prompts containing topics the Chinese Communist Party (CCP) likely considers politically sensitive, the likelihood of it producing code with severe security vulnerabilities increases by up to 50%.
> Poisoning open models would be of short-term benefit to China only if they could target US (and maybe EU + Commonwealth) specifically. Damaging anyone else would be a net loss and would erode the partnerships and alliances they are trying to build elsewhere. So it's a fire-once weapon with a huge risk of collateral damage.
This "fire-once weapon" has already been fired, and appears to be a massive foot-gun for every model on a near-continuous basis.
Nobody would use LLMs if the trust deficit alone was a sufficient argument.
> Much more plausible is simply making the models ideologically biased, but as history teaches us, preferring ideology or religion over science is a well-known path to ruin. It would be weird to simultaneously warn public not to use their own open models, so.
"Ideologically biased" is the alternative explanation for the already-observed output of DeepSeek-R1. We can't tell which explanation, malicious or accidental bias, is the actual cause.
Agree. I took a look at these last few months, did a write-up: https://languageops.com/blog/ai-safety-pdoom-local-vs-fronti... and I don't know if I agree or not on outlawing completely, but I think an age restriction *at least* like for alcohol, firearms and driving would be not unwise.
For now. One more ternary model type breakthrough, or MoE, engram thing (I don't fully understand those for the record, I just know they speed things up and use less VRAM) could see a few GB sized weights with quite the capabilities. On a gaming PC, savvy teens can already use them to cook up quite an interesting array of likely illegal items and substances. If it goes much further and runs on phones, you can assume word will get around that unlimited private AI is available and kids will run into all sorts of issues. Or mentally unwell people. I'm thinking a year or two down the line only.
It is not feasible. They never made much of an inroad against torrents and that is a much easier target than abliterated models. As the linked website shows; the process to abliterate a model can be as simple as
pip install -U heretic-llm && heretic Qwen/Qwen3.5-4B
let alone people just putting the weights up in a torrent. All assuming that someone even tried to ban abliterated models.
The torrents you are talking about are outlawed. Whether enforcement is working or not is another issue.
Torrents are legal until someone, at great expense and difficulty, proves otherwise (even then, jurisdiction and content dependent). At which point everyone involved will ignore the fact and carry on. That is a situation with enormous will, lots of money and ongoing enforcement effort to suppress the things.
And compared to torrents abliterated models are more complicated to identify, harder to suppress and there is a lot less reason for anyone to care.
I think that was the point being made? Outlawing something does nothing if enforcement is not feasible. The music and movie industries didn't crush torrents, they switched business models to streaming with prices being determined mostly by how much hassle was avoided by skipping the torrents.
5 replies →
Maybe it would be easiest for everyone if you clarified what country you live in, because abliterated LLM torrents are not "outlawed" in any of the major Internet-using countries that I am aware of.
Outlawed? Only in safetyist dreams
Just because something is easily available does not mean that it isn't easily banned. That doesn't make it go away, but it gives a dystopian government a lot of excuses to go after people breaking the law.
At the moment is there any legislature that is seriously pressing regulation to as you say outlaw .. or ban outright AI models that do not have guard rails built in? i know there is a lot of moves about this for things used in critical infrastructure.. but i thought no one is really saying ban these things completely.
They picked a good name for fighting that. The optics of trying to outlaw heresy probably aren't great. ;)
I feel like half the people in the US would zealously defend and support the government if it wanted to outlaw heresy.
This is the test. If the speech that's easiest to dislike is legal, then we all have free speech.
IMO math is free speech, and outlawing math is censorship.
Like they've outlawed drugs? Illegal weapons? Hacking?
I'm not sure what is your point. It reads as defeatism to me but I'm not sure.
Could you elaborate? Do you find it good or bad? What actions can be taken?
I'm not sure yet, tbh. Perhaps it does make sense to outlaw them eventually.
Then again, it will probably not stop someone who is determined. Same as with other legislation really.
Hes of the mind that american fascism will hold together long enough to be competent decesion makers
Good.
If you think closed source software/binaries only is bad, wait until you see how awful the state of the art is with a clear-as-mud bucket of matrix weights.
We know it's possible to train an LLM to secretly respond to certain trigger phrases, and last I checked these could only be detected with the assistance of whoever chose those phrases.
The trigger condition for such backdoors is not something anyone can do a systematic brute-force check for, for the same reason we had to invent LLMs in order to do natural language processing: combinatorial explosion.
Passing around open weight models from known sources is already asking you to trust those sources; because of how difficult this is to do correctly even without deliberately inserting such things, we still don't know if China has already put such trigger conditions into their models despite headlines such as these: https://venturebeat.com/security/deepseek-injects-50-more-se...
Regardless of if it was deliberate or not, we don't know if we caught all of these misbehaviours. We don't know how to.
And note, I'm not saying "and therefore you should trust the Big Name Models". If open weight models score 2/100 in this context, closed ones score 1/100.
You can actually discover those in open weight artifacts, reproduce them, study them and issue a security bulletin.
With proprietary hosted weights you can be specifically targeted and you would not be able to reproduce nor prove anything.
Poisoning open models would be of short-term benefit to China only if they could target US (and maybe EU + Commonwealth) specifically. Damaging anyone else would be a net loss and would erode the partnerships and alliances they are trying to build elsewhere. So it's a fire-once weapon with a huge risk of collateral damage.
Much more plausible is simply making the models ideologically biased, but as history teaches us, preferring ideology or religion over science is a well-known path to ruin. It would be weird to simultaneously warn public not to use their own open models, so.
I think the most plausible explanation for open models is simply that Huawei wants more customers and is willing to compete on the hardware front.
> You can actually discover those in open weight artifacts, reproduce them, study them and issue a security bulletin.
Finding unknown backdoors in models is NP hard.
> You can actually discover those in open weight artifacts, reproduce them, study them and issue a security bulletin.
No, you actually cannot. Not in general and without already knowing what the whole trigger pattern is. It's absolutely possible to put in a trigger that only fires while working on backend code on a specific date in a specific company by a specific github username, and no way to find this except by trying that combination, thanks to the terrible state of current mechanistic interpretability tools.
Remember: an AI model is not code. Solving this problem is as hard as the entire alignment problem.
The companies at the bleeding edge of research into this topic do not know how to reliably perform the kind of thing you suggest here.
The only reason we can point at DeepSeek-R1 and say the following, is because we can guess the magic keywords:
- https://www.crowdstrike.com/en-us/blog/crowdstrike-researche...
> Poisoning open models would be of short-term benefit to China only if they could target US (and maybe EU + Commonwealth) specifically. Damaging anyone else would be a net loss and would erode the partnerships and alliances they are trying to build elsewhere. So it's a fire-once weapon with a huge risk of collateral damage.
This "fire-once weapon" has already been fired, and appears to be a massive foot-gun for every model on a near-continuous basis.
Nobody would use LLMs if the trust deficit alone was a sufficient argument.
> Much more plausible is simply making the models ideologically biased, but as history teaches us, preferring ideology or religion over science is a well-known path to ruin. It would be weird to simultaneously warn public not to use their own open models, so.
"Ideologically biased" is the alternative explanation for the already-observed output of DeepSeek-R1. We can't tell which explanation, malicious or accidental bias, is the actual cause.
Agree. I took a look at these last few months, did a write-up: https://languageops.com/blog/ai-safety-pdoom-local-vs-fronti... and I don't know if I agree or not on outlawing completely, but I think an age restriction *at least* like for alcohol, firearms and driving would be not unwise.
Make sure they ban books with dangerous knowledge too.
Books with dangerous knowledge are banned.
See, the various banned porn varieties for an easy example
2 replies →
The hardware requirements are already quite restrictive
Hardware restrictions aren't restrictive for criminal organizations.
For now. One more ternary model type breakthrough, or MoE, engram thing (I don't fully understand those for the record, I just know they speed things up and use less VRAM) could see a few GB sized weights with quite the capabilities. On a gaming PC, savvy teens can already use them to cook up quite an interesting array of likely illegal items and substances. If it goes much further and runs on phones, you can assume word will get around that unlimited private AI is available and kids will run into all sorts of issues. Or mentally unwell people. I'm thinking a year or two down the line only.
Oh no, some run on iPhones
3 replies →