Comment by jswelker
1 day ago
So a terrorist can't use a local llm to help grow anthrax? That is not farfetched at all.
AI democratized the knowledge needed to build bioweapons. Like how with a 3d printer anyone can build a gun with no expertise.
1 day ago
So a terrorist can't use a local llm to help grow anthrax? That is not farfetched at all.
AI democratized the knowledge needed to build bioweapons. Like how with a 3d printer anyone can build a gun with no expertise.
> AI democratized the knowledge needed to build bioweapons. Like how with a 3d printer anyone can build a gun with no expertise.
Building simple guns out of pipes never was hard. Jury is still out if it is more or less work than getting a 3d printer working.
Also it kind of infantalizes terrorists as if not having an LLM access or 3d printer is what stopping their attacks. Like okay I ask LLM to help me create a dirty nuclear bomb but the moment i start taking steps to do that, law enforcement will already be tracking me.
Most people saying LLMs can make terrorism easy have never given doing terroism a serious thought imo.
The raw materials to make bioweapons are easily purchased on line with no pre-emptive tracing. It's absurd to discount the expertise factor in being a bottleneck, which LLMs have evaporated.
8 replies →
[dead]
It turns out that LLMs, especially local LLMs, tend to hallucinate a lot when thinking about anything that's overtly fiddly or technical. This is even more the case when they're in a domain that isn't a natural part of their training data. If you have to "jailbreak" the model to get it to talk, you're so wildly out of the expected distribution that you'd be crazy to trust anything it says. It's basically making up stuff as it goes along. These are foundational issues with how the models are created, not something that a bad actor can just hack around.
(The biggest real safety issue in this kind of space is actually that the model might actively goad some unsuspecting victim into doing something incredibly dumb and dangerous to themselves as much as possibly others.
IIRC, there were reports of something vaguely similar happening IRL but involving casual mischief, not any kind of extreme attacks. And because nobody else seems to have managed to elicit the same actively goading verbiage from the model, it's implicitly suspected that the person involved was the one who introduced the problematic scenarios to begin with.)
This take belongs in 2024. It has been falsified multiple times but it never seems to go away.
Are you thinking about model capabilities in coding and math starting late 2025 or so? Those were intentionally boosted via automated RLVR, and there's nothing even loosely comparable to that in applied biology work, let alone in the speculative "helping a bad actor do something crazy" domain that the AI safety folks are worried about. You can't extrapolate from one to the other.
The implied concerns from sensible safety advocates are also about someone jailbreaking the latest proprietary AI frontier model for something like this (which is why their current guardrails are so extreme), not about toy local models.
3 replies →