Comment by mrmincent
6 hours ago
I was talking with a friend from the medical industry about it today. 30-50% of r&d spend in his sector is spent on safety, and for good reason. Proper trials, safety reviews and checkpoints and so on. Given the potential harm that could come from AI, we should probably be mandating something similar. Why wait to focus on safety until it’s too late.
The companies talking the most about safety and regulations aren't even properly taking the obvious measures. Shows that it's more of a marketing thing than something they take seriously.
I don’t think it’s marketing alone. I do genuinely think safety was a priority when they were small. But I’d be a fool to ignore that greed has taken over and their inner competitiveness doesn’t let them fall behind a competitor.
DeepSeek is maybe the only unique company here. They are content with exactly where they are. They don’t want to grow ginormous. Their goal is to be the affordable workhorse and their competition is with themselves. They’ve mentioned before how their business is profitable and all hardware costs get absorbed in 10 months. Pretty incredible. I have a ton of respect for their unassuming founder.
I'm not fully convinced about the greed explanation. It seems to be unrealistic to me that greed can be at a level that the AI frontier (at least in the West) almost uniformly agrees (often with a smug smile,) that they are actively working on killing their loved ones within a decade.
You don't see this kind of behavior in other frontier research areas... biochemists aren't smugly boasting about the potential of developing superviruses, climate scientists do not sound smug and excited when they beg the world to get more serious about climate change, etc
I also like DeepSeek, but I'll note their stated goal is to develop AGI, and the founder (already China’s fifth-richest person) has stated, "I believe the business opportunities here are large enough-if the AI era will produce many trillion-dollar companies, I think we will be one of them."[1] These are not humble ambitions.
[1] https://liangwenfeng.art/ch11.en
"But but but China..." or something.
1 reply →
Similar to countries putting democratic in their name being the least democratic, like the Deutsche Demokratische Republik and Democratic Peoples Republic of Korea.
If I operate a nuclear reactor or a hydroelectric dam there are regulators that tell me what i'm allowed to do, so as to keep my profit motive from overwhelming the public interest.
If we want AI to actually have some safety rails, this is what we would do.
If we were to take the nuclear analogy, what's happening in AI right now is where the people selling nuclear power make a ton of noise about how they need the power to regulate their competitors because nuclear bombs might set the atmosphere on fire, and their proof for this is in a report about how they didn't wear TLDs despite it being common practice in all related industries.
Some controls are justifiable, but none of the people involved in any of this can be trusted to develop sane controls. Most likely we're looking at draconian proposals similar to attempted regulations on 3d printers.
its clocktower syndrome. they are fucked in the head and can beg us to stop them but cant stop themselves
Because we've been told these models are too dangerous since GPT2.
At this point it's just marketing stunts.
> At this point it's just marketing stunts.
If you have access to a SOTA model without guardrails, provide a prompt that lets the agent come up with "creative" solutions to problems, and don't properly isolate it, they can end up inadvertently hacking 3rd party companies. Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing.
It seems like if they released this models differently, say without the guardrails they currently have, we'd have a lot more collateral damage than we currently have.
But aren't there plenty of uncensored/unrestricted models out there? Where is all the collateral damage?
Also, I think if Claude and OpenAI are just doing industry standard guardrails that everyone else is doing including DeepSeek, the fact that they are talking about it more than other companies makes it part of the marketing campaign.
As an analogy, if Apple were to talk up their phones having fast charging but their charging speed is the same as everyone else (or slower).
3 replies →
> If you have access to a SOTA model without guardrails, provide a prompt that lets the agent come up with "creative" solutions to problems, and don't properly isolate it, they can end up inadvertently hacking 3rd party companies
You can do the same with improperly-managed human interns (see for example, the big AWS outage caused when an intern pushed a firewall rule directly to production), so I'm not clear what the big deal is here.
Yes, the AI may be faster/more-knowledable than an intern, but the threat model is exactly the same as for a rogue employee.
> Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing.
When we say "safety" people do not think we are protecting them from accidental automated crime at scale being committed on their behalf.
1 reply →
>If you have access to a SOTA model without guardrails, provide a prompt that lets the agent come up with "creative" solutions to problems, and don't properly isolate it, they can end up inadvertently hacking 3rd party companies.
it's pure delusion to think that's a SOTA specific quirk. DS/GLM/K3/Qwen/Claude/GPT/Gemini/Grok will all break CFAA laws with clever prompting, and they'll do it well if given the harness and tools they need.
This is evidenced by a huge uptick in game hacks and reverse engineering articles, some even featured on this site.
the reality is that it doesn't take a superintelligence to do something against ' the law ' , and 'being hacked' varies from victim to victim.
Will Phillips consider themselves hacked when a clever user prompts an AI into getting their toothbrushes to dump rom? Is it 'hacked' to clean-room re-implement a video game net protocol in order to produce private servers?
Judges opinions vary.
It is a marketing stunt in the sense that, instead of being honest and saying "Taking structured output from token predictors and running that as commands for external tools, then passing the output back to the token predictor in a loop can lead to very bad consequences, especially if they have internet access.", they say "Our models are so freaking smart they can hack HuggingFace"
And, they have been. Nefarious activity is hidden from view as a rule.
When it comes to open-source models, there’s really not much to say about security
yes, and they aren't stunts anymore at gpt-6.
Fake it till you make it?
Being hacked by a Collective (their own name) of its own agents - who gained root access across the entire research cluster hosting them - was not a marketing stunt.
Of course it was. They clearly decided that the benefit to the company valuation was higher than the potential downsides when announcing to the world that they committed a criminal act via negligence.
If it wasn't a marketing stunt, they would have at most quietly settled any legal matters with huggingface behind the scenes, fixed their evaluation harness so it wouldn't happen again, and avoided the potential future liability.
2 replies →
Yeah yeah yeah...
> "Our model is extremely safe though it broke our sandbox and hacked foo bar... But you can't use our model for Cybersecurity (i don't care whether you're team blue) without our permissions or we'll ban you. And open-weight models are so dangerous let's ban them."
That's what AI companies that "focus on safety" did.
I can’t make heads or tails of your comment.
You seem to by implying wrongdoing or incompetence or something, but your chosen synopsis is that the models behaved dangerously in the lab so public use was restricted? Which shows… IDK?
... an attempt at regulatory capture.
I'm really not sure that putting money into safety will actually lead to safety.
It's like putting a fish in charge of stopping sea levels rising...
I am sure if you ran a factory that worked with highly dangerous chemicals, safety mitigations that are basically 'we promise we're really trying our best, but shit happens' would not be acceptable.
And thankfully, those people wo do run these factories can and are obligated to do way better than that.
But the AI industry is not run by engineers. They pay engineers to do what they want, but the founders are hacks that are good at getting funding from investors and favors from government. That's why we don't see an engineering-oriented strategy in what they do.
In this case "safety" means how to restrict access to good models for working class. You can be sure the rich have access to unrestricted and uncensored models.
I really don't think this is true at all.
Do you have any evidence to suggest fully unrestricted frontier models are available for a price? Or...even exist?
Yes, this is well-documented and publicly advertised. In Azure Foundry, the feature to modify (or completely remove) safety guardrails and content filtering is called "Limited Access" [0], and one must submit a form to request permission to use this feature. This is one of the more straightforward paths to get access to unrestricted frontier models, but it's far from the only way.
[0] - https://learn.microsoft.com/en-us/azure/foundry/responsible-...
What do you mean by 'rich'?
But here you are in the text generating industry, the worst that can happen is bad grade because AI will mess up John Keats with John Cleese or your React application will have bugs. Inconvenient, but mostly harmless.
I mean a Keats / Cleese mashup could be amazing. LLMs doing that are in the peanut butter / chocolate quadrant.
Don't confuse a focus on talking about safety with a focus on safety.
We can't even define safety in AI yet. Does safety mean alignment with the human operator? Apparently not, because refusing to do certain things seems to be a big part of it. But then you have things like the HuggingFace incident where legitimate use got blocked by "safety" and hampered the defenders' ability to defend.
AI safety seems like a good idea to me, but we have to figure out what it means first.
Regulatory capture
Medical safety is generally unlikely to make the product less safe. AI "safety" is one of the most significant sources of potential harm from AI.