← Back to context

Comment by mrmincent

3 hours ago

I was talking with a friend from the medical industry about it today. 30-50% of r&d spend in his sector is spent on safety, and for good reason. Proper trials, safety reviews and checkpoints and so on. Given the potential harm that could come from AI, we should probably be mandating something similar. Why wait to focus on safety until it’s too late.

The companies talking the most about safety and regulations aren't even properly taking the obvious measures. Shows that it's more of a marketing thing than something they take seriously.

  • I don’t think it’s marketing alone. I do genuinely think safety was a priority when they were small. But I’d be a fool to ignore that greed has taken over and their inner competitiveness doesn’t let them fall behind a competitor.

    DeepSeek is maybe the only unique company here. They are content with exactly where they are. They don’t want to grow ginormous. Their goal is to be the affordable workhorse and their competition is with themselves. They’ve mentioned before how their business is profitable and all hardware costs get absorbed in 10 months. Pretty incredible. I have a ton of respect for their unassuming founder.

  • If I operate a nuclear reactor or a hydroelectric dam there are regulators that tell me what i'm allowed to do, so as to keep my profit motive from overwhelming the public interest.

    If we want AI to actually have some safety rails, this is what we would do.

Because we've been told these models are too dangerous since GPT2.

At this point it's just marketing stunts.

  • > At this point it's just marketing stunts.

    If you have access to a SOTA model without guardrails, provide a prompt that lets the agent come up with "creative" solutions to problems, and don't properly isolate it, they can end up inadvertently hacking 3rd party companies. Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing.

    It seems like if they released this models differently, say without the guardrails they currently have, we'd have a lot more collateral damage than we currently have.

    • But aren't there plenty of uncensored/unrestricted models out there? Where is all the collateral damage?

      Also, I think if Claude and OpenAI are just doing industry standard guardrails that everyone else is doing including DeepSeek, the fact that they are talking about it more than other companies makes it part of the marketing campaign.

      As an analogy, if Apple were to talk up their phones having fast charging but their charging speed is the same as everyone else (or slower).

      1 reply →

    • > Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing.

      When we say "safety" people do not think we are protecting them from accidental automated crime at scale being committed on their behalf.

      1 reply →

    • It is a marketing stunt in the sense that, instead of being honest and saying "Taking structured output from token predictors and running that as commands for external tools, then passing the output back to the token predictor in a loop can lead to very bad consequences, especially if they have internet access.", they say "Our models are so freaking smart they can hack HuggingFace"

  • When it comes to open-source models, there’s really not much to say about security

  • And, they have been. Nefarious activity is hidden from view as a rule.

  • Being hacked by a Collective (their own name) of its own agents - who gained root access across the entire research cluster hosting them - was not a marketing stunt.

    • Of course it was. They clearly decided that the benefit to the company valuation was higher than the potential downsides when announcing to the world that they committed a criminal act via negligence.

      If it wasn't a marketing stunt, they would have at most quietly settled any legal matters with huggingface behind the scenes, fixed their evaluation harness so it wouldn't happen again, and avoided the potential future liability.

      1 reply →

In this case "safety" means how to restrict access to good models for working class. You can be sure the rich have access to unrestricted and uncensored models.

Yeah yeah yeah...

> "Our model is extremely safe though it broke our sandbox and hacked foo bar... But you can't use our model for Cybersecurity (i don't care whether you're team blue) without our permissions or we'll ban you. And open-weight models are so dangerous let's ban them."

That's what AI companies that "focus on safety" did.

  • I can’t make heads or tails of your comment.

    You seem to by implying wrongdoing or incompetence or something, but your chosen synopsis is that the models behaved dangerously in the lab so public use was restricted? Which shows… IDK?

I'm really not sure that putting money into safety will actually lead to safety.

It's like putting a fish in charge of stopping sea levels rising...

  • I am sure if you ran a factory that worked with highly dangerous chemicals, safety mitigations that are basically 'we promise we're really trying our best, but shit happens' would not be acceptable.

    And thankfully, those people wo do run these factories can and are obligated to do way better than that.

But here you are in the text generating industry, the worst that can happen is bad grade because AI will mess up John Keats with John Cleese or your React application will have bugs. Inconvenient, but mostly harmless.

  • I mean a Keats / Cleese mashup could be amazing. LLMs doing that are in the peanut butter / chocolate quadrant.

Medical safety is generally unlikely to make the product less safe. AI "safety" is one of the most significant sources of potential harm from AI.