Comment by embedding-shape
3 hours ago
> At this point it's just marketing stunts.
If you have access to a SOTA model without guardrails, provide a prompt that lets the agent come up with "creative" solutions to problems, and don't properly isolate it, they can end up inadvertently hacking 3rd party companies. Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing.
It seems like if they released this models differently, say without the guardrails they currently have, we'd have a lot more collateral damage than we currently have.
But aren't there plenty of uncensored/unrestricted models out there? Where is all the collateral damage?
Also, I think if Claude and OpenAI are just doing industry standard guardrails that everyone else is doing including DeepSeek, the fact that they are talking about it more than other companies makes it part of the marketing campaign.
As an analogy, if Apple were to talk up their phones having fast charging but their charging speed is the same as everyone else (or slower).
An obliterated 30B model versus a 1T model without guardrails is like comparing an angry squirrel to a bear having a bad day. One hurts, the other hurts until it abruptly doesn't.
> But aren't there plenty of uncensored/unrestricted models out there? Where is all the collateral damage?
That come close to what SOTA GPT models are able to do? No, not even close. They're either "safety trained" and has bunch of guardrails, or aren't able to come up with 0days on the spot to escalate to root access on 3rd party infrastructure.
> doing industry standard guardrails that everyone else is doing including DeepSeek, the fact that they are talking about it more than other companies makes it part of the marketing campaign.
Yeah, that sounds reasonable to me, since all the top models currently have guardrails one way or another, but the amount they mention it in the press releases differs a lot.
> Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing.
When we say "safety" people do not think we are protecting them from accidental automated crime at scale being committed on their behalf.
I'm fairly sure most "safety" people consider "large scale automated crime" part of the threat model, as the agents could accidentally fall into such a trap, if optimized for some misunderstood goal.
It is a marketing stunt in the sense that, instead of being honest and saying "Taking structured output from token predictors and running that as commands for external tools, then passing the output back to the token predictor in a loop can lead to very bad consequences, especially if they have internet access.", they say "Our models are so freaking smart they can hack HuggingFace"