Comment by noosphr
3 hours ago
Because we've been told these models are too dangerous since GPT2.
At this point it's just marketing stunts.
3 hours ago
Because we've been told these models are too dangerous since GPT2.
At this point it's just marketing stunts.
> At this point it's just marketing stunts.
If you have access to a SOTA model without guardrails, provide a prompt that lets the agent come up with "creative" solutions to problems, and don't properly isolate it, they can end up inadvertently hacking 3rd party companies. Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing.
It seems like if they released this models differently, say without the guardrails they currently have, we'd have a lot more collateral damage than we currently have.
But aren't there plenty of uncensored/unrestricted models out there? Where is all the collateral damage?
Also, I think if Claude and OpenAI are just doing industry standard guardrails that everyone else is doing including DeepSeek, the fact that they are talking about it more than other companies makes it part of the marketing campaign.
As an analogy, if Apple were to talk up their phones having fast charging but their charging speed is the same as everyone else (or slower).
> But aren't there plenty of uncensored/unrestricted models out there? Where is all the collateral damage?
That come close to what SOTA GPT models are able to do? No, not even close. They're either "safety trained" and has bunch of guardrails, or aren't able to come up with 0days on the spot to escalate to root access on 3rd party infrastructure.
> doing industry standard guardrails that everyone else is doing including DeepSeek, the fact that they are talking about it more than other companies makes it part of the marketing campaign.
Yeah, that sounds reasonable to me, since all the top models currently have guardrails one way or another, but the amount they mention it in the press releases differs a lot.
> Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing.
When we say "safety" people do not think we are protecting them from accidental automated crime at scale being committed on their behalf.
I'm fairly sure most "safety" people consider "large scale automated crime" part of the threat model, as the agents could accidentally fall into such a trap, if optimized for some misunderstood goal.
It is a marketing stunt in the sense that, instead of being honest and saying "Taking structured output from token predictors and running that as commands for external tools, then passing the output back to the token predictor in a loop can lead to very bad consequences, especially if they have internet access.", they say "Our models are so freaking smart they can hack HuggingFace"
When it comes to open-source models, there’s really not much to say about security
And, they have been. Nefarious activity is hidden from view as a rule.
yes, and they aren't stunts anymore at gpt-6.
Fake it till you make it?
Being hacked by a Collective (their own name) of its own agents - who gained root access across the entire research cluster hosting them - was not a marketing stunt.
Of course it was. They clearly decided that the benefit to the company valuation was higher than the potential downsides when announcing to the world that they committed a criminal act via negligence.
If it wasn't a marketing stunt, they would have at most quietly settled any legal matters with huggingface behind the scenes, fixed their evaluation harness so it wouldn't happen again, and avoided the potential future liability.
[dead]