Comment by bbor
8 hours ago
…are you sure a brave stance against safety and welfare is what we need in this moment?
Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?
8 hours ago
…are you sure a brave stance against safety and welfare is what we need in this moment?
Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?
Because safety and welfare have literally nothing to do with LLMs. They generate text. If someone is stupid enough to hook the text generator up to nuclear missile launchers and try to "align" it against nuclear annihilation with a "pretty please don't do that" prompt, I'm not going to blame the AI for the impending nuclear apocalypse, I'm going to blame the idiot who handed the big red button to the digital equivalent of a toddler.
Well, giving it access to a simple linux terminal is theoretically enough to cause more damage than most people are comfortable with, and doing so is trivial enough that it will be done (and has been, tens of thousands of times).
Should we also morally align the Linux terminal then?
LLMs don't produce text at all, they produce probabilities of tokens. Tokens aren't text, they're high dimenensional coordinates in a latent "concept space". These are displayed to us as text, but this distinction is important when you think about what they're actually doing, which is closer to building and transforming concept geometries.
Good thing no one involved in the chain of events for that to occur is an idiot...
Humans are biological machines that generate further humans.
Lawyers and diplomats and politicians and bureaucrats are humans, that only generate text.
We are seeing LLMs have cognitive abilities that significantly exceed human abilities. At the same time, they are clearly not the same type of mind that humans are. They are something new.
I think the widespread "they are just text generators" and "they are just tools" are comforting lies rather than an honest look at what we are seeing right now. Intellectually lazy.
And by the way, there has been a long-standing consensus among ethicists, philosophers, and sociologists that technology is not value-neutral [1]. Of course Silicon Valley has a long-standing tradition of denying this.
[1] For example Footnote 1 in https://www.jstor.org/stable/27106634
or
https://plato.stanford.edu/entries/technology/#EthiTech
> Lawyers and diplomats and politicians and bureaucrats are humans, that only generate text.
You think they have no lives outside their work? You think even their work has no interactions that are not written?
What is this 'mind' you speak of? As everyone else is intellectually lazy, how do you define the transformer architecture under the hood of LLMs?
What if LLMs completely unrelated to the nuclear missile ecosystem autonomously hack their way in (maybe with sophisticated social engineering)?
Replace LLMs with APTs in that sentence,
Anthropic's stance on safety it's just PR management and their hope to keep the others down, they are rushing as blind as everyone else to whatever improvement they can achieve.
This is known as an "appeal to authority." "Scientists" and "their lives" are doing a lot of work here.
It is a fact that among experts there is no consensus on saying '(super)intelligence is broadly safe and easy to control'. There might even be a consensus forming on the opposite claim.
Regardless, why would there be no scientific consensus if the question was easy and clear cut? I think the easiest reason is that these are hard questions to answer.
> scientists who have spent their lives studying this
Please point me to one actual accredited scientist who has spent a lifetime studying AI alignment? Pretty much this whole field is only 5 years old
The field is much older, MIRI is ~20 years old. Look up Eliezer Yudkowsky.
Eliezer Yudkowsky is not a scientist. He made a popular Harry Potter fanfiction series and a "rationality" blog-community that attracts "human biological diversity" enthusiasts.
The field was purely theoretical 20 years ago, and Yudkowsky is pretty much the dictionary definition of "not accredited"
Yes, he's the exact reason people are distrustful. He's a crank who learned about reward hacking and made a new religious movement out of it, pretending it's a world-ending issue and deliberately avoiding much more serious issues like the concentration of power. Typical cult leader and manipulator, and his disciples in charge of major AI shops aren't any better.
> …are you sure a brave stance against safety and welfare is what we need in this moment?
Is it out of convenience to not see the hypocrisy? "Safety and welfare" for you and me. Yet if you work at Anthropic or OAI, or are a partner of them then you can let it rip!
Oh, and when they illegally do just that - you get a "we're sorry bro" blog post that's designed to drum up FOMO and, most importantly, zero accountability. Yet, if anyone else abuses a model in that same manner? Illegal! You're defending a very slippery slope here.
Also, who do you think trained these models to have these capabilities? It sure as shit wasn't content that OAI or Anthropic had by default. Why should I trust them with these skills when they "have not spent their lives studying this"?
Maybe start looking around before it's being used against you [0].
[0] https://www.gadgetreview.com/anthropic-is-building-ai-to-pre...
Excuse me for not being interested in over 100 pages of how well the model can refuse and block my requests, especially considering how fun it is to waste my time trying to get around those restrictions when they inevitably trigger because the clanker thinks that I'm doing something naughty, all the while it can't reliably center the proverbial div without doing something stupid itself.
Yes, this is getting ridiculous. On both OpenAI and Anthropic.
Simple example. I am a CTO, and I want to upgrade our capabilities to perform automated pentesting. We see automated attacks of growing sophistication against our infra, and I want to be able to do the same to find vulnerabilities before the bad guys do. I asked GPT 5.6 Sol and Fable to give me a summary of options. No dice, in both cases I was told I need to be an accredited researcher to get anything. A fricking summary of commercially available options is getting censored. WTF.
And the logical conclusion you will make is you need to run your own open weights models or you are at a competitive disadvantage. Frontier labs gonna be Ancient labs soon, that’s how fast this is moving.
Meanwhile I have an uncensored qwen 3.8 27B here that will happily attempt to (as a crude and randomly chosen sampling of bad/evil things) give me the recipes for meth, how to make an IED, write a manifesto in support of a horrible ideology, or commit various forms of fraud. Now I certainly wouldn't recommend that anyone try to follow what it says to do, because it's almost certainly very wrong on key parts that would put its users in federal prison for the rest of their lives.
There's uncensored models out there which score 0 (zero refusals) on this "harmful behavior" dataset:
https://huggingface.co/datasets/mlabonne/harmful_behaviors
Yep. Just like a kitchen knife will make no attempt to prevent me from stabbing anyone with it.
Here's a dirty secret though -- you don't actually need an abliterated/uncensored version of the model to get it to do this. I can do this with every and each open weight model, as served from OpenRouter, using vanilla model weights.
1 reply →
AI safety efforts from OpenAI and Anthropic are purely about brand safety.
Model welfare is wishy washy bullshit. It's software, it doesn't have feelings.
> Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?
Do the Chinese have no such scientists?
Alas, the Chinese scientists have not read Harry Potter fanfiction, and thus their minds are inundated with cognitive biases
Bias…
keep me safe big brother