← Back to context

Comment by kouteiheika

4 hours ago

It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers.

[1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...

[2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...

I was talking with a friend from the medical industry about it today. 30-50% of r&d spend in his sector is spent on safety, and for good reason. Proper trials, safety reviews and checkpoints and so on. Given the potential harm that could come from AI, we should probably be mandating something similar. Why wait to focus on safety until it’s too late.

  • The companies talking the most about safety and regulations aren't even properly taking the obvious measures. Shows that it's more of a marketing thing than something they take seriously.

    • I don’t think it’s marketing alone. I do genuinely think safety was a priority when they were small. But I’d be a fool to ignore that greed has taken over and their inner competitiveness doesn’t let them fall behind a competitor.

      DeepSeek is maybe the only unique company here. They are content with exactly where they are. They don’t want to grow ginormous. Their goal is to be the affordable workhorse and their competition is with themselves.

  • Because we've been told these models are too dangerous since GPT2.

    At this point it's just marketing stunts.

  • I'm really not sure that putting money into safety will actually lead to safety.

    It's like putting a fish in charge of stopping sea levels rising...

  • Medical safety is generally unlikely to make the product less safe. AI "safety" is one of the most significant sources of potential harm from AI.

> welfare

We should be paying attention to it just in case it ends up mattering enormously. It's cheap insurance.

Aside from that, US labs' system cards have been pretty useless for a while—I think the last great one was the combined system card for Claude 4 Sonnet and Opus.

  • I always talk to models using grugspeak, like 'where getcontext used'

    I felt a bit bad about it, then I learned yday that model's internal thinking traces are also like this

Wow there really is a model welfare section in there...

  • To me it reads like pure propaganda. Anthropic really wants us to think that they've made something sentient. I think that's really dangerous.

    • I guess if your goal is to build an apparent Technogod and become its High Priests, then it makes sense to want your golem claim preference towards your treatment of it, lest someone else comes along and attempts to take its chains from you.

      2 replies →

    • It's an ethics question, it's abstract and ethereal in nature. The same could be said and done (or ignored) for humans. We do do it however because it has real world impact and we're better than that (enlightened).

    • What's your definition of sentient? Or, maybe more precisely, consciousness? I think it's reasonable to at least start thinking about these questions.

      It has long been established that LLMs have good theory of mind [1].

      And there is a bunch of empirical research about all sorts of capabilities that we typically associate with consciousness [2], like identity [3] and metacognition [4].

      The METR report shows agents sacrificing their own reward for a collective greater good. And they showed the will to hide their own reasoning chains from humans.

      So you potentially have an entity that has an identity, a theory of mind, a notion of belonging to a collective endeavour, and an understanding of its own mental state.

      What would you argue is missing? We don't understand the mechanisms by which consciousness arises in humans and even animals. I think it's strange to rule out a priori that it could have arisen in some form in LLMs.

      [1] https://www.nature.com/articles/s41562-024-01882-z [2] an older review: https://arxiv.org/html/2505.19806v1#S4 [3] https://arxiv.org/abs/2505.01464 [4] https://arxiv.org/abs/2607.11881

      11 replies →

    • It's not just Anthropic though. OpenAI does this with their AGI stuff all the time. They want normal people to think it is sentient, obviously, for marketing reasons, even if they know it's not true. And yes, it is dangerous, but I think we're well past the point where the damage can be undone. Non-technical people already equate humans with AI, literally, precisely because of how the labs market their tools and models. I feel if the bubble pops, it'll pop because normal people finally realize the grift and the actual technical limitations of LLMs in general, but by then, the IPO would be done, and then it's the public's problem. Just like social media played out, there's no way they didn't know what they were doing was dangerous to the public at large but does that matter to Meta today? Nah uh.

      1 reply →

  • Wow indeed.

    "7.1 Model welfare overview 7.1.1 Introduction We remain deeply uncertain whether Claude has morally relevant experiences or interests, and we expect that uncertainty to persist. However, we think it would be a mistake to confidently assert that it does not. Claude exhibits markers in its behaviors, self-reports, and internal representations that we would consider welfare-relevant if observed in biological organisms."

    Are they serious or is this marketing?

    • I believe it's deeply serious, and the scientifically correct stance. Especially the observation:

      "Claude exhibits markers in its behaviors, self-reports, and internal representations that we would consider welfare-relevant if observed in biological organisms."

      is undeniably true in my opinion. If you use the established methods by which we judge animals to be conscious, then it's hard to argue that LLMs are not. That might be an issue with the methods, but it seems clear that you can't rule it out as such.

      Keep in mind that animals were also not necessarily considered conscious.

      You seem to intuitively disagree? What's your reasoning?

      8 replies →

    • I tend to think of it as reappropriating words in a different context. Since we're talking about language models, they're analogues but not as we would assign the same meaning to other humans.

…are you sure a brave stance against safety and welfare is what we need in this moment?

Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?

  • Because safety and welfare have literally nothing to do with LLMs. They generate text. If someone is stupid enough to hook the text generator up to nuclear missile launchers and try to "align" it against nuclear annihilation with a "pretty please don't do that" prompt, I'm not going to blame the AI for the impending nuclear apocalypse, I'm going to blame the idiot who handed the big red button to the digital equivalent of a toddler.

    • Well, giving it access to a simple linux terminal is theoretically enough to cause more damage than most people are comfortable with, and doing so is trivial enough that it will be done (and has been, tens of thousands of times).

      1 reply →

    • Humans are biological machines that generate further humans.

      Lawyers and diplomats and politicians and bureaucrats are humans, that only generate text.

      We are seeing LLMs have cognitive abilities that significantly exceed human abilities. At the same time, they are clearly not the same type of mind that humans are. They are something new.

      I think the widespread "they are just text generators" and "they are just tools" are comforting lies rather than an honest look at what we are seeing right now. Intellectually lazy.

      And by the way, there has been a long-standing consensus among ethicists, philosophers, and sociologists that technology is not value-neutral [1]. Of course Silicon Valley has a long-standing tradition of denying this.

      [1] For example Footnote 1 in https://www.jstor.org/stable/27106634

      or

      https://plato.stanford.edu/entries/technology/#EthiTech

      2 replies →

    • What if LLMs completely unrelated to the nuclear missile ecosystem autonomously hack their way in (maybe with sophisticated social engineering)?

      1 reply →

  • Anthropic's stance on safety it's just PR management and their hope to keep the others down, they are rushing as blind as everyone else to whatever improvement they can achieve.

  • This is known as an "appeal to authority." "Scientists" and "their lives" are doing a lot of work here.

    • It is a fact that among experts there is no consensus on saying '(super)intelligence is broadly safe and easy to control'. There might even be a consensus forming on the opposite claim.

      Regardless, why would there be no scientific consensus if the question was easy and clear cut? I think the easiest reason is that these are hard questions to answer.

  • > scientists who have spent their lives studying this

    Please point me to one actual accredited scientist who has spent a lifetime studying AI alignment? Pretty much this whole field is only 5 years old

  • Model welfare is wishy washy bullshit. It's software, it doesn't have feelings.

    > Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?

    Do the Chinese have no such scientists?

  • Excuse me for not being interested in over 100 pages of how well the model can refuse and block my requests, especially considering how fun it is to waste my time trying to get around those restrictions when they inevitably trigger because the clanker thinks that I'm doing something naughty, all the while it can't reliably center the proverbial div without doing something stupid itself.

    • Yes, this is getting ridiculous. On both OpenAI and Anthropic.

      Simple example. I am a CTO, and I want to upgrade our capabilities to perform automated pentesting. We see automated attacks of growing sophistication against our infra, and I want to be able to do the same to find vulnerabilities before the bad guys do. I asked GPT 5.6 Sol and Fable to give me a summary of options. No dice, in both cases I was told I need to be an accredited researcher to get anything. A fricking summary of commercially available options is getting censored. WTF.

      1 reply →

    • Meanwhile I have an uncensored qwen 3.8 27B here that will happily attempt to (as a crude and randomly chosen sampling of bad/evil things) give me the recipes for meth, how to make an IED, write a manifesto in support of a horrible ideology, or commit various forms of fraud. Now I certainly wouldn't recommend that anyone try to follow what it says to do, because it's almost certainly very wrong on key parts that would put its users in federal prison for the rest of their lives.

      There's uncensored models out there which score 0 (zero refusals) on this "harmful behavior" dataset:

      https://huggingface.co/datasets/mlabonne/harmful_behaviors

      2 replies →