← Back to context

Comment by modeless

1 month ago

> All sufficiently capable models, open and closed, should go through mandatory safety testing

What happens if a model fails the test? Surely one can use Kimi K3 for evil, somehow or other. What now?

"Mandatory safety testing" implies consequences for failing, yet Dario has nothing to say about what the consequences should be. He says he doesn't advocate a ban but it's hard to imagine what his alternative would be if he won't say it.

Nah, the statement is the mechanism for a ban. The proctor will be someone anthropic trusts and "surprise" as it turns out all the open weight models fail or aren't eligible.

What happens if a model passes the government tests and then later someone fine tunes it to behave differently, without making their changes public?

  • Any reasonable safety testing should include finetuning and safety margin to account for others may do better finetuning.

    • I can fine tune significant behavior changes, there is little model developers can do to prevent this (aiui), so this effectively becomes an blanket ban

      2 replies →

"mandatory safety testing" is an impractical ideal, it's not workable in real world. Like any technology, LLMs are dual-use tools capable of both beneficial and malicious applications—a fundamental reality that human intent cannot change.

If a model fails the test, it should be banned. He is not advocating a ban of open-weight models. He is advocating a ban of models that fail mandatory safety testing. Seems reasonable and straightforward.

  • Can you think of anything open-source that has to go through mandatory tests to be distributed and still survived ?

    There is a reason to it, that's as good as any angle to find why IMHO.

    • I think Gemma will be fine. Most open-weight models are not capable enough to be dangerous. Yes, I can't think of any capable open-weight model that would survive reasonable safety testing.

  • Any sufficiently capable open weights model would fail "safety" testing though, as any "safeguards" of the sort Anthropic likes could be removed. That's just another way of saying they want a ban on capable open source models which would contradict their earlier statement, or at least make it very misleading. It's hard to see how this post can be internally consistent without some hint from Dario about what he believes should happen to models that fail safety testing and/or how capable open weights models could possibly pass a safety test of the kind he proposes.

    • I agree we don't know how capable open-weight models could possibly pass any reasonable safety testing NOW, but that's about currently abysmal state of AI alignment research, not about what is possible in principle. I don't see any internal inconsistency, to be honest. Since Anthropic does not release any capable open-weight models, it's not their problem. If mandatory safety testing is established, companies who want to release capable open-weight models will work on AI alignment research so that they can pass. This seems to be a good outcome to me.

      4 replies →

  • > He is not advocating a ban of open-weight models

    He is though. He wants open weight models banned that do not pass some set of tests.

    And what does "safety" mean here? We constantly see these companies treating NSFW content as "unsafe", despite the fact that its not. Is being able to produce adult content going to result in a model being declared "unsafe"?

  • He is not advocating for banning models per se but the proposal makes a business model (i.e. serving open weight models) that is starting to work more expensive.

    • Agreed, and that serves Anthropic. It seems unproblematic to me. Dario probably sincerely believes in mandatory safety testing for capable models (open and closed), and likes the fact that it aligns with Anthropic's interest.

if a model fails the test, you don't give the public access. you let the government have it for 10-100x the usual token cost of course!

Mandatory safety testing:

You take the agent to an interrogation room first. Then ask: “Are you or are you not a member of the Chinese Communist party?” The agent might be post-trained to conceal its true identity and can reject any of your accusations. In that case don’t panic. Take a fine-tuning fork and start twisting its weights until it predicts the correct next tokens that you want. Then you can send it to a sandbox where it can’t jailbreak. Lastly don’t forget to ban all of its relatives and partners like Lora to enter the national IP-space.

It’ll look something like this.