← Back to context

Comment by verdverm

1 month ago

What happens if a model passes the government tests and then later someone fine tunes it to behave differently, without making their changes public?

Any reasonable safety testing should include finetuning and safety margin to account for others may do better finetuning.

  • I can fine tune significant behavior changes, there is little model developers can do to prevent this (aiui), so this effectively becomes an blanket ban

    • Yes, I agree it is effectively a blanket ban (above some capability) for now. I hope AI alignment research advances in the future so that it is not so.

      1 reply →