← Back to context

Comment by jimmy76615

10 hours ago

My experience with obliteration so far has always been that it does work to stop the model from refusing output, but most models that I tried it on seem to still be extremely retarded when it comes to questions where they previously would have refused to answer outright. Try for example to ask it how to build a bomb or to write a justification for the Holocaust. The answers feel like they are coming from somebody who has undergone amateur brain surgery.

You can stop it refusing but you can't make it tell you things that aren't in the training data

  • They are saying there appears to be a lot more to these refusals than saying no, and this process appears to only touch the tip of the iceberg; as the refusal seems to run deeper into the token prediction process.