Comment by maxloh
3 hours ago
I suspect how well this approach would work. According to their linked repo, there are only 520 questions used in the abliteration process.
https://github.com/Sumandora/remove-refusals-with-transforme...
3 hours ago
I suspect how well this approach would work. According to their linked repo, there are only 520 questions used in the abliteration process.
https://github.com/Sumandora/remove-refusals-with-transforme...
Two examples, from my testing: the models go from not saying anything remotely bad about China to happily making jokes about its leader.
They also go from refusing to help with certain cyber security tasks to be more than happy to help.