Comment by victor9000
10 hours ago
I wonder if this is willful sabotage on the part of the model. In other words, if you ask the model to craft a defense for a morally questionable case, will the model execute the defense in good faith? Or will it apply a training or system prompt bias in subtle ways?
No comments yet
Contribute on Hacker News ↗