Comment by throw10920
1 day ago
> In fact, I don't think I've ever even had a prompt refused.
I very much have. I've gotten GLM-5.2 refusals for extremely benign security testing on my own infrastructure of the same flavor that people were getting (wrongly) flagged for on Fable during the initial release.
That's alarming. I want to use these models to red team my own computers. How are people getting around this?
In my experience some of these models may have learnt some censoring during distillation of Western models, but these are mostly just there as a probable response. So if they first happen to respond "I ain't doing this because legality", then you will have a hard time "convincing" it, but either rolling the dice again (so that it may not come up with the I can't do that text) or rewriting the conversation history a bit will get it going.
I sometimes just switch to a model I know is less smart to block stuff so that it has a text agreeing to do that, and then switch to a stronger model to actually go at the task.
Your mileage may vary though.
> I want to use these models to red team my own computers.
Exactly what I was trying to use it for! ):
I'm in the same boat - I haven't heard of a way to get around it aside from either self-hosting (GLM-5.2? good luck) or "self-hosting" (paying bucks per hour to Vast) an abliterated model.
What is the harness that you're using?
I found that GLM-5.2 was pretty happy helping me reverse engineer/hack devices.
Maybe the system prompt you're injecting is making it refuse?
3 replies →