Comment by deadbabe
2 days ago
> Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is.
I have prompted out a lot of disturbing and inappropriate content with GLM-5.2, that would have left other American models blanched in the face or clutch their pearls. I think this is mostly a reference to Anti-CCP stuff.
In fact, I don't think I've ever even had a prompt refused.
> In fact, I don't think I've ever even had a prompt refused.
I very much have. I've gotten GLM-5.2 refusals for extremely benign security testing on my own infrastructure of the same flavor that people were getting (wrongly) flagged for on Fable during the initial release.
That's alarming. I want to use these models to red team my own computers. How are people getting around this?
In my experience some of these models may have learnt some censoring during distillation of Western models, but these are mostly just there as a probable response. So if they first happen to respond "I ain't doing this because legality", then you will have a hard time "convincing" it, but either rolling the dice again (so that it may not come up with the I can't do that text) or rewriting the conversation history a bit will get it going.
I sometimes just switch to a model I know is less smart to block stuff so that it has a text agreeing to do that, and then switch to a stronger model to actually go at the task.
Your mileage may vary though.
> I want to use these models to red team my own computers.
Exactly what I was trying to use it for! ):
I'm in the same boat - I haven't heard of a way to get around it aside from either self-hosting (GLM-5.2? good luck) or "self-hosting" (paying bucks per hour to Vast) an abliterated model.
4 replies →
It is cliche, but I haven't had good luck with having Chinese models openly discuss historical topics like Tienanmen Square. The US models don't seem to have a problem discussing history, even if it points an unglamorous light on the US government.
And there's a reason for that: the US government does not compel model trainers to train their models to paint them in a favorable light, while the PRC does.