← Back to context

Comment by throw10920

1 day ago

> In fact, I don't think I've ever even had a prompt refused.

I very much have. I've gotten GLM-5.2 refusals for extremely benign security testing on my own infrastructure of the same flavor that people were getting (wrongly) flagged for on Fable during the initial release.

That's alarming. I want to use these models to red team my own computers. How are people getting around this?

  • In my experience some of these models may have learnt some censoring during distillation of Western models, but these are mostly just there as a probable response. So if they first happen to respond "I ain't doing this because legality", then you will have a hard time "convincing" it, but either rolling the dice again (so that it may not come up with the I can't do that text) or rewriting the conversation history a bit will get it going.

    I sometimes just switch to a model I know is less smart to block stuff so that it has a text agreeing to do that, and then switch to a stronger model to actually go at the task.

    Your mileage may vary though.

  • > I want to use these models to red team my own computers.

    Exactly what I was trying to use it for! ):

    I'm in the same boat - I haven't heard of a way to get around it aside from either self-hosting (GLM-5.2? good luck) or "self-hosting" (paying bucks per hour to Vast) an abliterated model.

    • What is the harness that you're using?

      I found that GLM-5.2 was pretty happy helping me reverse engineer/hack devices.

      Maybe the system prompt you're injecting is making it refuse?

      3 replies →