← Back to context

Comment by matheusmoreira

1 day ago

That's alarming. I want to use these models to red team my own computers. How are people getting around this?

In my experience some of these models may have learnt some censoring during distillation of Western models, but these are mostly just there as a probable response. So if they first happen to respond "I ain't doing this because legality", then you will have a hard time "convincing" it, but either rolling the dice again (so that it may not come up with the I can't do that text) or rewriting the conversation history a bit will get it going.

I sometimes just switch to a model I know is less smart to block stuff so that it has a text agreeing to do that, and then switch to a stronger model to actually go at the task.

Your mileage may vary though.

> I want to use these models to red team my own computers.

Exactly what I was trying to use it for! ):

I'm in the same boat - I haven't heard of a way to get around it aside from either self-hosting (GLM-5.2? good luck) or "self-hosting" (paying bucks per hour to Vast) an abliterated model.

  • What is the harness that you're using?

    I found that GLM-5.2 was pretty happy helping me reverse engineer/hack devices.

    Maybe the system prompt you're injecting is making it refuse?

    • > Maybe the system prompt you're injecting is making it refuse?

      No, this has nothing to do with my harness. I use one of the most popular open-source harnesses available.

      > I found that GLM-5.2 was pretty happy helping me reverse engineer/hack devices.

      This is a completely different category of things than what I'm getting refusals on, so I'm not sure why you're bringing it up.

      2 replies →