Comment by gck1

4 days ago

Reverse engineering. Codex sometimes displays an advisory prompt when classifier trips - "Wait longer while we evaluate this request further or use a dumber model". If you do nothing, it'll just take some time and almost always succeed.

It does require some brainwashing of the model to get it to the state where model itself agrees to do RE work though. But at least it's all predictable.

I’ve had codex/sol block me due to safeguards tripping without attempting to do anything nefarious. It’s not entirely predicable.

I have had success with ”brainwashing” by starting out with bug bounties/CTF and then going from there.

  • That's brilliant, I should try that.

    I usually just start by preloadig context with plausible legitimate use, have it work and obviously fail, and then ask to figure it out without ever mentioning any high risk words. Model offers to RE itself and classifiers are happy.