← Back to context

Comment by maCDzP

3 days ago

I have had success with ”brainwashing” by starting out with bug bounties/CTF and then going from there.

That's brilliant, I should try that.

I usually just start by preloadig context with plausible legitimate use, have it work and obviously fail, and then ask to figure it out without ever mentioning any high risk words. Model offers to RE itself and classifiers are happy.