Comment by mrbungie

2 years ago

They do work detecting LLM outputs that are sampled "naively" (when the model/user is really not trying to pass it as human output).

I copied a prompt translated from spanish to english using ChatGPT Plus in a GPT-4o Azure OpenAI Service endpoint. It did work in Spanish but didn't run in english because the default AOS Content Filters detected a jailbreak intent. It was quite weird.