Comment by mrbungie
2 years ago
They do work detecting LLM outputs that are sampled "naively" (when the model/user is really not trying to pass it as human output).
I copied a prompt translated from spanish to english using ChatGPT Plus in a GPT-4o Azure OpenAI Service endpoint. It did work in Spanish but didn't run in english because the default AOS Content Filters detected a jailbreak intent. It was quite weird.
No comments yet
Contribute on Hacker News ↗