Comment by bakugo

9 hours ago

This isn't very effective on any models released in recent years. With older ones, you used to be able to influence writing style significantly by just putting examples in the context, but newer models have gone through so much assistant RLHF, they really want to revert back to their default "assistant voice" during their turn.

You can still influence their writing style in a broad manner that might look correct at a glance, but the repetitive little patterns that give it away will always be there - if it was that easy to get rid of them, don't you think the AI labs themselves would've done it before releasing the models?

I think nowadays defeating the detection probably looks like finetuning a smaller LLM and getting it to paraphrase the text from the other one (or just using a more obscure finetune: it'll probably have its own cliches and habits but it will be at least different). As an added bonus this also likely removes the fingerprinting from the output as well. But I think most people are not going to bother with this.