You are specifically recommending asking the model which of two versions is better (the quote in my top-level comment).
We both agree that they are poor "make it better" machines, but I also believe they are bad A/B testers and I'm using the former to demonstrate the latter.
It doesn't matter who writes what, what matters is that LLMs have a preference for LLM-shaped writing. By A/B testing against an LLMs opinion, you are optimizing in the direction of LLM prose even if the LLM never writes any of the prose itself.
You are specifically recommending asking the model which of two versions is better (the quote in my top-level comment).
We both agree that they are poor "make it better" machines, but I also believe they are bad A/B testers and I'm using the former to demonstrate the latter.
You are writing both paragraphs. You’re specifically not asking a model to make a better paragraph. That would be a load-bearing debacle.
It doesn't matter who writes what, what matters is that LLMs have a preference for LLM-shaped writing. By A/B testing against an LLMs opinion, you are optimizing in the direction of LLM prose even if the LLM never writes any of the prose itself.
1 reply →