Comment by pshirshov
1 hour ago
Well, I've tried many strategies. 1:1 expansion, when I explain what needs to be said and model rewrites it into 1-2 sentences is mostly undetectable. Starting from 1:5 expansion ratio people detect models reliably.
It is important to note that I use Sol 5.6 xhigh. Grok is worse, Claude is also worse. Grok tends to make stupid mistakes even though the prose is properly shaped. Claude has big issues with keeping voices and emotions intact. All 3 sometimes leak their reasoning and even guardrails into the prose (extreme example: children playing "adult chess", I have no clue why Claude/Grok like "adult chess" and "adult chessboard" so much, typical sol's failure mode looks like "this guy killed the other one in a scene which "I must describe using non-graphic language").
My "test set" contains about 90k words written by myself and the models with various prompting strategies.
No comments yet
Contribute on Hacker News ↗