Comment by jtorsella
2 hours ago
Did you read the paper you are linking to? It records a 0% false positive rate for evaluation of human-authored controls, and fewer than 5% of the hybrid and humanized papers had their AI levels overestimated by pangram. It is not completely clear from the data, but it seems that depending on whether the n=2 overestimates were “100% ai generated” assessments, the study you link says pangram’s 100% AI assessments were right anywhere from 95-100%
Still, I guess if you would care you could let the LLM perform loops with pangram tests, changing the text till it scores well enough.
My citation was correct. To question to ask yourself: For my use case, is it okay that Pangram can't reliably tell the difference between "100% AI generated" and "AI assisted"?