Comment by uoaei

5 years ago

NLP researchers are tearing their hair out about this right now, since people are posting mountains of GPT/etc.-generated text online with no easy way to distinguish whether it's of human or other origin.

Reminds me of a scifi book where the Internet-analogue is so corrupted with junk deliberately injected by filtering services so that they can sell you the filters that it's impossible to use "naked".

I think it's either Neal Stephenson or maybe Stephen Baxter, but I'm not sure which book it was an aside in (it's not Fall, I haven't read that yet, though that appears to have a similar idea).

Well... you can read it.

I might be wrong, but so far I think it's been pretty easy to tell if text came from a human or a deep-learning system.

Granted, that probably doesn't scale well.

  • That has been true up until very recently, but lately there has emerged an uncanny valley that has confused the distinction between "underpaid freelance (ESL) writer farm" and "shoddy but roughly convincing neural language model" so that you may be convinced the blogspam you may happen upon across the internet may just as easily be computer-generated as anything else.

    • If GPT-3 et al. are producing results on par with "underpaid freelance (ESL) writer farm" then wouldn’t the next step be to focus only on better writing? i.e. books, newspapers, magazines, etc.

      Obviously the corpus (perhaps 1 trillion to to 10 trillion words?) will be exhausted by a sufficiently large model so there’s an upper bound.