← Back to context

Comment by Lerc

10 hours ago

How can you tell if people can accurately identify AI generated text?

If a person reads AI generated text and does not notice, they by definition will not know about it.

There have been numerous cases of people accessing human created content as being AI.

There are instances where it seems relatively uncontroversial that it is AI generated, but without knowing both the amount of AI content people are exposed toand the amount that they register I don't think you can draw a conclusion of the overall state.

I don't think this is about edge cases where someone has successfully disguised the writing to some degree: the current crop of LLMs have some pretty blatant (and frankly annoying) habits by default, ones that are hard to miss once you have read a decent amount of their output. If I had to describe them broadly, I would say they are a collection of habits which are common in certain kinds of persuasive and emotive writing, but are usually applied way out of proportion to the topic at hand, which tends to make the result quite grandiose, overly dramatic, and tiring to read: a LLM will often write a TODO app README like it's a cross between a thriller novel, a political speech, and a bombshell news article. There's lots of specific tics (and just by sheer volume and uniformity almost any habit an LLM picks up is going to rapidly shoot into cliche regardless of its own merit) but this is the general effect which I think is objectionable independent of the source of the text.

I do think the sensitivity to it can vary a lot: it depends a lot on how much and how closely you read the text, and how much exposure you have to LLM writing. Certainly it seems like a lot of people just don't really notice, or at least don't care much.

  • > the current crop of LLMs have some pretty blatant

    This is just "em dash redux." Except now we've moved on to accusing anyone who does "It's not X. It's Y." of being AI. In six months, it'll be "use of the word 'petrichor'" or something.

    • Idk, it's more like "your writing is cliché and I don't feel like reading it because I've already read something that sounded similar countless times and it wasn't worth the read". The source of the clichés being an LLM. And maybe now humans are writing the same way as LLM output, I still am not going to read all that, sorry. If I see a sea of clichés, I'm going the other way.

      I'm also not reading pumpkin spice murder mysteries for a similar reason. I'm also not reading stories where everybody clapped. Actually, I'm already familiar with petrichor, so unless someone has surrounded the word "petrichor" with non-cliché prose, I'm also not going to read all that.

    • I do think there is a tendency to over-index on one or two particularly straightforward tells, and for any given feature of LLM writing you can find places where people do also use that feature (they had to learn it from somewhere, and in a lot of cases it is good writing practice — for the context in which it is used). But I'm not talking about just that, but also the general tone issue: it's bad writing regardless because it's in most cases just not appropriate for the context it's been written in.

      (TBH I think the biggest likelihood for false positives comes from heavy LLM users picking up their tics: it's a natural tendency and I've already seen a few cases where it seems like that has happened).

    • I went to /show and chose a random project with a GitHub repo. Here's the README:

      https://github.com/ucsandman/declick/blob/main/README.md

      Please let me know if you

          (a) believe this is human prose
      
          (b) enjoy reading this prose
      
          (c) would enjoy reading 100 READMEs like this.
      

      As for invoking petitio principii and questioning other commenters' logical coherence [0], can you politely shove the argumentum ad Latinum up your ass?

      [0] https://news.ycombinator.com/item?id=49582713

      1 reply →

Pieces that people aren't revolted by will be fine. Readers aren't revolting because of a flood of high quality writing though.

  • Some people are still losing their shit over em-dashes, with no other tells, and humans can't use the not x; y construction anymore either, regardless of any other merit to the writing.

    LLM writing is verbose and meandering, but people are making a much bigger deal over this stuff than necessary for virtue signalling purposes. You don't want to read someone else's LLM writing? Get a summary of the page from yours. No time wasted, no pretentious posturing, and you don't make the error of assuming because the piece was written by an LLM that there was no thought put into the subject or there's no value in what is being communicated.

This is very handwavy and dismissive. It is pretty safe to assume that must of us catch it most of the time because the simple fact is so many people just copy and paste whatever the LLM outputs without even trying to edit it or mask that they used one. We’ve all seen so many examples of the exact same cadence and verbiage that we’ve learned how to identify it pretty reliably. The ones who are “slipping past us” are actually putting in the work make not just pasting raw LLM outputs, which is the real issue here. If somebody has edited it meaningfully after the fact then it’s not the same crime.

  • > It is pretty safe to assume that must of us catch it most of the time because the simple fact is so many people just copy and paste whatever the LLM outputs without even trying to edit it or mask that they used one.

    This statement does not logically cohere. "We can spot it because so many people make it easy to spot." You don't see how this is just petitio principii in action?

    • What is your point? It’s not that complicated. Obvious slop is obvious. Maybe there are some humans out there who sound like Claude but I’m not going to force myself through 900 slop blog posts on the off chance that one of them might actually be written by a human.

      Maybe some people stop reading LLM slop purely because it violates their moral principles or whatever but most people bail out because slop is mentally painful to read. If you are a human and you write like today’s AI find a different writing style, not because reads like AI, but because it reads like shit.

AI writing just means "writing I don't like" now. Just like Nazi means whatever and whoever I politically disagree with. Words have lost their meaning.

  • You're right that Nazi doesn't mean Nazi anymore. It means neonazi / white supremacist / white nationalist, which is a much broader group of people that, for some baffling reason, are under the impression that people don't care about their fascism and racism anymore.

    • Clearly not, or this wouldn't be something people discuss at all.

      There are lots of people who belong to the above groups, sure, but at least here in Germany Nazi is now applied to basically anyone who doesn't vote green, it's ridiculous.

  • Just like "violence" can now mean speech you don't agree with, "genocide" means military action you don't agree with, and nobody seems to know what "woman" means anymore, Maybe the solution is to stop using words.