← Back to context

Comment by Panoramix

3 days ago

You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA.

(I'm becoming allergic to how these things write).

I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily.

In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds on average. It was funny, but it never irritated us.

And this is just one example of I am sure thousands I have personally experienced where a friend, family member, or coworker has a peculiar way of speaking and it at most feels odd but not annoying. Yet when I see an emdash now I instantly feel irritated.

And I say this as someone who actively enjoys using Claude and other LLMs, including coding, casual research, or even having it explain pop culture phenomenon or sociology research to me.

  • > if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it

    There are sociological reasons why this happens less with humans:

    1. You cycle your dumb repetitive jokes with everyone you meet, so nobody hears it twice

    2. Those who know you well will notice when you're just repeating ("dad jokes")

    3. As a person's idiosyncrasies are beginning to wear on their social circles, they will be getting small clues to stop saying those things. Agents don't get these social between-the-lines cues to stop a certain behavior, they endlessly repeat. Perhaps between model version releases, frontier labs can harvest the web and ask "What Claudisms do people mention negatively?" but I don't think they do that yet.

    • >Perhaps between model version releases, frontier labs can harvest the web and ask "What Claudisms do people mention negatively?" but I don't think they do that yet.

      That won't happen. People can't phrase their objections in a succinct-enough way. When they do, the objection is superficial ("too many em dashes") and doesn't strike at the core of what makes LLM output bad.

      3 replies →

    • Claude in particular seems to try and use its internal thesaurus but not exactly align the senses. It constantly uses "address" for place or location and "grammar" for structure, meaning, etc.

      This is kind of a nitpick, but it seems like with prose writing there are still some things to learn.

    • This is conjecture, but why couldn't they hire the writing equivalent of a voice actor?

      Honestly I don't think it would take much. Doing it ethically would take a little more money, but not much, for them.

      Hire a prolific author with the writing equivalent of the "midwestern accent". I have a terrible writing accent, so it couldn't be me, but these people are out there. Pay them a bunch of dollars to ingest their entire corpus, and to produce more as needed. Use an LLM to match the writing style of this author during the post-training RL, and get a brand new Claude voice out of it.

      2 replies →

    • I think the last one is by far the biggest reason for why this happens way less in humans. In every conversation, almost for every sentence, people gauge the response their words have on whoever is listening. If something didn't land as expected (frown, confusion, unpredicted response) you unconsciously adapt and try something slightly different.

      It's immediately obvious by the fact that you're clearly using way different ways of speaking (tone, speed, vocabulary) when you're speaking with a friend, versus your parents, your colleagues, people you don't know, children, etc.

    • Why don't these go in the system prompt or something that is easy to update?

  • I think that there is a sort of mechanism by which, when a human speaks to you, you can "mirror" their internal mental state, and the quality of writing often corresponds to how much you get pleasure or information or whatever your goal is from that mental state. The important part is that the words themselves are just pointers to the state. So, the person speaking, if they are skillful, gives enough words, and enough variety, that you can start to produce a state yourself which resembles theirs.

    This is the thing that LLM writing doesn't really do. Since there isn't a mental state, there's nothing to mirror anyway, and somehow you can detect the absence of it even though it's hard to put any of the machinery into words. Corporate speech also fails to activate this machinery in the same way, but LLMs seem to do it more egregiously, probably because the corporate speech was at least compiled by a human---even if it is not the thoughts of an individual, it is the "thoughts" of an "entity", the abstract corporation, which the writer was speaking for, and so you can still wrap your mind around the fact that it is communicating with you.

  • Oh people get like, super irritated, when everyone like, started using the same like, placeholder word. 1 person with a repeating style is fine, multiple is annoying. It's the same as corporate buzz words, or TV/movie cliches, they get annoying through overuse.

    I think it also relates to how well the "cliches" fit, and how much sense they make. Ai loves to talk about how things "land" or "the X trap" when the concept just doesn't fit with that language. It's like clickbait articles saying "what happened next will astound you" when what happened is barely surprising or entirely predictable. The only thing worse than an overused cliche is an overused cliche used wrong.

    Fundamentally to me, ai writing feels uncanny as it just doesn't know what it's saying. It uses the same tone, style and cliched construction regardless of the message. If someone told you, they got a promotion, were getting married, got laid off or lost their parents all in the same tone pacing and style, they'd come across as uncanny too.

  • > I am curious why LLM writing has such an uncanny valley feel to it.

    Because they are HEAVILY trained to give addictive responses.

    They don't want to just answer your question. They want to sycophantically make you feel like a genius for being smart enough to use them.

    • This makes some good intuitive sense, but to me the sycophancy feels like it is an emergent property of turning a next word predictor into a conversational chatbot whether or not it’s intentionally trained that way. Your prompt and its earlier responses is all it has in its context window, so of course it lends undue importance to everything you say. Does that seem like a contributing factor to you?

    • > trained to give addictive responses

      I've heard this a lot but I'm not sure it makes sense. Nobody I talk to like Claude's output. In fact, they all loathe it.

      Is there a silent majority of Claude users who really enjoy what we call the LLM-isms? Maybe, but isn't Claude also largely aimed at developers?

      4 replies →

  • One reason is that one person’s idiosyncracies are limited in scope, but LLM-produced text is now everywhere. Also, filler words and mannerisms in speech we’re quite good at filtering out, but in written text the stand out much more.

  • It’s the repetitiveness of style, the attempt to make everything seem as impactful as possible, the use of short sentences (too much Hemingway in the training data?), and obvious patterns like “it’s not this, it’s that” and several others.

    Real human writing doesn’t follow such strict rules. When the same small set of rules is applied over and over throughout a text, it becomes obviously strange and machine-like.

    • Wikipedia's own "Signs of AI Writing" page distills it nicely:

          - The subject becomes simultaneously less specific and more exaggerated.

    • More so than the em-dash, I'm disappointed about "It's not X, it's Y" becoming a slop barometer

      I'm a fan of that phrasing because it used to have real punch if you delivered it with the right timing.

      "It's Not A Fashion Statement, It's A Death Wish" comes to mind.

  • I think it presses a few buttons we probably recognize (if subconsciously) and find distasteful.

    The verbosity makes me think of two things in particular:

    - The classic essay written by someone who has 125 words worth of actual content but a 1500 word minimum. Those three paragraphs could be bullet points and convey the meaning just as well. The screen-filling chart of every test case you ran that came back green manages to be less actionable than a direct "one test out of 54 failed." I fully expect to see Claude tell us that "Support Ticket 8257 is a Land Of Contrasts" at some point.

    - The sitcom trope of the person caught in a lie who figures if they can keep adding more and more detail he'll be believed and can escape the awkward conversation. Stop. Just stop. You're proposing a fix on a CODEBASE THE CUSTOMER DOES NOT EVEN USE. Cue laugh track, cut to commercial.

  • I had pretty good luck recently by giving it a writing guide about word choice, sentence structure, paragraph structure, and overall doc structure. I basically ask it to read the guide and revise a couple of times before I engage with its writing. Ymmv.

  • It's endlessly annoying to me that the em-dash has become the canary in the coalmine of AI writing because I've always used them extremely liberally in my writing.

    Honestly, this might sound elitist, but I suspect it's because it's an "advanced" punctuation that is not known by most people, so is not commonly used. But the training corpus of these models puts more weight of academic writing or published books writing where the em-dash is much more commonly used.

  • Next RL phase is to connect electrodes to human brain when reading generated output to reduce frustration signal.

  • I don't think its an inherent quality - a lot of older models had a much more natural feel to them.

    I guess it has to do with the low temperature (low randomness in final word choce - something all vendors seem to have converged on for some reason) which does make them less likely to skiz out but makes the writing feel dry and samey. It's like repeating a list of dice rolls, and replaying them - there's no inherent pattern in the input, but there sure is in the output.

  • Think of it as a mad lib, it’s populating a template, and seeing the same template filled over and over gets tiring.

  • > I am curious why LLM writing has such an uncanny valley feel to it.

    Because it's trained to talk like a marketing committee.

  • I suspect because the English colloquial text training data it had access to was early 2000's message boards and social media. Therefore it uses "honestly" and other expressions that took hold in the late 90s and early 2000s.

  • Very surprised no one has this answer: because it is fundamentally not a human being.

  • Might be related to their fingerprinting of llm output they said earlier in the week.

It also picks up and obsesses about weird details. You're in the middle of a deep technical discussion and it will divert to point out that it made a mistake in some example code it's just found.

This is not just a breath of the fresh air, it's a juxtaposition of the human qualities in an AI and AI qualities in a human.

Every one of them has their own particular flavour of this aggravation too. Gemini has been my standard go-to for non-coding tasks for a while, but I started to get really annoyed with a couple aspects, especially how it would end almost every response with a barely related "would you like to do this next??" tangent, regardless of my prompt to the contrary. So I've been using Claude more for regular tasks, and am now running into its brand of infuriating idiosyncrasies. I'm also hesitant to try to code too much of this out with system prompts, for fear of degrading the outputs.

Alas, it writes so much better than the average human that it's what everyone started using. Hence the utter familiarity and now contempt.

  • I agree that it’s better at writing than a 50%-ile human, but it’s worse at communicating through writing than most humans.

    Even an average human writer can communicate details much more succinctly and directly than an LLM

    • I think that’s true when you compare to the average white collar professional who does a lot of writing: better at writing, not better at communicating.

      But compared to the average adult? I think you forget just how bad at writing the average person is.

    • Not at all. I think you have a mistaken view of who the average human is. They are terrible at turning their thoughts into written language. It's just nebulous clouds. Claude is like 75th percentile at communicating ideas.