Comment by mlsu

2 days ago

Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual content doesn't matter.

- "Introduction that rephrases your prompt."

- "3 paragraphs, with one section of bullet points"

- "The Twist"

- "The Bottom Line"

It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane observation about California burritos, is phrased in exactly the same way. This is obviously an artifact of post-training but it's also kind of how you can tell that this thing is a lot closer to a blindsight scrambler than real intelligence.

You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA.

(I'm becoming allergic to how these things write).

  • I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily.

    In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds on average. It was funny, but it never irritated us.

    And this is just one example of I am sure thousands I have personally experienced where a friend, family member, or coworker has a peculiar way of speaking and it at most feels odd but not annoying. Yet when I see an emdash now I instantly feel irritated.

    And I say this as someone who actively enjoys using Claude and other LLMs, including coding, casual research, or even having it explain pop culture phenomenon or sociology research to me.

    • > if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it

      There are sociological reasons why this happens less with humans:

      1. You cycle your dumb repetitive jokes with everyone you meet, so nobody hears it twice

      2. Those who know you well will notice when you're just repeating ("dad jokes")

      3. As a person's idiosyncrasies are beginning to wear on their social circles, they will be getting small clues to stop saying those things. Agents don't get these social between-the-lines cues to stop a certain behavior, they endlessly repeat. Perhaps between model version releases, frontier labs can harvest the web and ask "What Claudisms do people mention negatively?" but I don't think they do that yet.

      10 replies →

    • I think that there is a sort of mechanism by which, when a human speaks to you, you can "mirror" their internal mental state, and the quality of writing often corresponds to how much you get pleasure or information or whatever your goal is from that mental state. The important part is that the words themselves are just pointers to the state. So, the person speaking, if they are skillful, gives enough words, and enough variety, that you can start to produce a state yourself which resembles theirs.

      This is the thing that LLM writing doesn't really do. Since there isn't a mental state, there's nothing to mirror anyway, and somehow you can detect the absence of it even though it's hard to put any of the machinery into words. Corporate speech also fails to activate this machinery in the same way, but LLMs seem to do it more egregiously, probably because the corporate speech was at least compiled by a human---even if it is not the thoughts of an individual, it is the "thoughts" of an "entity", the abstract corporation, which the writer was speaking for, and so you can still wrap your mind around the fact that it is communicating with you.

    • > I am curious why LLM writing has such an uncanny valley feel to it.

      Because they are HEAVILY trained to give addictive responses.

      They don't want to just answer your question. They want to sycophantically make you feel like a genius for being smart enough to use them.

      7 replies →

    • One reason is that one person’s idiosyncracies are limited in scope, but LLM-produced text is now everywhere. Also, filler words and mannerisms in speech we’re quite good at filtering out, but in written text the stand out much more.

    • It’s the repetitiveness of style, the attempt to make everything seem as impactful as possible, the use of short sentences (too much Hemingway in the training data?), and obvious patterns like “it’s not this, it’s that” and several others.

      Real human writing doesn’t follow such strict rules. When the same small set of rules is applied over and over throughout a text, it becomes obviously strange and machine-like.

      4 replies →

    • I had pretty good luck recently by giving it a writing guide about word choice, sentence structure, paragraph structure, and overall doc structure. I basically ask it to read the guide and revise a couple of times before I engage with its writing. Ymmv.

      2 replies →

    • Next RL phase is to connect electrodes to human brain when reading generated output to reduce frustration signal.

    • I don't think its an inherent quality - a lot of older models had a much more natural feel to them.

      I guess it has to do with the low temperature (low randomness in final word choce - something all vendors seem to have converged on for some reason) which does make them less likely to skiz out but makes the writing feel dry and samey. It's like repeating a list of dice rolls, and replaying them - there's no inherent pattern in the input, but there sure is in the output.

    • I suspect because the English colloquial text training data it had access to was early 2000's message boards and social media. Therefore it uses "honestly" and other expressions that took hold in the late 90s and early 2000s.

    • Think of it as a mad lib, it’s populating a template, and seeing the same template filled over and over gets tiring.

    • > I am curious why LLM writing has such an uncanny valley feel to it.

      Because it's trained to talk like a marketing committee.

    • Very surprised no one has this answer: because it is fundamentally not a human being.

    • Might be related to their fingerprinting of llm output they said earlier in the week.

  • It also picks up and obsesses about weird details. You're in the middle of a deep technical discussion and it will divert to point out that it made a mistake in some example code it's just found.

  • This is not just a breath of the fresh air, it's a juxtaposition of the human qualities in an AI and AI qualities in a human.

  • Every one of them has their own particular flavour of this aggravation too. Gemini has been my standard go-to for non-coding tasks for a while, but I started to get really annoyed with a couple aspects, especially how it would end almost every response with a barely related "would you like to do this next??" tangent, regardless of my prompt to the contrary. So I've been using Claude more for regular tasks, and am now running into its brand of infuriating idiosyncrasies. I'm also hesitant to try to code too much of this out with system prompts, for fear of degrading the outputs.

  • Alas, it writes so much better than the average human that it's what everyone started using. Hence the utter familiarity and now contempt.

The second bullet point, down to the comma in the middle of the sentence, is what has been driving me absolutely batty of late. It's a surefire tell that I cannot seem to beat out of my outputs. It CONSTANTLY does it, even when you say not to.

Between that and the insistence on "this, not that" structure makes me want to install the caveman skill and use it even for non-code workflows.

I've noticed that ChatGPT (whatever model the free version uses by default) likes to phrase answers as though it's correcting me, even when my question doesn't contain any assumptions.

  • Something I've noticed quite a bit on the paid plans as well is that it starts its answers with "I mostly agree..." or "almost correct...," then goes through the list of points I made without actually disagreeing with any of them.

    I assumed this is a system prompt or RL that nudges it to be always skeptical but then it still has like all the models the urge to appease the user.

  • Interestingly, I've been noticing almost the opposite issue. 5.6 Sol frequently starts its responses with "Yes" even when my prompt doesn't contain a yes-or-no question.

    • I've been noticing the same thing since 5.4- it starts with "Yes" yet I haven't asked a question. I think it might be related to the reasoning, like it's answering its own questions?

      1 reply →

  • Claude will sometimes claim that my prompt contained implicit assumptions, then argue against them.

    That can be annoying, but it also happens in debates between people, and sometimes the implicit assumptions are real and (sorry) load-bearing. So it can be an appropriate conversational tack.

  • Doesn’t just mean the model has picked up on what happens when two graybeards who still wear cargo shorts meet and one utters a declarative sentence? :-)

  • One thing I despise about ChatGPT is how it goes on 4 paragraph tangents about how slight details are wrong, and then always ends with a paragraph of bolded fucking text rephrasing some statement in my original question, but with a huge amount of hedging to exclude minute counterexamples.

At this point, I'm basically telling models to not write any English text or prose. Only write code. They are great at writing code. Not so great at writing good English. In software projects, lengthy comments and docs are an anti-pattern: the software should instead be written to do the right expected thing so that you don't have to think about it. I don't want all these tokens polluting my context, either.

  • Agreed, think the first user rule I ever put into Cursor was "Don't write code comments unless absolutely necessary to explain something that couldn't just be inferred"

    • i am having fun imagining an agent existentially freaking out while trying to parse this directive since inference is the entirety of their world

Yes! Ive started getting a feel for AI writing on blogs. It feels slightly verbose and involves "reveals"

"It's not the naked man on your lawn waving a chainsaw that's scaring you. It's the burrito you ate for lunch: it went down easy, but now it's coming for you"

It's the same way how every AI generated poster looks exactly the same. As if there is a single underlying prompt that describes the template of the poster/long-form article, and it does not dare deviate from that.

  • Isn't there? Like everyone using $MODEL is starting from the same base system-prompt. Then our user input is a small bit on top of that core mode. Like what would happen if everyone asked Mikey to paint their ceiling - they'd all be similar and therefore boring.

Am I crazy for thinking that this is a pretty big regression compared to past models? I remember being blown away by GPT 4.5, and I kept using it up until they decommisioned it. I think claude 3.7 sonnet was pretty good too. Gemini seems to be the best one right now for actually talking. Opus is top tier for code but when i talk to it I want to rip my hair out. GPT-5.6 is doing best for me right now among the powerful models.

Perhaps this is related to their new "invisible watermark" concept which would probably require rather contrived language patterns to make possible.

  • I love this theory. "We've invented a new invisible watermark that can detect whether code is LLM written."

    The watermark: counting instances of 'load-bearing seam', 'the hard truth', 'and that's the whole point'.

  • If it's using Aaronson's approach it shouldn't have any noticeable affect on generations. When it picks between options weighted by probability after the generation of logits, it still follows the probability mass, it just uses a known pseudorandom seed so that when you go back and look at the exact choices you can fingerprint it.

    • That's exactly right. And as is well understood, a good pseudorandom generator, despite being fully deterministic, is extremely hard to distinguish from randomness, unless you have the algorithm and key (internal state). Quite smart, really.

  • I think opus was released before they included it on model. Its hard to say, but from what Ive read it doesn’t seem like it would have that drastic of an effect.

    I had the same thought though.

    • The watermarking is independent of the model. The model itself has the probability weights to determine the next token. The watermark is similar to things like temperature and top_p/top_k in that the watermark adjusts the probabilities in a deterministic way that changes over time to hide tells from word choices.

The more I use the AIs the more I feel like I did at the end of reading blindsight. The things are undoubtedly “intelligent” by any practical definition of the word, but they are not aware.

This is also why I advocate against using AI as a writing partner. No matter the argument you lay out, a fresh context window will always have the “a few good things and a few bad things” feedback. There is no higher order opinion to align with.

A related theory here is that Opus is heavily RL’d to be a sub-agent.

If Fable is the primary interlocutor then perhaps there is less pushback on the obtuse language.

Indeed perhaps the convoluted language acts as a kind of Neuralese between models deriving from the same pretrained base.

> The aesthetic is that of an expert slowly revealing an insight to the user.

Ah, that's it! Thank you. I wonder if they are training it to talk like this because that's what their customers actually want? They want a machine genius to lead them.

  • It’s the TED Talk playbook: the crafting of a lecture given by an expert to laypeople to maximize attention, engagement and satisfaction. Every piece of prose is built to pack in as many TED Talk mic-drops/expectation-subverting insight bombs as possible.

The model is generating tokens one by one and that sentence structure allows it to keep its options open rather than committing at the beginning of the sentence

In your reading, what is the distinction between blindsight’s scrambler and real intelligence? My reading is that it’s just as real, and draws out the disadvantages a sense of self constrains intelligence with

  • Sure. I should have been more precise about what is 'real intelligence' here.

    What I mean is that blindsight's scramblers are aliens that cannot share human values. Their structure is completely different to ours, their qualia (or whether they even have it) is impossible for us to understand. In short, they do not have a soul. When Claude does this "slowly revealing a dramatic insight" thing that it does, it does that not because it has judged itself through some introspection as having an insight to share. It does not even know what an insight is or is not. It is not sharing anything, because it is not capable of sharing, because it does not have a soul.

    The aesthetic structure of its replies is a pattern, a constraint on the token distribution, like the color of noise.

    It's my bad to use the word 'intelligence' because it's so overloaded. Will Claude will act as a therapist or produce value or produce a work of art? No. It cannot, because it does not have a soul. I leave it freely open to interpretation whether having a soul is required for "real intelligence." But what I've noticed is that "intelligence" in these discussions is mostly used to denote some capability to produce [economic/social] value. In my mind value is a relational thing, a thing of human feeling.

    • I think I understand what you are trying to convey but I fear you've made the same mistake again, this time with "soul" instead of "intelligence."

      I think what you are getting at is that they are deterministic automata. They are machines. We have introduced randomness to add variation but it is an artificial randomness that simply perturbs the path traversed.

      When we choose words it isn't because of a token distribution, nor because we rolled a die. We choose words because we feel a certain way, the external world, our body and senses are all connected as one system. These machines don't experience moods or get tired or feel better after a good night's sleep. They don't know their audience, we're all the same to them. We have no personal relationship nor can we establish one, as presenting some arbitrary background is not the same thing as a fluid, evolving relationship that accumulates through experience over time. There are no scars or fond memories.

      If these things can truly be intelligent, to abuse your use of the word, then at least we are quite far from holding them correctly.

      8 replies →

    •     Will Claude will act as a therapist or produce 
          value or produce a work of art? No. It cannot, 
          because it does not have a soul. 
      

      I tentatively agree, although I'm only tentative because I don't think it's an interesting question.

      Here's what I do think is interesting. You!

      I mean... yes, you, too but not you specifically. The plural "you" that the english language lacks.

      And so here's what I think is the actual interesting question. Might AI help you create art? Or be a therapist? Or something else interesting and worthwhile?

      Maybe AI won't write the next great guitar solo. I'm pretty sure it won't. But might it help you learn to play guitar? Help you fix your broken guitar amp? Help you understand some tricky parts of guitar playing? Help you work through some tricky tabulature where you can't tell if you're playing it wrong or if the tab is just bad?

      I don't know. But that's my angle for finding any of this interesting.

    • “It is not sharing anything, because it is not capable of sharing, because it does not have a soul”

      soul.md though, just saying

I think the Opus 5 formula is to be the little professor treating your ideas like an essay for grading, or like a buyer analyzing merchandise for purchase.

This twist often involves loosely related or even unrelated bugs or even non-issues, and Claude bringing those into the conversation at that point breaks my mental processing of the response.

Any tips on how to avoid that would be highly appreciated!

It's an artifact of the reasoning process I think. Setting reasoning effort to none works better when trying to change writing style ime

And this is infuriating. I don't want to read all this gibberish anymore. It's making me hate what software engineering has become.

...I mean, on the whole, I'm glad it's detectable. I imagine they could have post-trained it to not be detectable.