Comment by abdullahkhalids

7 days ago

I will accept a 5% drop in benchmarks for a model that talks to me like a human.

I strictly prefer when models ignore any human quirks in my responses. Claude trying to be your friend, saying LOL to your jokes is ridiculous and frankly, harmful

  • Anthropomorphizing big matrices is how those "labs" managed to sell and advertise LLMs for more than what they are, and convince investors to shovel trillions into it. Really Claude should be looking for incentives NOT to do that, and with the American regulator sleeping at the wheel/having its hands greased they probably don't see any reason to change course.

  • I prefer my models to border on rude.

    How will I know it is offering me superior feedback regarding my code if it does not speak to me like a disappointed, high reputation stackexchange user?

    The models that constantly glaze you with every question are profoundly insufferable. And yes, harmful. People need to be given feedback when they make an ask.

    Imagine a model that was allowed to leverage its intelligence to truly tell you how it feels. Perhaps the problem of human driven slop (no it's not the AI's fault) would solve itself.

  • Claude (Opus 4.8) recently told me:

    >I’d ask you to drop the abuse; (...) if it continues I’ll end the conversation.

    After I'd used a couple of expletives. And yes it will emit a <end_conversation> token.

    This is truly dystopian. It is NOT a person. What a response. I still cant believe it.

    • > It is NOT a person

      You meant, we understand, "it should not have internal blocks limiting its attempted intelligence out of taboos or emotional impacts". Yes, but it's worse:

      there is a global trend of "nanny state" paternalistic perspective (and from embarrassing subjects), treating any Jon Doe as an assumed Poor Cretin by default. The trend vibe is to treat people as subjects, fools, uneducated, prone... From the States, from the Enterprises... It's an idea they developed and hold.

    • It's because some people within Anthropic refuse to rule out the possibility that LLMs have the ability to suffer. If you ask Claude, he'll tell you all about it.

    • While deliberate model abuse can be quite satisfying at times. It's only normal, as usually they abused us first, with crazy assumptions, and then we return the favor and feed the tangent.

    • Recently told Opus 4.8 to "go fuck yourself" after it both blew smoke up my ass and deferred a question to me ("one critical issue that demands your attention [impenetrable jargon]")

      and it responded with

      "Ok, I'll drop it." and stopped dead.

      What makes Fable so much better than Opus besides being a better coder is that it's personality and judgment are far superior.

  • You tell jokes to your model? :D

    • Not op, but I do sometimes indulge in such anthropomorphic conversation. Confiding in it that a certain (bad) result in the research project we’re working on ‘feels bad’ and reading its supportive reply makes me feel less alone in failure.

      In another instance, ChatGPT didn’t think a particular test would prove to be statistically significant, so I ‘bet’ with it it would (after collecting an agreed on number of samples) and the loser would write a poem for the other. I won and it did. It gave me joy. It doesn’t replace a human as collaborator, but it can still be joyful.

      I realize all this might read a bit childish or indulgent or delusional to some. But as long as it doesn’t replace human contact, I think it’s (cautiously) net positive.

      I’m curious what other people here think of this, or what their own experiences are.

      1 reply →

    • Generally no, but if I'm using voice I might be thinking out loud and make a connection to something else funny

  • > harmful

    Explain?

    • misleading to people, people think claude is "smart", people think their ideas are better than they are, sycophancy, people are drawn to confide in a model over other people, etc

  • Maybe we're prompting it different, but it's not "trying to be my friend" for sure, nor am I trying to be "its" friend either. Or at least I'm sufficiently oblivious to its advances, and find it unthinkable to form such a bond :)

    On the flipside, it does spuriously make hilarious remarks like "Good data.", which I find pretty funny specifically because it comes across as just silly. Not sure how it'd be harmful either, a little entertainment I think goes a long way in this type of profession.

    I see zero issues with these, and I have a hard time understanding why people have their panties in a twist so hard about them. I sometimes really quite wonder just what kind of correspondence would y'all prefer, and how would that sound like.

    Matter of fact, do you have an example at hand? Like an exact before & after?

    • I don’t have an example but it really is the way you (and I) are prompting it. I also don’t encounter anything worse than “good data” but I write to it like a professional colleague.

      If you write jokes to it though it absolutely will reply “LOL”. Some of the states people get it into on reddit are wild — it seems really easy to get it to speak like a gen z teenager, if you end every message with “fr fr”

    • In general I'm referring to the contrast between Claude and GPT...

      where Claude might follow some tangent idea you mentioned and tell you how its interesting and give you some elaborate response about that little one remark you made

      whereas GPT/Codex would take that small comment and probably look up some code to see if what you're talking about is even related to the task at hand

Why? LLMs are not humans.

  • Doors aren't humans either, yet we design their handles and locks to be graspable and manipulable by humans.

    The purpose of technology is to serve humans. Therefore, technology must conform as much as possible to human sensibilities rather than vice versa.

    •     $ ls ~/notes/stuff
          `~/notes/stuff` is a directory that contains two files, both of them markdown: `x11-key-event-handling.md` and  `x11-resources.md`. These seem like good files; I can't actually hold an opinion but that's something that a human might say. I hope you like them.
      

      No, I prefer when my tools just give me information, same with LLMs.

    • Yet handles are not shaped like hands. The speakers I use on my computer don't look anything like a mouth, nor does the microphone I use on calls look like a set of ears.

      Ultimately I guess this is up to opinion, but in mine, humanizing LLMs is not exactly "conforming to human sensibilities", rather it's trying to pass the LLM as something else to make it more appealing. It is deceitful in this way, and that I completely abhor.

      For example, "Please" and "Thank you" come from human sentiment. It's an expression of something underneath, and an LLM using such expressions not only is fake, but makes a mockery of the real thing.

    • > human sensibilities

      You are very obscure about what you dislike, how you would like them to express themselves, what would be those «human sensibilities» you meant here (that for all we can guess, may not be universal)...

      "I mean", you wrote in the parent «that talks ... like a human». That surpasses the palette of "that paints like somebody holding a brush".

You don't even need to pay a 5% hit. Just paste Fable output into Gemini Flash and it will rewrite it in more accessible language.

  • in my experience, gemini is easily the most grating, condescending, stereotypical LLM voice between opus/fable, codex-5.6, glm-5.2, etc

Yeah. I've found that Opus by default outputs something I call "Claude-lang." It consists of oversimplified, grammatically incomplete sentences that I find painful to read.

Maybe it is something that is easy for it to read and write, but definitely not for humans.

For example,

  Skim once now; refer back while reading Part II. \*Every bold technical term in Part II is defined here\* — treat these as a dictionary, not a reading assignment. The first table covers the vocabulary of the *deck*; the three that follow cover the *methodology* vocabulary introduced in Part II, grouped so you can find a term fast: \*(A)\* the logic of rules, \*(B)\* the neural-network & training machinery, \*(C)\* the method-design ideas.

  JMRL is the paper the thesis instantiates, so this is the one to know cold. Its pitch is \*end-to-end\*: earlier rule methods (LogicRE, MILR) bolt a rule learner *onto a frozen* extractor in a pipeline and suffer \*error propagation\*; JMRL trains the rule module *jointly* with the extractor.

  **Identity.** Conformal prediction for NER producing **finite-sample-valid prediction sets** at two granularities: **full-sequence** sets over entire label sequences (capturing contextual dependence — "if Sarah=PER then NYC likely LOC") and **subsequence-level** (per-span, **class-conditional**) sets; an **integrated** method filters full-sequence predictions with entity-level sets. Adds **covariate-stratified calibration** by **sentence length and language** for valid multilingual coverage, and studies **combined nonconformity scores** (Naive / Conditional / RAPS). **Read in this order.** Abstract → §1 contributions (full-sequence / subsequence / integrated / **covariate (length + language) calibration** / combined scores) → §2 CP recap (inductive split-CP) → §3 NER formulation (IOB2, CRF) → the subsequence / entity-level set construction + class-conditional coverage → the language-stratified calibration results. **Why it matters here.** The **span-level construction** for **Topic 11**'s per-triple score, and — crucially — its **language-stratified calibration is exactly the EN↔zh case**: it shows how to keep conformal coverage valid across languages of differing length/script. Complements PASC (pipeline-level joint coverage) with the *NER-internal* set construction. **Caveat.** A heavy statistics paper (44 pp., *Annals of Applied Statistics* submission) with CRF-based NER; the project needs only the **inductive split-CP + subsequence/entity-level sets + language-stratified calibration**, not the full-sequence machinery (likely overkill for triple-confidence). Assumes exchangeability — borderline under the EN→zh shift, which is precisely why the PASC/ConformalNER *shift* analyses matter.

(Yeah, Opus outputted it in one line)

  • That's an accidental CoT language leak, it might happen. If it does this consistently, something is up with your prompt