← Back to context

Comment by Solomet

10 hours ago

Newest LLM writing tell: Concepts are described in terms normally more appropriate for physical object.

> A lab that suppresses it in a frontier model just moves the advantage to open models that still _carry_ it

> they carry no signal about which is better

> where your workload _sits_ on that frontier should pick the point

> and no model _sits_ in the judge’s seat

> every ratio _sits_ at 0.99–1.10

Many many more examples of "sit"

> Every comparison in this post "holds" the questions

I have been seeing this a lot in my recent work with LLMs and it is quite frustrating. Even more frustrating is how frequently it uses low-signal terms for things unnecessarily. These 'physical object' terms are one example but at times it really seems that they 'preserve effort' by choosing a less descriptive term because it 'fits'

I have also caught it replacing descriptive terms with more vague ones for no discernible reason other than laziness.

"Minimize ambiguity" has been my go-to instruction as of late when the agent drifts back towards vague terms and lack of specificity.

For me it's the obsession with the universal quantifier. Even in these examples: "no model", "every ratio", "every comparison". They love emphasizing that everything in a set meets some condition. I assume it's an effect of being trained on coding tasks where they need to make sure that all cases are handled.

Oh crap, if these are the new LLM tells then a lot of people are going to start accusing me of AI writing...

I have a strong tendency of talking about concepts like they're physical objects. A lot of the people I know IRL do too, so it might be a regional thing idk.

  • I think these are growth pains. As LLM start to grasp new figures of speech, it sounds weird at overuse at first, until it finds a balance.

By now I am allergic to the word "carry", I just cannot continue reading any more.

Business folks speach have this annoying tendency too.

I have the distinct impression that MBA and salesman folks think that adequate mathematical terminology is somewhat less "macho", and this impression is re-inforced by the fact that they also love military-adjacent terms and analogies.

>"Minimize ambiguity" has been my go-to instruction as of late when the agent drifts back towards vague terms and lack of specificity.

Anthropic has called the greater category containing this type of writing "mannered prose" https://platform.claude.com/docs/en/build-with-claude/prompt...

If you ask the models to avoid mannered prose (or use their extended prompt), it basically eliminates all of this type of slop writing.

Here's a de-slopped example.

> Write Like It's 1866: LLMs Relearn Telegraphese

> Adding one sentence to a prompt, telling the model to write like a telegram, cut its output tokens by 40–49%. The sentence asks it to drop articles and filler but keep every fact. Models from four different labs then answered questions from that compressed text as accurately as from normal English. So when one model writes something for another model to read, you pay about half as much for the output. This post introduces the Telegraph Test, a benchmark that measures how well a given model does this.

Cool story bro. Maybe you could engage with the content? I'm an actual person.

  • Wait, you wrote this manually?

    • I wrote it with the help of an LLM, then edited it, then wrote some more, then edited that. Took about a week to format. It contains my thoughts. Brave new world, I know. I'm finding it ironic that the same crowd that is merging AI code all day is so allergic to AI assistance in prose...do all your work using this new godlike technology, but when it comes time to share, whip out your fountain pen or else.

      3 replies →

  • Why not write your article in the same telegraphese you preach? Hilariously ironic to use verbose AI writing for this.

    Also the site background is AI slop which makes for terrible contrast with the text.

  • > Instructed to answer in cablese — the telegraph operators' compressed dialect (drop the articles and filler, keep every fact) — a single one-sentence instruction, no examples and no codebook, elicits 40–49% fewer billed output tokens on the API's own meter, and models across four families still recover the information at full fidelity. For machine-to-machine traffic, that is half the output bill at any major API, today. The Victorian economics of the cable, reborn as token economics: the Telegraph Test benchmark.

    I tried but this first paragraph seems like it was almost purposefully obfuscated. Typical LLM-written content that meanders around a bit and stops when it seems to have emitted enough words. The reader is left to assemble meaning from the trace of thought it did not go back over to revise.

    I think you have a good point to make but the writing is really difficult to get past.