I wrote it with the help of an LLM, then edited it, then wrote some more, then edited that. Took about a week to format. It contains my thoughts. Brave new world, I know.
I'm finding it ironic that the same crowd that is merging AI code all day is so allergic to AI assistance in prose...do all your work using this new godlike technology, but when it comes time to share, whip out your fountain pen or else.
The problem is that the prose becomes difficult to read because of its lexical quirks and verbiage. One AI generated piece of prose in isolation, fine; but people have developed a 'smell' of it and a mental association with low-quality work.
It might contain your thoughts, but the writing definitely needs work.
With code, people often think just getting it to run correctly is enough (which is often wrong, but that's another discussion.) With writing, the success criteria are more complex.
Even just the first sentence is unnecessarily cryptic, with inconsistent grammar. For example, the parenthetical clause, "the telegraph operators' compressed dialect," is surrounded by an em-dash and a comma. It should be the same character on each side.
One issue you may be having is forgetting that readers don't have all the context about your task that you have in your mind. Writing maximally concisely, which seems to be what you or the model were going for, means that readers have to try to reconstruct that context. Presumably you weren't trying to write the article itself in telegraphic style.
Here's a possible rewrite of that first sentence:
> When LLMs are instructed to answer in "cablese"—the compressed dialect historically used by telegraph operators—models produce 40-49% fewer billable output tokens through the API. Despite this compression, models from all four tested model families still recover the information at full fidelity.
The original "Instructed to answer" formulation is known as a compressed participial construction. It's grammatically valid, but overly concise for the first sentence of an article. Being explicit about what's being instructed helps the reader avoid having to read the whole sentence before determining what its subject is.
The ", and" pattern is something I often catch myself using, but it's often better to break up the sentence instead, and that's certainly true in this case. It also makes it easier to call out the result you're describing more explicitly, with "Despite this compression".
The "elicit ... on the API's meter" is just a very weird construction, and does sound quite Claudish. I doubt "elicit" is the correct word here. A prompt elicits a response from a model, but the model itself doesn't "elicit" anything when responding to a prompt.
If you want to improve your writing, I recommend reading. Not just non-fiction - reading fiction can be a great way to improve your writing skills. It trains your brain in what's correct, essentially - assuming what you're reading is well-written!
However, I'm as allergic to slop in my code as in my reading. I'd hit "request changes" on this blog post- LLM assistance is as irrelevant to that as it is in a PR.
I'll commit to the bit: here's a PR review on your blog post. Just my humble opinion and I'm no writer myself, so feel free to disregard all of it. It was written by hand, minor rewrite according to AI critique :)
---
> Instructed to answer in cablese — the telegraph operators' compressed dialect, models elicit 40–49% fewer billed output tokens on the API's meter, and models across four families still recover the information at full fidelity.
suggestion: that's a run on sentence, with a "—" signifying a new sentence, but this second sentence doesn't have a clear flow. Try vocalising this sentence, where are your breathing pauses? I can't vocalise it clearly. It's a very minor thing! Try vocalising this minor rewrite: "Instructed to answer in cablese — the telegraph operators' compressed dialect — models elicit 40-49% fewer billed output tokens. Models across four families still recover the information at full fidelity."
That's much easier to vocalise. Usually, that makes it easier to read, too.
> The Telegraph Test benchmark measures how well models compress using this technique as well as how much information they can retreive from it afterward.
nit: reword. "how well models compress" is a bit awkward here. something like "The Telegraph Benchmark measures the model's ability to compress, as well as...". To me "Telegraph Test benchmark" is a confusing term, "Telegraph Benchmark" is clearer, and still terse(er!)
praise: otherwise this is a good abstract-like introduction.
nit: You did clearly edit this by hand, because you misspelled "retreive".
> Cross-family matrix (readers = foreign models answering from GLM-5.3-Flash’s records; writers = GLM-5.3-Flash answering from theirs) :
suggestion: add a paragraph. This is not a good introductionary text, what kind of questions did you test on, what's the goal here?
> GLM-5.3-Flash itself: 48.4% savings with the lowercase instruction, in-family recovery 1.09. No comparison in the matrix favors plaintext; every ratio sits at 0.99–1.10.
question: why start with "<one of the models> itself?" It's unclear why you're talking about that one in particular here, and makes it hard to follow your point.
> The condition ladder — same questions, one variable at a time:
question: what is a condition ladder?
> The register is not a construct we invented (LLMs were handed compressed records cold and read them at parity); the capability was already in the weights, inherited from a century and a half of people writing under metered bandwidth.
suggestion: rephrase. What's "the register", as in the tone the LLMs speak in? explain your terms.
> Every model tested can do this.
suggestion: rephrase, that's not a grammatically correct sentence. e.g. "all models we tested can do this", or "every tested model can do this" if you wish terseness.
One such sentence is obviously not a problem at all! but too many, and you'll lose readers.
> Notice what this adds up to: Result:
praise: good centerpiece, that's your central thesis
> No new hardware, no training, no API change, one sentence of instruction.
nit: very AI coded language, "no <x>, no <y>, ..." is cliché. Not a problem obviously, but I thought I'd bring your attention to it.
---
That's sorta where I lost interest in this PR review bit. the point here is that it's harder to read and engage your content. Also, if it looks too AI, you'll lose readers who assume you didn't put effort in.
> Instructed to answer in cablese — the telegraph operators' compressed dialect (drop the articles and filler, keep every fact) — a single one-sentence instruction, no examples and no codebook, elicits 40–49% fewer billed output tokens on the API's own meter, and models across four families still recover the information at full fidelity. For machine-to-machine traffic, that is half the output bill at any major API, today. The Victorian economics of the cable, reborn as token economics: the Telegraph Test benchmark.
I tried but this first paragraph seems like it was almost purposefully obfuscated. Typical LLM-written content that meanders around a bit and stops when it seems to have emitted enough words. The reader is left to assemble meaning from the trace of thought it did not go back over to revise.
I think you have a good point to make but the writing is really difficult to get past.
If you call your ideas "content", then there is no point in reading it anyway.
Wait, you wrote this manually?
I wrote it with the help of an LLM, then edited it, then wrote some more, then edited that. Took about a week to format. It contains my thoughts. Brave new world, I know. I'm finding it ironic that the same crowd that is merging AI code all day is so allergic to AI assistance in prose...do all your work using this new godlike technology, but when it comes time to share, whip out your fountain pen or else.
The problem is that the prose becomes difficult to read because of its lexical quirks and verbiage. One AI generated piece of prose in isolation, fine; but people have developed a 'smell' of it and a mental association with low-quality work.
It might contain your thoughts, but the writing definitely needs work.
With code, people often think just getting it to run correctly is enough (which is often wrong, but that's another discussion.) With writing, the success criteria are more complex.
Even just the first sentence is unnecessarily cryptic, with inconsistent grammar. For example, the parenthetical clause, "the telegraph operators' compressed dialect," is surrounded by an em-dash and a comma. It should be the same character on each side.
One issue you may be having is forgetting that readers don't have all the context about your task that you have in your mind. Writing maximally concisely, which seems to be what you or the model were going for, means that readers have to try to reconstruct that context. Presumably you weren't trying to write the article itself in telegraphic style.
Here's a possible rewrite of that first sentence:
> When LLMs are instructed to answer in "cablese"—the compressed dialect historically used by telegraph operators—models produce 40-49% fewer billable output tokens through the API. Despite this compression, models from all four tested model families still recover the information at full fidelity.
The original "Instructed to answer" formulation is known as a compressed participial construction. It's grammatically valid, but overly concise for the first sentence of an article. Being explicit about what's being instructed helps the reader avoid having to read the whole sentence before determining what its subject is.
The ", and" pattern is something I often catch myself using, but it's often better to break up the sentence instead, and that's certainly true in this case. It also makes it easier to call out the result you're describing more explicitly, with "Despite this compression".
The "elicit ... on the API's meter" is just a very weird construction, and does sound quite Claudish. I doubt "elicit" is the correct word here. A prompt elicits a response from a model, but the model itself doesn't "elicit" anything when responding to a prompt.
If you want to improve your writing, I recommend reading. Not just non-fiction - reading fiction can be a great way to improve your writing skills. It trains your brain in what's correct, essentially - assuming what you're reading is well-written!
I think AI assistance in writing is fine for what it's worth. Perhaps you'd like https://sockpuppet.org/blog/2026/09/17/how-to-write-with-an-... or https://www.seangoedecke.com/how-i-use-llms/#proofreading-fo...?
However, I'm as allergic to slop in my code as in my reading. I'd hit "request changes" on this blog post- LLM assistance is as irrelevant to that as it is in a PR.
I'll commit to the bit: here's a PR review on your blog post. Just my humble opinion and I'm no writer myself, so feel free to disregard all of it. It was written by hand, minor rewrite according to AI critique :)
---
> Instructed to answer in cablese — the telegraph operators' compressed dialect, models elicit 40–49% fewer billed output tokens on the API's meter, and models across four families still recover the information at full fidelity.
suggestion: that's a run on sentence, with a "—" signifying a new sentence, but this second sentence doesn't have a clear flow. Try vocalising this sentence, where are your breathing pauses? I can't vocalise it clearly. It's a very minor thing! Try vocalising this minor rewrite: "Instructed to answer in cablese — the telegraph operators' compressed dialect — models elicit 40-49% fewer billed output tokens. Models across four families still recover the information at full fidelity."
That's much easier to vocalise. Usually, that makes it easier to read, too.
> The Telegraph Test benchmark measures how well models compress using this technique as well as how much information they can retreive from it afterward.
nit: reword. "how well models compress" is a bit awkward here. something like "The Telegraph Benchmark measures the model's ability to compress, as well as...". To me "Telegraph Test benchmark" is a confusing term, "Telegraph Benchmark" is clearer, and still terse(er!)
praise: otherwise this is a good abstract-like introduction.
nit: You did clearly edit this by hand, because you misspelled "retreive".
> Cross-family matrix (readers = foreign models answering from GLM-5.3-Flash’s records; writers = GLM-5.3-Flash answering from theirs) :
suggestion: add a paragraph. This is not a good introductionary text, what kind of questions did you test on, what's the goal here?
> GLM-5.3-Flash itself: 48.4% savings with the lowercase instruction, in-family recovery 1.09. No comparison in the matrix favors plaintext; every ratio sits at 0.99–1.10.
question: why start with "<one of the models> itself?" It's unclear why you're talking about that one in particular here, and makes it hard to follow your point.
> The condition ladder — same questions, one variable at a time:
question: what is a condition ladder?
> The register is not a construct we invented (LLMs were handed compressed records cold and read them at parity); the capability was already in the weights, inherited from a century and a half of people writing under metered bandwidth.
suggestion: rephrase. What's "the register", as in the tone the LLMs speak in? explain your terms.
> Every model tested can do this.
suggestion: rephrase, that's not a grammatically correct sentence. e.g. "all models we tested can do this", or "every tested model can do this" if you wish terseness.
One such sentence is obviously not a problem at all! but too many, and you'll lose readers.
> Notice what this adds up to: Result:
praise: good centerpiece, that's your central thesis
> No new hardware, no training, no API change, one sentence of instruction.
nit: very AI coded language, "no <x>, no <y>, ..." is cliché. Not a problem obviously, but I thought I'd bring your attention to it.
---
That's sorta where I lost interest in this PR review bit. the point here is that it's harder to read and engage your content. Also, if it looks too AI, you'll lose readers who assume you didn't put effort in.
Why not write your article in the same telegraphese you preach? Hilariously ironic to use verbose AI writing for this.
Also the site background is AI slop which makes for terrible contrast with the text.
I use the tools I study, nothing more or less.
Since when is writing blog posts themselves part of "studying"?
2 replies →
> Instructed to answer in cablese — the telegraph operators' compressed dialect (drop the articles and filler, keep every fact) — a single one-sentence instruction, no examples and no codebook, elicits 40–49% fewer billed output tokens on the API's own meter, and models across four families still recover the information at full fidelity. For machine-to-machine traffic, that is half the output bill at any major API, today. The Victorian economics of the cable, reborn as token economics: the Telegraph Test benchmark.
I tried but this first paragraph seems like it was almost purposefully obfuscated. Typical LLM-written content that meanders around a bit and stops when it seems to have emitted enough words. The reader is left to assemble meaning from the trace of thought it did not go back over to revise.
I think you have a good point to make but the writing is really difficult to get past.