Comment by cnity
1 day ago
I've noticed that most people seem to consider the core problem of Claude's output as "too verbose" but I don't think this actually cuts to the heart of the matter at all. It's almost, in some weird way, the opposite: like the text is far too _dense_. It tries too hard to invent odd terminology to try to condense stuff, but it doesn't tell you up front that it is going to call your company wide error-handling mechanism a "flare" (or some other such strange term).
Exactly, it’s absurdly dense, it’s almost impossible to follow. An it always omits the subject of each sentence.
Ngl I think this is partially an artifact of it having a better grasp of English than almost everyone
Frequently its choice of a particular word is perfect and gives me the vocabulary to talk about the task at hand the way I want
Like it’s tuned to just be “maximally dense” instead of “dense/technical where you can handle it and simple where you can’t”
It doesn’t know where your language strengths/weaknesses are, so it can’t communicate to you like a fellow human does.
Human explaining something: are you familiar with phlox gabrania? no? let me give you some background first
Claude explaining something: gedarkin load bearing phlox gabrania seam. Also, you didn't ask about cheesecake but let me tell you about phlox gabrania cheesecake woles.
Yeah I don't the problem is verbosity as such, as I frequently have to ask to explain how it reached a certain conclusion and in particular what the empirical evidence for it is, at which point it too frequently reconsiders its answer.
It's just that the details it parrots are often irrelevant and wrapped in a way that makes them seem relevant.
my hypothesis is that its trying to hide the thinking process so people can't train models on the output, try to learn anything complex using AI, its basically imposible, its like its actively fighting giving you the main rationale
Kind of both. On the one hand, it is “verbose” in the sense that it will tell me every little nit that it can think of while doing a task, it will tell me a narrative about its thought process, and it will tell me every other detail it can think of. But it does so in a way that tries to be incredibly dense to the point that I have to struggle to figure out what it is saying. I wonder if there are any “legibility benchmarks” that one could use to determine what prompts work best?
I find it to be both as well, as in "packed full of information, but most of it is worthless". Sentences so dense I have to read them three times, assembled into a five paragraph essay of "honest caveats" and "things worth knowing" in response to the simplest yes-or-no questions.
I quite like https://surgehq.ai/benchmarks/hemingway-bench
I wish this had non-model comparisons. If Opus 5 is in the top ten, it’s clear that the entire benchmark is somewhere between “Tom Clancy” and “Dan Brown” and about 1,000 new model releases away from Hemingway.
When you see, “Wow, Fable is number one”, you might think it’s a good writer, but that’s not what the benchmark says.
Seems to me a bit insensitive or logarithmic. Fable is way worse than some of the others in this list, but only 10-20% higher score.
There are no "best" prompts. Its a random BS generation machine that you can at times direct enough to get stuff done for you. The output will almost always have varying levels of BS that you have to clean up with various levels of effort.
"<Country> doesn't have a largest city because all of the cities in <Country> are small"
yes, but how else would you know that "flare" was the load bearing part of that statement? /s
I'm going to argue to my boss that our KPI for the next quarter should be the number of load bearing seams discovered. I'll await the promotion.