Comment by throwaway894345

1 day ago

Kind of both. On the one hand, it is “verbose” in the sense that it will tell me every little nit that it can think of while doing a task, it will tell me a narrative about its thought process, and it will tell me every other detail it can think of. But it does so in a way that tries to be incredibly dense to the point that I have to struggle to figure out what it is saying. I wonder if there are any “legibility benchmarks” that one could use to determine what prompts work best?

I find it to be both as well, as in "packed full of information, but most of it is worthless". Sentences so dense I have to read them three times, assembled into a five paragraph essay of "honest caveats" and "things worth knowing" in response to the simplest yes-or-no questions.

I quite like https://surgehq.ai/benchmarks/hemingway-bench

  • I wish this had non-model comparisons. If Opus 5 is in the top ten, it’s clear that the entire benchmark is somewhere between “Tom Clancy” and “Dan Brown” and about 1,000 new model releases away from Hemingway.

    When you see, “Wow, Fable is number one”, you might think it’s a good writer, but that’s not what the benchmark says.

  • Seems to me a bit insensitive or logarithmic. Fable is way worse than some of the others in this list, but only 10-20% higher score.

There are no "best" prompts. Its a random BS generation machine that you can at times direct enough to get stuff done for you. The output will almost always have varying levels of BS that you have to clean up with various levels of effort.