Comment by vidarh

2 days ago

I have a document generation task that I used to run with 4.8. This morning after it switched to 5, the documents were consistently 30%-40% longer for the same prompt... Not evaluated whether they are actually better or worse yet, but what was interesting was how consistently more verbose it was.

Review fatigue is already draining my energy levels at work to a point where I've been using agents less because it's just not possible to be productive when I have to read hundreds of lines of markdown every couple minutes.

Other people might just turn to automation blindness and click OK without verifying but I refuse to just let anthropic go rampant in my codebase.

I'm sure that's entirely unrelated to the fact that they charge per token.

  • Heh. For this specific task, if the output is the same quality or better per token output, it'd actually be positive, but I have plenty of tasks where extra verbosity would be undesirable as well, so I'll be keeping a close eye on this.

    Currently has Opus running a comparative test of itself against Kimi K3 on one project, and it keeps finding that Kimi is competitive for multiple stages of the pipeline, so I may well end up reducing my use overall anyway...