← Back to context

Comment by Aurornis

5 hours ago

You should read this person's full article to understand what these charts are showing https://x.com/Lon/status/2101034933284417614

If you thought this was a repeated test of the same problems showing fluctuating performance, it's not. They set up a MITM proxy between Claude and the servers and ran analysis on the work they were doing.

So those ups and downs in the charts, which they plotted with sub-daily resolution, are just as much a function of their work changing from day to day. It's like plotting the miles per gallon of your car and blaming the gas station when the number goes up and down, without admitting that some days you drive to the grocery store on surface roads and other days you drive up a mountain on the freeway.

> The corpus analyzed in Charts 1-5 comes exclusively from Fable 5, at xhigh and max effort levels, during sustained production work across a diverse set of projects and workloads. Data was aggregated from transcripts and live wire logs

The analysis (which feels very vibe-slop) gets worse from there. In the second half they take thinking token counts for ARC-AGI-2, thinking problems designed to stress LLMs, and compare their average thinking-tokens-per-turn counts to that!

If you don't realize why this is so flawed: ARC-AGI-2 is a benchmark meant to collect problems thought to be extremely difficult, nearly impossible, for LLMs. If your goal was to cherry-pick a mislead example which would produce the highest number of thinking tokens, this is it!

Your daily coding work should not be producing a proportional number of thinking tokens on every invocation while it reads through some source code or edits a couple lines in a file.

You don't want to maximize the number of thinking tokens. You want problems solved accurately with the minimum number of tokens.

Confirmation bias runs deep on this topic so I assume few people read the analysis before posting, but as far as experiments go it's basically useless. Are they changing something on the server? I don't know, but this analysis isn't useful for answering that question.

Thank you for your kind words, Aurornis.

Yes, this is my production corpus, across 65 usage days, two subscription accounts, three machines, 25 project groups, and 213 sessions, across 43,261 invocations and 7,583 turns. Use your own data if you want to prove or refute what was seen in my corpus.

The "ups and downs in the chart" were not plotted with sub-daily resolution. Specifically, the two-month temporal chart uses a 3.5-day Gaussian bandwidth which is meant to reduce short-term noise while retaining broader changes.

Additionally, a separate episodic analysis identified multi-day changes in delivered thinking. And those episodes were predictive of held-out work. The more projects pulled into an ensemble, the more predictive they were of delivered thinking tokens for held-out projects during the episode.

The point of the benchmarks is to establish a baseline for what thinking-token counts one should expect from specific effort levels using published numbers, not what every response should be delivered. If you looked closer, you would see that P90 invocations were still delivered 13x thinking tokens below that level.

So you don't have to provide a generous interpretation of my workload if you don't want to. Remove all of the zero-thinking token responses, redistribute those samples across the distribution, and then tell me if it magically shifts right and starts delivering anything close to published numbers. Only -46- out of 36,374 July and August invocations even broke 16k thinking tokens - only 0.13% of the total. You are welcome to present what percentile you think is a fair comparison to make here.

If you would like to denigrate a month of my time as vibe-slop, that is your prerogative. You can even be dismissive of my workload, if you want, even if my background should tell you otherwise. But if you want to knock a month of someone's time, do it with your own data to at least help move the conversation forward.

  • > Use your own data if you want to prove or refute what was seen in my corpus.

    I don't think you understand. What you posted is highly dependent on your corpus. I can't "refute" anything because it's not available and it's the major variable in the experiment.

    > The point of the benchmarks to establish a baseline for what thinking-token counts one should expect from specific effort levels using published, not what every response should be delivered. If you looked closer, you would see that P90 invocations were still delivered 13x thinking tokens below that level.

    I think you're missing something from that first sentence, but I assume you're talking about the comparison to ARC-AGI-2 published thinking tokens?

    It should be blindingly obvious that you do not want your thinking token counts to be as high as a benchmark that was designed to push LLMs to their limit.

    > Only -46- out of 36,374 July and August invocations even broke 16k thinking tokens - only 0.13% of the total. You are welcome to present what percentile you think is a fair comparison to make here.

    What point are you even trying to make?

    Again, you don't want invocations to be burning 16K thinking tokens except for rare problems that 1) must be solved in one step and 2) are designed to be entirely self-contained thinking in that step.

    You're trying to compare development work to a benchmark that encapsulates complex thinking into a single step.

    Coding work is iterative and works in incremental steps: It runs commands, reads more files, checks the web. Thinking tokens should be low for your turns.

    ARC-AGI problems have an input and an output. They look like this: https://arcprize.org/tasks/b5ca7ac4 They have more thinking tokens because that's the entire state. They get one output and it's constrained.

    > But if you want to knock a month of someone's time, do it with your own data to at least help move the conversation forward.

    It is fair to discuss a published analysis. Saying that only people who bring their own month of equivalent analysis (which conveniently would take another month to produce) are allowed to critique it is just a cheap trick to shut people down.

    If you post big claims, they are open to analysis and review by others

    • After reading what you had to add to this discussion, I have only two things to leave you with:

      1/ if the point of setting xhigh or max effort on a "reasoning model" is -not- to have additional reasoning tokens applied to the problem, then I fear we are all using AI wrong 2/ the decorum with which you "review" someone's work in public is entirely your choice. the only "cheap trick" on the internet is being a pseudonymous ass.

      1 reply →