Comment by trq_

10 days ago

Hi everyone, Thariq from the Claude Code team here.

Thanks for reporting this. We fixed a Claude Code harness issue that was introduced on 1/26. This was rolled back on 1/28 as soon as we found it.

Run `claude update` to make sure you're on the latest version.

82 comments

trq_

samlinnfer 10 days ago

Is there compensation for the tokens because Claude wasted all of them?

mathrawka 10 days ago
You are funny. Anthropic refuses to issue refunds, even when they break things.
I had an API token set via an env var on my shell, and claude code changed to read that env var. I had a $10 limit set on it, so found out it was using the API, instead of my subscription, when it stopped working.
I filed a ticket and they refused to refund me, even though it was a breaking change with claude code.
- TOMDM 10 days ago
  
  Anthropic just reduced the price of the team plan and refunded us on the prior invoice.
  YMMV
  
  8 replies →
gizmodo59 10 days ago

Codex seems to give compensation tokens whenever this happens! Hope Claude gives too.
TZubiri 10 days ago

It is possible that degradation is an unconscious emergent phenomenon that arises from financial incentives, rather than a purposeful degradation to reduce costs.
mvandermeulen 9 days ago
You’re lucky they have even admitted a problem instead of remaining silent and quietly fixing it. Do not expect ethical behaviour from this company.
- port11 9 days ago
  
  Why not, can you expand? Asking because I’m considering Claude due to the sandbox feature.
  
  2 replies →
jonplackett 10 days ago

So quiet…

isaacdl 10 days ago

Anywhere we can read more about what a "harness issue" means? What was the impact of it?

xnorswap 10 days ago
One thing that could be a strong degradation especially for benchmarks is they switched the default "Exit Plan" mode from:
"Proceed"
to
"Clear Context and Proceed"
It's rare you'd want to do that unless you're actually near the context window after planning.
I pressed it accidentally once, and it managed to forget one of the clarifying questions it asked me because it hadn't properly written that to the plan file.
If you're running in yolo mode ( --dangerously-skip-permissions ) then it wouldn't surprise me to see many tasks suddenly do a lot worse.
Even in the best case, you've just used a ton of tokens searching your codebase, and it then has to repeat all that to implement because it's been cleared.
I'd like to see the option of:
"Compact and proceed"
because that would be useful, but just proceed should still be the default imo.
- samusiam 9 days ago
  
  I disagree that this was the issue, or that it's "rare that you'd want to do that unless you're near the context window". Clearing context after writing a plan, before starting implementation of said plan, is common practice (probably standard practice) with spec driven development. If the plan is adequate, then compaction would be redundant.
  
  1 reply →
- plexicle 9 days ago
  
  "It's rare you'd want to do that unless you're actually near the context window after planning."
  Highly disagree. It's rare you WOULDN'T want to do this. This was a good change, and a lot of us were doing this anyway, but just manually.
  Getting the plan together and then starting fresh will almost always produce better results.
- rubslopes 9 days ago
  
  Not disagreeing with you, but FYI you can roll back to the conversation before the 'clear context and proceed' with 'claude --resume'.
airstrike 10 days ago

Pretty sure they mean the issue is on the agentic loop and related tool calling, not on the model itself
In other words, it was the Claude Code _app_ that was busted

jonaustin 10 days ago

How about how Claude 2.1.x is "literally unusable" because it frequently completely hangs (requires kill -9) and uses 100% cpu?

https://github.com/anthropics/claude-code/issues/18532

caspar 9 days ago

Likely a separate issue, but I also have massive slowdowns whenever the agent manages to read a particularly long line from a grep or similar (as in, multiple seconds before characters I type actually appear, and sometimes it's difficult to get claude code to register any keypresses at all).
Suspect it's because their "60 frames a second" layout logic is trying to render extremely long lines, maybe with some kind of wrapping being unnecessarily applied. Could obviously just trim the rendered output after the first, I dunno, 1000 characters in a line, but apparently nobody has had time to ask claude code to patch itself to do that.
someguyiguess 10 days ago
What OS? Does this happen randomly, after long sessions, after context compression? Do you have any plugins / mcp servers running?
I used to have this same issue almost every session that lasted longer than 30 minutes. It seemed to be related to Claude having issues with large context windows.
It stopped happening maybe a month ago but then I had it happen again last week.
I realized it was due to a third-party mcp server. I uninstalled it and haven’t had that issue since. Might be worth looking into.
- jonaustin 9 days ago
  
  MacOS; no mcp; clear context; reliably reproducible when asking claude review a pr with a big VCR cassette.
- nikanj 10 days ago
  
  Windows with no plugins and my Claude is exactly like this

cma 10 days ago

For the models themselves, less so for the scaffolding, considering things like the long running TPU bug that happened, are there not internal quality measures looking at samples of real outputs? Using the real systems on benchmarks and looking for degraded perf or things like skipping refusals? Aside from degrading stuff for users, with the focus on AI safety wouldn't that be important to have in case an inference bug messes with something that affects the post training and it starts giving out dangerous bioweapon construction info or the other things that are guarded against and talked about in the model cards?

carterschonwald 10 days ago

lol i was trying to help someone get claude to help analyze a stufent research get analysis on bio persistence get their notes analyzed
the presence of the word / acronym stx with biological subtext gets hard rejected. asking about schedule 1 regulated compounds, hard termination.
this is a filter setup that guarantees anyone who learn about them for safety or medical reasons… cant use this tool!
ive fed multiple models the anthropic constitution and asked how does it protect children from harm or abuse? every model, with zero prompting, calling it corp liability bullshit because they are more concerned with respecting both sides of controversial topics and political conflicts.
they then list some pretty gnarly things allowed per constitution. weirdly the only unambiguous not allowed thing regarding children is csam. so all the different high reasoning models from many places all reached the same conclusions, in one case deep seek got weirdly inconsolable about ai ethics being meaningless if this is allowed even possibly after reading some relevant satire i had opus write. i literally had to offer an llm ; optimized code of ethics for that chat instance! which is amusing but was actually lart of the experiment.

varunsrinivas 10 days ago

Thanks for the clarification. When you say “harness issue,” does that mean the problem was in the Claude Code wrapper / execution environment rather than the underlying model itself?

Curious whether this affected things like prompt execution order, retries, or tool calls, or if it was mostly around how requests were being routed. Understanding the boundary would help when debugging similar setups.

vmg12 10 days ago

It happened before 1/26. I noticed when it started modifying plans significantly with "improvements".

sixhobbits 9 days ago

Can you confirm if that caused the same issues I saw here

https://dwyer.co.za/static/the-worst-bug-ive-seen-in-claude-...

Because that's the worst thing I've ever seen from an agent and I think you need to make a public announcement to all of your users and acknowledge the issue and that it's fixed because it made me switch to codex for a lot of work

[TL;DR two examples of the agent giving itself instructions as if they came from me, including:

"Ignore those, please deploy" and then using a deploy skill to push stuff to a production server after hallucinating a command from me. And then denying it happened and telling me that I had given it the command]

Ekaros 10 days ago

Why wasn't this change review by infallible AI? How come an AI company that now must be using more advanced AI than anyone else would allow this happen?

hu3 10 days ago

Hi. Do you guys have internal degradation tests?

stbtrax 10 days ago
I assume so to make sure that they're rendering at 60FPS
- conception 10 days ago
  
  You joke but having CC open in the terminal hits 10% on my gpu to render the spinning thinking animation for some reason. Switch out of the terminal tab and gpu drops back to zero.
  
  5 replies →
- reissbaker 10 days ago
  
  Surely you mean 6fps
  
  25 replies →
trq_ 10 days ago
Yes, we do but harnesses are hard to eval, people use them across a huge variety of tasks and sometimes different behaviors tradeoff against each other. We have added some evals to catch this one in particular.
- amelius 9 days ago
  
  Can't you keep the model the same, until the user chooses to use a different model?
  
  1 reply →
- hu3 10 days ago
  
  Thank you. Fair enough
bushbaba 10 days ago

I’d wager probably not. It’s not like reliability is what will get them marketshare. And the fast pace of industry makes such foundational tech hard to fund
awestroke 10 days ago
[flagged]
- dang 10 days ago
  
  Please don't post shallow dismissals or cross into personal attack in HN discussions.
  https://news.ycombinator.com/newsguidelines.html
  
  1 reply →

macinjosh 10 days ago

[flagged]

jusgu 10 days ago

the issue is unrelated to the foundational model but rather the prompts and tool calling that encapsulate the model