← Back to context

Comment by mpoteat

4 days ago

Sorry folks, this is a rollout artifact, we needed a way to turn this off remotely via feature flags if it broke something, and with telemetry off you don't get those. It's already been fixed as part of v2.1.281 releasing today.

The mod is source available here: https://github.com/anthropics/claude-code/tree/main/mods/age...

Apologies again folks, this was a fully human error on my part - I should've found a better way to launch with a kill-switch.

The AGENTS.md support was implemented via our new extensibility system for CC, called Mods, which is launching soon-ish. A mod is a plugin with a new type of hook, which we call a function hook.

If folks play around with it, I would love feedback on the relevant issue: https://github.com/anthropics/claude-code/issues/91870

Mods allow quite a bit more customizability and control. I really believe in the idea.

  • I skimmed a couple pages of the docs at:

    https://github.com/user-attachments/files/31802150/EXTERNAL....

    Might I gently suggest that you have a model at least as capable as Opus 5.5 translate that from Claudish to English? Or, even better, have an actual human work on the docs a bit? As it stands, they are fairly egregious, and they seem to devote at least as much space to little AI-generated quips that convey no meaning than to actually explaining what’s going on.

    Also, maybe a human should decide whether these are “function” hooks or “module” hooks. All of this marketing calls them “function” hooks, but the json config seems entirely unaware of this.

    (Has anyone else noticed that half the sentences in Claudish aren’t merely weird: they are noun phrases and not sentences at all? I’m pretty sure that any decent pre-LLM NLP-based grammar checker would correctly flag half the sentences in Claudish. Also, whatever variant of Claude wrote this thing can’t even capitalize around semicolons consistently with itself, let alone consistently with how English has been written for at least a century.)

    edit: Fixed the link. Thanks, kaszanka.

    • Ouch! Much of this verbiage was dictated by me personally; I just have a fairly distinct register some might consider inscrutable. My English teachers in grade school always said the same :)

      Rest assured I'll inject a bit less soul into the official docs once Mods are launched; re your feedback on the JSON key, what would you recommend?

      23 replies →

    • I’m curious… would you talk to them like this in person? The only ‘egregious’ thing I read here was your reaction. Maybe you could consider acting more human?

    • Writing documentation is one of the most often mentioned uses of LLMs. I suppose if Anthropic wouldn't be doing it it would put into question why anyone else would.

    • > they are noun phrases and not sentences at all? I’m pretty sure that any decent pre-LLM NLP-based grammar checker would correctly flag half the sentences in Claudish. Also, whatever variant of Claude wrote this thing can’t even capitalize around semicolons consistently with itself, let alone consistently with how English has been written for at least a century.

      Probably just a case of a company hoping their scale can change the societal standard faster than they can be bothered to match the standard.

      You'll talk like 2023 unsupervised TikTok generators and you'll be happy.

  • The only difference between CLAUDE.md and AGENTS.md is the filename. Do you really need a whole plugin system to support that use case? Feels like this could have been a one line change.

    • Don't you run bizantine ralph loops on remote environments with codex security checks, coupled with jev, grok, open router and a fully independent openclaw (on a maxed out mac mini inside a caveau in an undisclosed location, with open telegram) to change constants? You're going to be left behind.

      11 replies →

    • The plugin system itself was probably already in the making, and they just chose to implement this tiny feature as a plugin to try it out.

      As for the only difference being the file name, that's an untested assumption. Up until now, Claude hadn't supported AGENTS.md, and it's a simple application of Hyrum's Law that somebody, somewhere, was taking advantage of that to give one set of instructions to Claude, and a different set of instructions to some other provider. The correct behaviour in the presence of both files is not obvious, either.

      Having a kill switch for the change is a perfectly reasonable safety measure in case something goes horribly wrong.

      36 replies →

    • Yes, but at the same time, it’s also a good, simple use case to test a new plugin system. I can totally understand that.

    • At $FAANG, ~all changes go through feature flags.

      There are processes to make changes outside of feature flags, but they have enough friction that it's easier to just use a feature flag.

      This level of paranoia is consistent with the blast radius of changes breaking Claude users.

      3 replies →

    • I'm not sure why they didn't just extend plugins, but having mod support is definitely a plus for everyone.

      AGENTS.md seems to simply showcase what mods are capable of.

    • We've just had a file called `.rules` that is symlink as `AGENTS.md`, `CLAUDE.md`, etc to support the 3 or so common ones used around our projects.

    • With apologies for not just testing this myself (currently AFK), doesn't it still work to have a CLAUDE.md file containing just `@AGENTS.md`?

    • Go download the leaked source code from earlier this year, search for all occurrences of the string `CLAUDE.md`, and be horrified.

    • One line change? Pffft. That means you're still looking at the code, you're behind the times.

    • > The only difference between CLAUDE.md and AGENTS.md is the filename. Do you really need a whole plugin system to support that use case? Feels like this could have been a one line change.

      I don't think this is a reasonable assumption. The document format in CLAUDE.md is whatever Anthropic specifies, where AGENTS.md is a common ground format that is expected to be supported by any agent, be it from Anthropic or not.

      https://agents.md/

      You might argue that differences are small or negligible, but that is just an expectation.

      6 replies →

  • >via our new extensibility system for CC, called Mods, which is launching soon-ish. A mod is a plugin with a new type of hook, which we call a function hook.

    An extensibility system called mods, which is a plugin with a new type of hook that we call function hook?

    I can't tell if this is real, or you are making fun of overengineered AI solutions.

    Is this real?

  • I've developed several plugins for different harnesses, and I needed some upstream change for most of them.

    The deciding factor for me whether or not I will work on the feature of the plugin is whether I (or rather, my agent) can look in upstream source and evaluate if it can be done with minimal upstream change, which I then contribute. And generally, even if no upstream change is needed, agents work so much better when they can read the code.

    So why not just make Claude code open source? Considering also that source code was leaked once anyway.

  • > The AGENTS.md support was implemented via our new extensibility system for CC, called Mods,

    AKA "we need 100~ish files wrtitten in the most horrible Clean Code style replete with no two files agreeing on the same naming of the same feature... to read one of two files, one of which has been a de-facto industry standard for over two years"

  • Will you extend your plugin to read skills and rules from `.agents`? Or should we write our own plugin/mod for that?

  • This may or may not be related, but if you're working on CC, who do we have to annoy to make Anthropic stop trying to force use of arbitrary Bash commands instead of the actual tool calls built into the harness? (https://github.com/anthropics/claude-code/issues/90450, https://github.com/anthropics/claude-code/issues/89251, etc) It's deeply infuriating at times that there's this full system of hooks, permissions, etc that's unusable at times because CC keeps trying to make the model not use any of it.

So what Anthropic calls "telemetry" is really "telechangeability"? That's sits with me even less comfortably than did the idea that some features are only available with telemetry enabled.

  • I’d just assume good intent here. Feature flags, telemetry, and fast rollouts / rollbacks are standard practice in software. Have a look at chrome://flags perhaps.

    I fully believe GP that there was zero intent to gate this behind collecting telemetry. Sounds like a little tech debt and a little oversight, and the simplest explanation is that it is.

    • You can have flags work in a pull fashion, though. At boot, and every 5 minutes, query an endpoint which returns flag values. You don't need the client to send any information about what the user is doing.

      2 replies →

  • That's just how big companies roll out software changes for software that auto-updates.

    It's much preferable to be able to instantly fix it if the rollout of a new feature goes wrong than have everyone who installed the broken version bring stuck with problems until the company realizes the issue and rolls forwards with a fixed version.

    https://martinfowler.com/articles/feature-toggles.html

    • I've used feature flags extensively throughout my career as a SaaS developer, but I've never considered their usage in desktop software. Not sure how I feel about it.

      1 reply →

  • Not necessarily. They probably have an integrated service that handles some telemetry and also feature flags, like Braze. They toggled the whole thing off based on telemetry settings. It's a bug I've made before too.

  • Seems like they overloaded whatever they use for telemetry to do feature flags.

    It’s not a crazy conspiracy. They messed up, it’s fine.

Ouch.

I've had to send such messages, but internally at work, not on HN!

Have a great day, human.

Really loved this Tibo-level responsiveness, if Anthropic can keep it up with this level of service, I am pretty sure a lot of people will just ditch their ChatGPT subscription and just move to Claude.

  • Why would I ever do that to myself? My experience with Codex/GPT is fantastic, while my impression of Claude/Opus is that it's longwinded, patronizing, token-inefficient, stops to ask stupid questions every other minute, overcomplicates simple tasks, often poor engineering overall. I don't use it but this is what I see my partner run into who has access to both and compares them often. She has the same assessment.

    • Cuz OpenAI has been secretly downgrading models on many accounts, including mine lately. I paid $200 a month since like gpt-5.4, and since Astra released I found the model is somehow acting strange, it is until I checked X I have discovered that OAI is giving Luna level models when I am requesting Sol/Astra, or some piece of s** that is even worse than Luna. I basically had to ran every session with a Pelican test to determine if that session is safe. So I just spun up my Claude $20 and figured that now I can get all the work done just with Opus 5. Let me show you a pelican, by "gpt-6-sol". Cutting usages is one thing, but secretly downgrading models to a level that is not reliable anymore is the last straw. I am not saying other frontier labs (I am talking about you Anthropic) isn't doing this, but their version of downgraded/quantized/reduced effort model is at least usable, probably just slightly dumber, OAI's differences is day and night. https://imgur.com/a/PDbYdOQ

      4 replies →

  • On the other hand, if Anthropic is to follow the industry standards, this would never have happened in the first place. It's not like the feature gates are the frontier of software development.

    • There are two or three relevant companies in this space in America and this is the one of them that kicked off the whole terminal agent harness thing in getting market adoption. It's perfectly fine for neither of these companies to follow industry standards while they're figuring shit out

  • What we need is a low level but constant drumbeat against openai in general. In general the AI situation is overleveraged and underpoliced, with the occasional hints of AI gone wild. If openai were to just be left to die, we could let that financial mess unroll and bail out the leftovers, I don't like bailouts anymore than the next guy but with this administration its almost a guarantee if things go south because this adminstration can charge administrative fees of maybe $20-30 billion (which goes to trump), get Sam Altman to serve one or two years in a cushy resort type fed place for the hugging face hacking and put openai's processes on github as a premium feature, say $10000 a month to access (which again goes to trump).

    I know I know, why are we giving money to trump? Its because he's going to take it anyways so can't we at least apply some window dressing?

  • Sure a fast response on HN would make people switch. Try better rates, infra, limits etc.

  • Do yourself a favor: ditch both and go local.

    • Local is becoming ever increasingly scarce and cost prohibitive. It's kind of bleak out there right now. A minimum bar to entry for decent local AI (something that can run a 27B tier model with some reasonable context) is going to set you back a year or five worth of AI API token costs.

      1 reply →

  • Because they respond to HN threads about their products? Which are likely Claude hooks monitoring for activity in the first place? Come on...

    At least make an argument for switching vendors based on the quality or price of their service.

    • After I was mildly disappointed with GPT-6 Sol and Luna not improving intelligence and only cutting the price, I'm running Opus 5.5 today.

      After the last month or so in the Codex app, I was pleased with the Claude app.

      It might be a case of the grass always being greener on the other side, but this is what stands out:

      After 3-4 hours of usage, the weekly usage limit moved by only 1%.

      Compared to Astra where I can watch the limit draining live, this is a great improvement.

      I'd estimate it 3x cheaper, and that's with a 450k context limit instead of the 258k in Codex.

      So far Opus 5.5 appears less prone to stopping for no apparent reason at checkpoints in the middle of a longer task.

      It doesn't open an internal browser with a useless comparison page, where it then proceeds to add notes despite no one having asked for it.

      It is a breath of fresh air: I get the response in the chat, while the Codex app recently loves randomly opening artifacts instead.

      Opus 5.5 xhigh made great progress on the task, more so than Astra High, but that could be random chance.

      Oh, and the 'Auto' mode actually works and does not force me to instead run 'Full access' like in the Codex app, lest it blocks even 'git push'.

      1 reply →

  • > ...if Anthropic can keep it up with this level of service...

    fuckin laughable, literally invoked a laugh from me in real life.

    I hope customers aren't so stupid that they think a chatty developer on twitter/hn/mastodon/screaming-in-the-wind/wherever (or any other public-facing-place) means shit about customer service, and that goes towards ANY company where the primary customer service is an LLM.

    Anthropic is the only company where it took (!) 9 weeks (!) to convince to hand over a 4 dollar refund for book-keeping errors on their side that caused an inappropriately early account deactivation due to time zone issues on their end, while all the while telling me that they don't offer refunds. It took stacks of evidence and argument, and that was after spending two weeks in their system trying to convince every level that I was worth a human.

    For me personally it'd require Dario to resort to armed mugging to see another buck out of my wallet. I'm not alone.

    tl;dr : being able to convince the powers that be on highly active industry forums (hacker news, twitter, mastodon..?) to act right using the power of peer shaming doesn't good customer service make. That said -- I do appreciate the direct response/statement from mpoteat;

    ..I just don't appreciate the good actions of a decent individual being too broadly interpreted as the do-good customer-centric nature of Anthropic .. an element I do not believe exists there.

Thanks for admitting to the mistake, but my understanding was that coding was fully solved now?

Hey, it seems you work there. I had a tangential question. Did any Anthropic exec threaten to do something unsavoury if any engineer ever tried to not name the claude cli binary as the version number itself? Because if they did, I'd understand. Or if you dare change it all the vibe-coded ts/react/etc dominoes will go for a fall in unison? I'd understand that too.

> a fully human error

Would be interesting to know how much time you/your team spent on that design decision

  • The correct design was in the AGENTS.md but they didn't have telemetry on

I can understand honest mistakes, but like, usually when I develop anything with agents (which I assume is what you're doing internally at Anthropic), they're almost too enthusiastic about trying to add test cases to the point where they sometimes try to glue together things in ways that are structurally impossible in the actual code in order to try to test that behavior. I'm honestly a bit mystified that adding a new feature didn't get bundled in with tests that the feature works for arbitrary configurations.

Shouldn't feature flags be independent of telemetry on/off? I would be real surprised if folks who have telemetry disabled would be upset if "required-for-software-to-function-normally" functionality like feature flagging itself was gated behind telemetry disablement.

The same guy writing readme's for my vibeslopped toy projects is also the readme writer at Anthropic, what a coincidence ;)

It was obvious this was the reason, it’s a very easy thing to forget about

is there reason I can't update claude to 2.1.281? I just run claude update

> claude update Current version: 2.1.280 Checking for updates to latest version... Claude Code is up to date (2.1.280)

How are you planning to turn this off remotely when telemetry off?

  • If it’s like other CC features that depend on telemetry being enabled, disabling telemetry will turn off the feature.

> Sorry folks, this is a rollout artifact

Aka: "an issue even a junior would've spotted if we didn't rely on Claude of 100% of our tasks"

telling people Anthropic remote-control their users computers isn't the smartest thing to do

  • Lol that's literally the entire point of the software, you install it just so the program can talk to a remote server to make changes on your local computer.

is it common to deploy software with a remote kill switch installed?

  • yes? feature flags have been a thing for a long time.

    • how strange.

      as i said in another comment. i don't want someone to toy with my software remotely. that seems wrong to me!

      i do not like others to decide that they know what is best for me. and then force it on me without my consent.

      i will decide if i like your changes. if i do like your fix, i will install it.

      in my car, do not remotely turn off my air conditioning. don't turn off my AGENTS.md.

      6 replies →

"Rollout artifact"? This is Claude-speak isn't it? I have never ever heard anyone call a bug like this a "rollout artifact" before.

  • I don't think "bug" is the correct term. They put a feature behind a feature flag, and feature flags don't work if you turn them off (via telemetry). That's "Working As Designed™".

    • It's clearly an unintended interaction. Nobody intended for the telemetry switch to control whether it reads AGENTS.md (hence mpoteat's apologetic response). Whether you consider that to be a "bug" or not really wasn't the point of my message.

      2 replies →

  • Feels pretty normal to me, IDK. "Rollout" is definitely what was happening here, and this leftover problem can be described as an "artifact" most generally - I guess otherwise it'd be a... just "problem" or "mistake"? Cause "bug" doesn't really fit. Plus, the Claudism here would definitely involve "soak", and possibly even "wall-time" lol

    IMHO it's worth keeping in mind that Anthropic employees are some of the least likely to casually pass off artificial prose as authentic, given the company's ethos/brand/cover story (depending on how cynical you are). To them this is all getting pretty high stakes pretty damn quickly; based on my usage of full strength Opus 5.5 today, I can't even imagine what working with their full internal stack must feel like. If they were willing to let the machines speak for them, they'd all be melancholically lounging around home by now instead of coming in to work!

    ...I am refusing to consider the fact that they probably are still WFH because of Salesforce forcing their shared security contractor to strike. Call that a mental health ignorance on my part :)

  • I was probably an agent that made the change, and the same agent that commented here on HN.