← Back to context

Comment by DanielHB

18 hours ago

Claude ran npm update (update all dependencies to the latest version compatible with the semver specified) in my repo without telling me when trying to fix some problems. Given that only updates the dependency lock-file I didn't notice and it caused several hours of debugging for me.

It is quite sneaky how LLM output can sometimes bypass human verification like that. No one is going around checking every single line change in auto-generated files. Someone could easily sneak a malicious dependency in there through some online tutorial that the LLM searches for.

> No one is going around checking every single line change in auto-generated files.

There's a simple fix for your particular case: commit your lock fine (which you should do) and always review the diff (which you should also do). (:

  • You've made me realise a good signal for bug hunting: Search repos with lock files listed in their .gitignore.

    It's the sort of terrible practice that someone might be frustrated into taking after a nasty merge conflict, and signals a willingness to cut corners.

    • My lockfile was not gitignored, I had made significant changes to package.json so I was expecting diffs in the lockfile.

      I just don't usually read lockfile diffs and claude inadvertently updated a few dozen packages to new minor versions without me noticing. In fact I only realized the problem after I looked at the lockfile diff.

      2 replies →

    • > You've made me realise a good signal for bug hunting: Search repos with lock files listed in their .gitignore.

      What would be the point of that? Do you just go around hunting for bugs in random repos?

      9 replies →

  • That's a manual step, not a solution. the solution is just boring basic file permissions. Treat Claude as semi hostile user. If you don't want them accessing your files, set the permissions to exclude them (like require sudo).

    I already do this for my unit tests, because Claude will "fix" the tests so they'll pass.

  • I made changes to my dependency lists in the same code where Claude ran npm update. The lockfile diff was a few hundred lines after I undid what Claude did.

    And yes, eventually I did check the lockfile changes and spotted the problem. I just usually don't check the lockfile that throughly.

    • > I made changes to my dependency lists in the same code where Claude ran npm update.

      ...but was it in the same commit? Two "update lockfile" commits, one yours and one Claude's should have made this obvious, no?

      Here's another useful rule of thumb: never mix your changes with the agent's changes. Agent always starts with a clean repository (no pending, uncommited human changes). You always start with with a clean repository (no pending, uncommited agent changes).

      Personally I have this in my `AGENTS.md`:

          ## Commit early, commit often
          You are allowed and encouraged to produce small, self-contained commits.
          Never `git push`; I will always review and rebase the full history and do the push myself.
          Commit messages should be *short* and on-point. They're there for *me* to review your work, and *not* a public historical artifact.
      

      So my workflow is usually this: start agent with a clean repository, tell it to do a thing, it works in the background, then once it's finished I come back, review, rewrite and clean up half of what it wrote, then maybe iterate some more with it, and finally do an interactive git rebase to get a clean commit history.

      4 replies →

I cannot tell you how much time I have saved by stopping Claude and asking, "what are you doing?"

At least 50% of the time, Claude "realizes" it already has all the information but is doing something that's unnecessary for the current work, stop, and tell me the previous step has completed.

People complain about approval prompts etc and have Claude run in fully autonomous mode. Outside small bug fixes, I just never find that useful. It helps me immensely to see what commands Claude is running to understand where the work is going.

  • Also Claude often cooks up atrocious overengineered ideas but responds pretty well to being guided hands on to the desirable scope.

When testing Claude code in auto mode in a fresh sandbox with a docusaurus website freshly cloned, I asked:

Can you see the docs folder with the git project?

It was in auto mode. So it immediately saw a docusaurus site (good) but instead of stopping there, it installed nodejs from a static binary download (no root access so only way), ran npm install, started the dev server and confirmed the project worked.

That is some crazy amount of leeway for an intent based classifier. I'm not surprised it's full of holes, and it seems to be 100% by-design.

Worse: Claude installed packages by just typing versions into package.json instead of running `pnpm install x`, then when running `pnpm install`, discovering that the package versions are too new and incompatible due to the default `minimumReleaseAge`, then proceeding to circumvent this by disabling `minimumReleaseAge` and running a full package update :)

Maybe I don't understand you correctly but if your lock file isn't in Git then you have bigger security issues than LLM output, given the last years NPM worms. Unless you're a single developer and the file on disk is the primary source of truth.