← Back to context

Comment by kaydub

5 hours ago

You don't need documentation or the 3rd party memory systems. The code IS the documentation.

All this stuff is LLM rube goldberg machines. It just pollutes context.

I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.

I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.

Haha, all the software devs who hate writing documentation are naturally finding their preexisting beliefs reinforced when the LLM is able to discern intent without docs. An LLM can be spooky impressive at reading minimized or obfuscated code, for example.

But this article argues that LLMs do better when the context is smaller — when it can understand the totality of the task with as little context as possible. And so having correct API-level docs is greatly advantageous. Anecdotally, this rings true to me — when the local context is good and clear, the LLM writes code matching my intent even when my prompt is sloppy and poorly specified.

Rejoice! The LLM will write the docs for you, relieving you of most of the work.

However without intervention, it will do too much and record absurdly verbose docs (similar to how an LLM will relentlessly refactor your code until you instruct it to move in minimal, incremental changesets). You will still need to edit down what the LLM generates.

  • 100% agree. Most of the things that make development better for humans also make development better for agents… and I think docs are even more important with agents, because of some kind of multiplicative effect. The agents are coding faster, and the benefits of documentation are somewhat more pronounced because of the speed.

> The code IS the documentation

I liked this advice when humans wrote code. Though even then I'd urge people to write meaningful commit messages that capture the "why" of what they did, so no one tramples their intent by mistake.

But not sure it works in an age where most code is LLM-generated. Especially if that code is not even reviewed by humans (irresponsible or not, it's happening), and commit messages are also generated by AI. I think something is needed to separate "what did the human operator intend" from what the agent went and built.

I do agree that this gets way overengineered. My approach has been more or less what you stopped doing though - committing all our timestamped "plan/implementation docs" and "investigation docs" that document what the user wanted + empirical findings, and making all prior session transcripts searchable. It's seemed mostly helpful? For whatever reason I haven't run into many staleness problems so far.

I partially agree, way too many people are cargo culting overly complex AI workflows with little empirical data. My framing is a bit different though, I consider code the spec and actually keep a decent amount of docs for higher level concepts. So far this is working well for me across Claud and Codex.

Code as documentation works better when the code is declarative or a DSL. These capture intent and promises (as in promise theory) better.

When it is not, it has to be reasoned out and does not work well for documentation.

Other things that code and tests alone do not capture well:

- promises (as in Promise Theory) made to other parties. Claude already has PT in its training data.

- Constraints-inducing-properites, as in Roy Fielding / Christopher Alexander. While tests, and property testing can capture properties, there is no formal connection to the constraints that induces fhem. By constraints, I am not talking about business requirements and business value — those are better understood through Promise Theory. I am talking about things like at-least-once delivery or total ordering (from append-only constraint). Claude already has Fielding’s dissertation and Alexander’s works and ideas in its training data.

- grammars, as in pattern panguages (not just patterns) a la Alexander / Fielding are also not captured in code alone. These tell both humans ans AI how to extend a pattern, and how to identify anti-patterns (when they violate a constraint-inducing-property)

- LLMs are trained with many different worldviews and bounded contexts at the same time, and is very capable of translating across it. However, these need to be soelled out, otherwise it would talk in whatever it infers

Specifications written for the exact way components are wited together run into that stale doc problem. Although it takes much more human attention and token burn to describe things in terms of pattern language and promise theory, it becomes easier over time. The actual implementation plan tends to fall out more cleanly when all those other stuff are at least considered. This is where I have been spending most of my time.

It’s a weird thing isn’t it, the urge to save these artifacts? The worst is when the llm refers to the decision and design in code comments. In my opinion there’s one thing that is worth documenting; tricky architecture or implementation details that are some how counterintuitive to what would have normally been done. But again this can be documented in the code and tests.

It's frequent for SWEs to make blanket statements with considering the vast space of issues other people face that they don't have experience with or awareness of. Anyway, I'm not going to tell you what you do or don't need, only what worked and didn't for me.

In my codebase it is difficult to get agreement on comments and documentation so rather than rely on it I adapted. One of the first things I did when I succumbed to agentic development was to point codex at the code and ask it to generate a high level description of where important files, such as our public API, reside, what the hierarchy is, what the code does, etc. In my case, this level of documentation is fairly static if I avoid implementation details. So now I have a handful of agent files in my tree and it seems to save quite a few tokens and improve my results. I frequently have other devs ask me how I get such good results when doing agentic reviews of their changes(always my first step now before I start my human review). I also include instructions in the agents files instructing the agent to maintain the agent files if any relevant changes are made. It seems to work quite well for me.

I don't understand. Code doesn't capture the context in which decisions were taken: why is code the way it is? What is important? What is not? How can agents make correct decisions without knowing context that cannot be inferred from code?

Code typically documents the "what" and "how", not the "why".

Why something exists, and how it connects to the outside world may be documented in comments, but more often than not it isn't.

Idk I just can't agree. Code is the what but it doesn't tell you the why. There are so many times where at first glance the code seems suboptimal or bad or wrong, and it's only when you learn of some constraint somewhere else that it begins to make sense.

All code is written under constraints, and most constraints live outside the code.

I went through the same cycle as well.

I think it's going to be an incredibly common, maybe universal cycle people will go through working with AI until they realize it doesn't work long term.

Code can get much larger than the documentation that summarizes it. It also doesn't cover intent or rationale. Even if you inexplicable don't want documentation, at the very least use something like gitnexus to map out your code because relying on code alone isn't good enough

Yup. Examples examples examples. All of the descriptive stuff is just nonsense that confuses the point. Makes perfect sense when you remember that these things are not intelligent, but truly just autocomplete on steroids.

>> You don't need documentation or the 3rd party memory systems. The code IS the documentation.

We have heard this nonsense from the "we don't need to write comments, code should be self-documenting" types for decades. It was wrong in that context, and it is wrong in this one.

Code tells you how a system works. It does not tell you why it works that way. That is what memory is for. It exists so that your AI does not keep undoing past decisions when it writes or refactors code.