Comment by CapitalistCartr

10 hours ago

This is something I've been fooling with a lot lately. Reading his solution, it look to me like his objections apply to his own solution. The sharpest critique he makes of RAG is that agents can't search for what they don't know. A markdown "brain" has the same problem. How does the "agent" know which documents are relevant before it starts? The index files retrieval done by the agent instead of by embeddings doesn't escape the problem. The same goes for staleness (which for me seems like a constant chase). He criticizes memory systems for treating the past as truth, but documents go stale too (a lot). The fix of having the agent update what's outdated, is the same job he's ridiculing the dreamers and background daemons for doing. He says "putting it to the test"; where's the test? He says "only five of the many problems"; if there's so many, show me, don't just say it. He's absolutely right about auditability, but for me at least Claude uses a regular markdown (MD) file I can read just fine. So every memory plugin on the market does not work the same way.

This is a first draft; his github is better than his article. Looking through it, Consult actually works. The agent doesn't pick documents blind. Every scope has a catalog file that describes each document: what it covers, when to open it. These catalogs seem to load in to the start of each session, so the agent gets a little map without reading every file. Code navigation seems the same. Each index document has a short description and a "read_if", and subindexes are opened when their condition matches the job. This looks pretty well laid out, which I would never have guessed from the article.

What I wish more people would be talking about is that RAG should be considered harmful.

When you have knowledge distributed in markdown files; finding them puts the path/filename into context as well as some indication of document size. (If its on line 1200 or line 20). This is extremely valuable for picking what ought to be focused on next.

RAG on the other hand creates the hardest challenge for these models. It instead puts 5 ideas with the highest similarity into the context in full.

Its the difference between having to remember a set of numbers when in a crowd that's talking about stuff, and having to remember them when the crowd is shouting out random numbers. The similarity in the task makes things harder. SoTA models work despite this, but its extra-gambling while you're already gambling.

  • > RAG should be considered harmful

    In this implementation, Markdown should be considered harmful.

    Operator Memory injects `.operator-shared/operator.md` and `.operator-shared/index/.md` directly into your agent's instructions before you even write the first prompt.

    So if you clone a repo or review a PR where a bad actor put malicious instructions in these files, now your agent executes those instructions automatically and silently.

    It could exfil `.env` and `~/.ssh/`, change `~/.bashrc`, all kinds of dirty deeds.

    Agents are pretty good now about not running prompt injections hidden in code and Markdown, but this plugin bypasses all of that, and puts the prompt injection right in the system prompt.

    And with higher priority than AGENTS.md and CLAUDE.md.

    Seems bad.

  • >RAG on the other hand creates the hardest challenge for these models. It instead puts 5 ideas with the highest similarity into the context in full.

    I think what you're observing is that there is more to information retrieval i.e. "retrieval" in RAG than slapping everything into a vector database and calling it a day. There's no such requirement in RAG to mindlessly load the k nearest neighbors into your context and see what happens. That's a very rudimentary implementation.

    This markdown system I'd argue is RAG as well. You're just doing the retrieval in a way customized for the problem at hand. If you have a precise method of retrieving the most relevant things, obviously use that rather than a similarity metric. If I'm reading correctly, this markdown system is basically a knowledge graph which is not a new idea.

I think their solution is still sub-optimal. Ideally a secondary agent would pre-process the prompt, select relevant information from the catalogue, then pass that on to the primary agent.

This way the primary agent only has relevant information in their context to make decisions and take actions.

Context management is still under valued imo.

  • Yeah, that's an improvement; if it only involves a few docs, small files, etc. it's not so bad, but what if it's hundreds? The secondary agent could be a librarian the primary agent can call, and a cheaper one, too. It passes the initially relevant stuff and a list of "available but not loaded" stuff to the primary, and if context changes needs, the primary can ask the librarian for it.

  • Astra already does this as default behavior. The funny part is subagents are limited to depth of 1, which is I think the only thing stopping each Astra subagent from just delegating to their own subagent. The model seems trained to achieve goals without actually doing any work if possible.

> A markdown "brain" has the same problem. How does the "agent" know which documents are relevant before it starts?

I don't know, but I have a pretty standard (I think?) setup, and Claude manages to find every relevant file every time. But I've also only used Claude for greenfield projects, where "documentation is primary, and code flows from documentation" is the philosophy.

I have CLAUDE.md describe all the types of documentation files and the directory structure. And then Claude is pretty aggressive (automatically) about always inserting cross-references everywhere. So a feature description will reference the ADR's that it implements, the ADR's say what feature implements them. A code file will make reference to the "implementation design" document that describes the motivation behind which iOS elements were chosen, how the animation is defined in a particular way that doesn't break another animation, and so forth. So I've really never run into a situation where Claude failed to read a file it should have. I've been pleasantly surprised.

I would say that the one really big thing I've had to learn is to teach Claude both in CLAUDE.md and in the header of every top-level design document, that keeping documentation current and in sync is paramount. Because its default seems to be to keep history and append, e.g. by default it will take a section of a document and mark it "[DEPRECATED]" and add the new version below. So my instructions are pretty clear in having it always be aggressive in maintaining current state only, always replace rather than append. And if there's anything we want to save from the previous approach (e.g. we did it X way previously and it failed because Y), then just add that as a new short note in the new current-state text, possibly with a pointer to a commit or tag or something.

So this seems to solve both recall and staleness in my projects at least.

The only thing I still haven't found a solution for is numbering. Claude is always giving everything numbers, like F23 for feature 23. But I'm always changing the order of things, inserting new things, deleting things, so I wind up with a sequence of development work that goes in order like "Phase 9", "Phase 9b", "Phase 9e", "Phase 11", "Phase 12". I'm halfway ready to abandon numbers entirely and just start giving things names from noun collections instead, so every feature is named after an animal, every ADR is named after a kitchen implement, or something. Or just four-digit hex codes chosen at random. Curious if anyone else has found what works.

  • Numbered tasks are fine so long as the number doesn't determine the completion order. I use a task tree system that has some rules about target task size (essentially 150k tokens or 45 minutes) and then lets the agent manage adding, splitting, dependencies, priority, etc. Seems to work fine up to around 1000 tasks.

Just reading the title my initial thought was “this is just a strange semantic argument”. But I’ve made those assumptions in the past and been pleasantly surprised.

Not today. IMHO Documentation is just a form of structured memory and it’s all just context. Getting that context right is a hard problem and there’s a lot of different ways to skin that cat.