Comment by Alifatisk
10 hours ago
I wish there was a clean way to compact the conversation into a prompt with all necessary context for a new fresh conversation.
10 hours ago
I wish there was a clean way to compact the conversation into a prompt with all necessary context for a new fresh conversation.
I'm fed up with compaction. I want my agent to get compacted but also retain full access to the prior conversation via search and tool calls - I want it to know "the requirements for X were discussed in detail previously in conversation C51E31CE-C985-4633-A749-DCC9805A7FEB" and have a tool that lets it dispatch a subagent to find those details again.
Do any of the coding agents have this already?
Hydra is an agent in spirit, but actually wraps other agents through ACP. It drives its own async compaction algorithm that gives the agent it wraps full access to history that it can search as an MCP server after compacting.
https://github.com/smagnuso/hydra-acp
This one does agentic search instead of compaction: https://github.com/wedow/harness
Generally works quite well but it allows nested subagents and that can use a lot of tokens if the agent prompts are too open-ended.
All messages are persisted as separate markdown files under a session directory, which makes them much more friendly to grep and such than the common jsonl files are. Agents can search them rapidly without messing around with jq.
Also written in Bash so it runs on any old potato you have laying around.
Claude Code makes agents reasonably aware of where their log files/history/etc are and get stored. Generally they’ll work with them without explicitly being told (especially to recover broken sub agents, corrupted sessions, etc) to do so.
I think the more general problem is that compaction is just a bandaid: you HAVE to dump context to keep going and searching back for it is more expensive than if you had just kept the right context. The better a job the harness does at filtering out junk, the more likely compaction is to remove context that might have been, forgive me, “load bearing”
IMO the default Claude Code / Codex (which to my understanding is almost continually-compacting?) compaction has got much better over the part few months. If you spam sub agents then context will naturally nest, and you can just resurrect them as needed without polluting the main thread.
“Search session history for when we discussed XYZ”
https://github.com/nerdyaustin/memory_mcp
You can skip all the sync stuff, not necessary at all
Both CC as well as grok build for me seem to know about the log file location and they just read off the context from there
Create your own protocol. I created a "wind down session" protocol my agents use that takes detailed notes in a "next_session_prompt.md" file that covers what was done this session, what is still open, and where they need to pick up the next session.
You can refine the protocol as you realize what's working and what isn't. I've been using that for months and it rarely drops important things now.
Have you tried Matt Pocock's "handoff" skill?
https://github.com/mattpocock/skills/blob/main/skills/produc...
> Write a handoff document summarising the current conversation so a fresh agent can continue the work. > […]
I do this a lot and you have to be really careful to clean these up or qualify/steer agents around them. They’ll often be very emphatically confident about some assumption or implication they made, and if another agent stumbles upon them they’ll get mislead.
They don’t really know what they’re handing off or what you’re trying to actually do, so in a sense it’s not a grounded task for them. Actually, if you think about it, any scenario in which a handoff doc might be valuable is probably almost always better as a subagent thread, because you are paying the same amount of read/write tokens but you can clear things up synchronously.
I’ve found two-way message passing (each get their own write file, they read each others) to work much better because the communication is more grounded in actual coordination/work. You can also give each an inbox so that multiple can write to it. If you do the “progressive disclosure” right it scales subquadratically because they only read/write to others when it’s relevant to what they’re working on.
But IMO “write a handoff” is a trap, as a human you end working in some kind of robot-graffiti codebase full of junk, and it ends up being a booby trap for agents literally within days.
1 reply →
Just make a new slash command with those instructions as the prompt.
There is, but your wish of "clean" is ambiguous.
Here's a view for "clean:"
1. Every chat should have a context used/remaining measurement so you know when you have to ditch the current chat for a fresh one.
2. Every chat should analyze and categorizes each element of context by how useful it is towards the overarching goal of the chat.
3. Every chat has a handoff button with a "usefulness" slider (say 1-5) that shows the total size of the context based on its setting.
4. The handoff automatically creates a new chat with the desired amount of context and a prompt to get it back to where you were.
That said, I am newb and so there is some reason why these non-deterministic LLMs can't do this :-/
With clean I mean the opposite of how I currently do it, which is by asking the model to compact the whole thread into a prompt which will act as context for next model.
My way of prompting this varies and every time I receive the blob of output, I can’t fell how well it managed to capture the necessary details. This way feels lika a dirty way to transfer knowledge from one conversation to another.
You said there is? What’s the options?
My largest issue is that when I'm looking for this I'm already dangerously close to autocompaction. And what I really want is a prompt which manages to preserve the most important parts of the chat log. And my opinion of important will not be the same as Claude, so we'll need to iterate on what that handoff really is.
1 reply →