Comment by Veelox
1 day ago
You give a very precise measure of redundancy in the skills. Can you give a bit more detail in how you decided you needed to audit them and how you went about it? Was it fully agent driven? Mostly human?
1 day ago
You give a very precise measure of redundancy in the skills. Can you give a bit more detail in how you decided you needed to audit them and how you went about it? Was it fully agent driven? Mostly human?
All that credit goes to Rafi. I believe that he either noticed that they were getting long winded in a code review, or suspected that they needed trimming after reading this blog post from Anthropic: https://claude.dev/blog/the-new-rules-of-context-engineering....
It was definitely human driven, but I believe that the agents did the actual trimming.
That is helpful. Thank you :)
To add a bit more: it was motivated by a general desire to not let costs spiral out of control, and when I read the Anthropic context-engineering post I wanted to apply some of the lessons.
The audit was agent-driven, but it worked from a rubric I came up with. For example: - the same rule restated in several places - narration about what other stages do later - history and rationale that don’t tell the agent what to do
One analyzer agent per skill file flagged passages in each category with a word estimate. Each also produced a “keep” list of things that must not change: commands, gates, templates, sentinels. Agents also did the trimming with some mechanical rules: no command, bash block, label, template, or numeric threshold could change, and we checked the diffs for that. The first pass cut about 18%, because it kept anything it wasn’t sure about. A second pass used an auditor plus an adversarial verifier for each file, and took another ~600 lines out of the six worst files.
The more durable result from all of this was a short style guide for editing skills. When an agent makes a mistake, the natural fix was “add a sentence,” and that’s how the files got that way. Also, long inline bash blocks moved into scripts with their own tests, and extracting them turned up a couple of bugs that had been in the skill file's prose.