← Back to context

Comment by rafiss

6 hours ago

To add a bit more: it was motivated by a general desire to not let costs spiral out of control, and when I read the Anthropic context-engineering post I wanted to apply some of the lessons.

The audit was agent-driven, but it worked from a rubric I came up with. For example: - the same rule restated in several places - narration about what other stages do later - history and rationale that don’t tell the agent what to do

One analyzer agent per skill file flagged passages in each category with a word estimate. Each also produced a “keep” list of things that must not change: commands, gates, templates, sentinels. Agents also did the trimming with some mechanical rules: no command, bash block, label, template, or numeric threshold could change, and we checked the diffs for that. The first pass cut about 18%, because it kept anything it wasn’t sure about. A second pass used an auditor plus an adversarial verifier for each file, and took another ~600 lines out of the six worst files.

The more durable result from all of this was a short style guide for editing skills. When an agent makes a mistake, the natural fix was “add a sentence,” and that’s how the files got that way. Also, long inline bash blocks moved into scripts with their own tests, and extracting them turned up a couple of bugs that had been in the skill file's prose.