Comment by gregwebs

1 day ago

Agreed, and this seems better.

My thought though has always been that I don't want there to be agent-only designated documentation.

I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.

I've found that most of LLM generated docs are diluted and unfocused. Spending 3 paragraphs on some quirk of a library, then 2 simple examples of a command.

I write README and docs by hand, thus only important stuff goes in there. If it is not worth my attention to write down, it's not worth writing down.

Smaller context, easier to consume.

I've got basically a loop of docs, test (TDD) and code. Starting with and IMPLEMENT-<plan name>.md. I ask to revise as TDD, then loop.

Once qwen3.8-flash-next showed up, it cañ go "forever" with dynamic context pruning.

It's fascinating for local coding.

  • I’ve been trying to nudge agents (both Claude and GPT) into a red/green/refactor TDD loop and I’m increasingly unconvinced it’s worthwhile. TDD works for humans because it forces us into a pattern of doing the simplest thing that could work, then thinking about how to refactor that. An LLM can be prompted to do that without the skeleton of a failing test, and I plan to experiment with other methods of pushing them towards the same ends.

    • The value of TDD is it does ensure that tests are written and that will help with future regressions.

      However, I found the same problems with agents that it has with humans (usually with humans it isn't TDD but code coverage requirements). The problem is the tests are now written to satisfy a bureaucracy rather than to properly verify the code. And there are real studies that point out the low value of these types of bureaucratic testing requirements. For example, when models are just required to write a test file, they don't produce a better result: https://arxiv.org/abs/2602.07900

      I wrote a /verify skill that focuses on properly verifying code and I am finding it works a lot better. I do need to revise it now- in practice certain parts of the skill are doing all the heavy lifting and others are more dead weight. But it does seem to be properly orienting the agents towards finding defects. https://github.com/gregwebs/skills-sdlc/blob/main/skills/ver...

      3 replies →

    • I'm working on local models; qwen3.8-Flash-Next seems to be passed to rubicon.

      I'm working on javascript, and now I'm only doing this via typescript. I've successfully got this going:

      1. Design a feature in plain language with the coding agen (opencode) and write a document for it.

      2. Restart the context (or I use /compact to flush any errant details)

      3. Pull the new plan into context and ask the agent to revise it follow Test Driven Development.

      4. Depending on how big it is, the agent places it into multi stages, each with it's own document.

      It then loops through the stages. It might be entirely based on the language you're using, but this loop seems strong enough.

      Then when there's bugs, errors, anything, we update the doc, add more tests, then revise the code.

      It's quite possible the harness you're using isn't setup properly. For me to do this locally, I had to hack on dynamic context pruning which I described here: https://news.ycombinator.com/item?id=49906637#49907641

  • A lot of agentic software development feels like cowboy coding, set it up, let it rip and see what it figures out.

    The slightest amount of guidance, input from experience can make a huge difference.

    • yeah, to some extent. I've been pulling repos to play around but I've now got a hard fork of paseo; I ripped out anything related to github. I upgraded the relay with security features to prevent abuse. I'm in the process of adding the first "real" large feature update that'll let the agents work over the relay with one another which will let me add compute wherever I find it. In the process of adding a browser plugin that will let me pull up pages for context to work on whatever (I don't trust even agents for searching the web).

      And it's quite fascinating. Still dont think there's trillions of TAM out there if Qwen3.8-Flash-Next runs fast enough on $3k (before memory cartel) pricing.

      2 replies →