Comment by Garlef

18 hours ago

I think they even more so need deterministic feedback:

I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.

https://habit-hooks.com/

I'm using it to foster IOSP (integration operation segregation principle) for example.

Deterministic feedback is precisely how frontier models are trained. It’s called RLVR. You let the agent run on a problem and then calculate a deterministic score of how well it did. Repeat 1000x times and you can “brute force” a good solution. (Which includes all thinking traces and you add it to your training data.) And then a Chinese model can copy your advance for 1000x less compute. Which is why US labs call this not learning, but a distillation “attack”. It’s an attack on the business model.

This looks pretty cool but I'd want to be able to setup a bunch of my own project specific code smells, so it's not just a few generic rules. Is that the approach you've taken?

Then for example you could write your own hook and convert existing code smell documentation which agents ignore into a format that works with the hook.

  • > but I'd want to be able to setup a bunch of my own project specific code smells

    That's what I did; I did not use the library I linked to ~ It served only as an inspiration.

    Instead, I let the agents create custom lint rules (using eslint, pylint, ...) and add custom coaching error messages based on where I want to take my codebase.

The readme reads like AI slop. Why would one believe the tools would prevent AI slop?

  • Here's a link to a video where the creator explains the idea

    https://youtu.be/6AgndHSkHFI?t=238

    (You could of course argue that you don't like the direction the rules push the agents in given in the example - some people don't prefer small functions everywhere - but that's not the point: The lint-hooks work in pushing the agent in the desired direction; If one desires something else they'd simply need different rules)

This sounds interesting but (from the homepage too) I don't understand how it's different than having it run any other linter?

  • The difference is that the error messages contain instructions on how to resolve the issue.

    Just flagging big functions will make the agent write small functions ~ but not necessarliy in a good way

    (for example the agent might just cut `doOneThingAndTheOther` in half and call the second half `doOneThingAndTheOther2`)

I suspect strategies like this might be even more powerful with less intelligent agents. Could a 40B model outperform a 400B model with good feedback and instructions?