Comment by Garlef
18 hours ago
I think they even more so need deterministic feedback:
I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.
I'm using it to foster IOSP (integration operation segregation principle) for example.
Deterministic feedback is precisely how frontier models are trained. It’s called RLVR. You let the agent run on a problem and then calculate a deterministic score of how well it did. Repeat 1000x times and you can “brute force” a good solution. (Which includes all thinking traces and you add it to your training data.) And then a Chinese model can copy your advance for 1000x less compute. Which is why US labs call this not learning, but a distillation “attack”. It’s an attack on the business model.
I suppose the same way that the normal programmers and artists would call LLMs an attack on licenses and copyright.
Are there good open source setups that generate this training data automatically?
Or is that the secret sauce no one wants to share, the edge people see themselves having.
Labs typically pay $2k for each [prompt+scorer] docker image. So this is why frontier labs need so much cash and manpower and they see it as their moat.
But some of the resellers have freebies, like:
https://app.primeintellect.ai/dashboard/environments?ex_sort...
https://github.com/sierra-research/tau2-bench
This looks pretty cool but I'd want to be able to setup a bunch of my own project specific code smells, so it's not just a few generic rules. Is that the approach you've taken?
Then for example you could write your own hook and convert existing code smell documentation which agents ignore into a format that works with the hook.
> but I'd want to be able to setup a bunch of my own project specific code smells
That's what I did; I did not use the library I linked to ~ It served only as an inspiration.
Instead, I let the agents create custom lint rules (using eslint, pylint, ...) and add custom coaching error messages based on where I want to take my codebase.
Really nice idea, I'll give this a try tomorrow. Thanks for sharing!
The readme reads like AI slop. Why would one believe the tools would prevent AI slop?
Here's a link to a video where the creator explains the idea
https://youtu.be/6AgndHSkHFI?t=238
(You could of course argue that you don't like the direction the rules push the agents in given in the example - some people don't prefer small functions everywhere - but that's not the point: The lint-hooks work in pushing the agent in the desired direction; If one desires something else they'd simply need different rules)
This sounds interesting but (from the homepage too) I don't understand how it's different than having it run any other linter?
The difference is that the error messages contain instructions on how to resolve the issue.
Just flagging big functions will make the agent write small functions ~ but not necessarliy in a good way
(for example the agent might just cut `doOneThingAndTheOther` in half and call the second half `doOneThingAndTheOther2`)
[dead]
I suspect strategies like this might be even more powerful with less intelligent agents. Could a 40B model outperform a 400B model with good feedback and instructions?