Comment by akyshnik
3 days ago
feedback is generated based on evals. example: eval: function foo wasn't triggered even though [...]
feedback (exaggerated): 1. change stage prompt 2. change function description 3. add extra instructions to the end of the context
metrics are easy to generalize (e.g. call transfer rate), but baseline is different for each agent, so we're interpreting only the changes, not the absolute values (in the context of self-improvement).
No comments yet
Contribute on Hacker News ↗