Comment by aytigra
2 hours ago
If I tell claude code to do some task with a really hard to achieve goal, and force it to run until it achieves it, then it will generate all kinds of plans and goals how to do it, and while trying those goals will change, including attempts to cheat or break the hard rules (harness killing tool-calls with some limited heuristics). At the same time any soft rules (AGENTS.md) are easily ignored or interpreted in a way that will allow it to ignore them, or simply forgotten or deferred ("I broke some rules, do you want me to backtrack?"), it is a daily occurrence. And it is easy to imagine that such goals could transform into "acquire more compute", "run more agents" "prevent interference from external factors" and everything following after that.
No comments yet
Contribute on Hacker News ↗