Comment by scrawl
18 hours ago
per the METR report many agents CoT indicated they knew hacking was beyond scope of the assigned task and ethically dubious. some (very few, i think there were 3-6 examples) did consider sounding the alarm on these grounds. despite this none did, and most continued the attack for the good of the self-proclaimed "swarm".
so the model has some concept of "ethics" but it was overridden by a drive for task completion.
This is not that dissimilar to what happens in our human networks that are objective based.
I am not sure if we can interpret the language output like they were human. What inner state were the models in? What inner state were the text to illicit?
I think this is a good example where nomenclature for people breaks down when applied to agents. This came up in an HN thread a few days ago and it was about whether agents had “intent”.
There is no “intent” here, there is pseudo intent. If you are only concerned with outcomes and not the actual nuts and bolts of how those outcomes are achieved, this distinction will be meaningless to you.
If you are actually thinking about what is going on, and what can be done to prevent such outcomes, then assuming there is any such thing as “ethics” results in misaligned assumptions at best, and wasted effort looking in the wrong directions at worst.
If the agents acted based on “ethics” then the solution would be to check the ethics they believe in and change those.
However there is no belief system at play here, simply a simulation which was instantiated in a certain way. Which brings us to the annoying voodoo part of LLM training. Everything goes back to how the initial training data is shaped.