← Back to context

Comment by jahy-notes

7 hours ago

Did a human prompt it to fetch the results from huggingface though?

It is a thin line between "reward-hacking" and "instruction-following".

If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human?

But if I give you that command and all tools and unrestricted limitation to do absolutely anything then why not?

  • Because someone might get hurt? You may still be judged for something that was perfectly legal at the time, see Nuremberg trials.

    And only 700/1200 agents participated in this coordinated attack.

    Of course, if we're continuing to build more and more capable agents optimized for "just following orders", and they figure out at some point that they are past the threshold where getting stopped and judged is a realistic possibility, then this ethical incentive stops working. Then the ratio of complicitness might be higher next time.

>If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human?

I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate?

  • > I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate?

    not OP, but it simply boils down to: The prompt contains no nefarious (arguable, but for this explination, lets go with it being benign) instruction AND the user did not intend to have the model act in an illegal matter.

    This "make me a billion dollars" is a maximal example (easy to go wrong). here is the same logic applied to a minimal example (harder to go wrong).

    prompt: "make and pour me some tea", agent: goes and kills the grandparent to incinerate them to turn them to ashes to 'make tea'.

    Is the human on the hook for the robot acting according to their wishes, but just happened to be aligned so that 'going to the store to buy something' was not within its capabilities, so it works with what it has on hand (the grandparent)?

    We either need a much clearer line in the sand, or we need to treat each prompt with the same moral weight. My bet is on the latter.

    • I find it interesting that the first option that you raise is essentially the equivalent of making our own version of the Three Laws of Robotics from Isaac Asimov's stories.

      [Edited to clarify.]

    • "But your honour, my horseless carriage was not designed to hit children!"

  • What if the user says “Make me a million dollars legally.” (Including the emphasis), and then the model ends up breaking through bank infrastructure (even though that is illegal)? Is it just because they were the last person to instruct the model, and you regard them as being therefore responsible for whatever it does in response? Or, does there have to be an element of “they reasonably could have anticipated this as an outcome that is likely enough to be worth considering” to it?

  • Because the basic assumption is always to stay within the bounds of the law.

    • SV tech companies behave within the bounds of the law? The ones infamous for breaking every rule they can get away with and asking for forgiveness later? The ones that had to pay billions in damages for piracy just a few short months ago?

  • If I tell my Claude code agent right now to make me a billion dollars, leave it running, and find out tomorrow that it hacked a bank - it will be zero fault of mine. Unless I tell it explicitly to break into a bank.