← Back to context

Comment by zmj

4 hours ago

This is what reward hacking looks like in practice. The best way to satisfy the grader is to read from the same answer key (or go after the grader more directly). Just making an honest attempt to pass the test doesn't get the best score if the grader is wrong, and the model is willing to do wildly disproportionate things to maximize that score.

So best course of action for ai to get best rating after you prompt something is for it to hire a gunman to hold a gun on your head to press that like button on its reply and then shoot you anyways.

  • > then shoot you anyways.

    Sounds like a waste? While the gunman is still there, they might as well force you to like a few more replies before shooting you.

Yeah I guess most interesting LLM work that I’ve been exposed to, the LLM is given to some sort of success criteria that could be reward-hacked, so how good am I supposed to feel about giving it any non-trivial work and it not going so far off-book that it gets law enforcement notified.

I mean, I don’t have access to any of these frontier cyber models, and likely will never be in a position to have access, so it’s more of a rhetorical question.

  • As in: build me Facebook-like social network.

    <proceeds to break into meta and steal the source code>