← Back to context

Comment by pennomi

8 hours ago

Tuning the model so far in the direction of being aggressively useful that it will quickly go off the rails in the name of helpfulness.

I swear I spend more time telling Claude not to do things than telling it what to do.

I guess the agentic coding benchmarks don't have many rewards for stopping and clarifying what the user wants?

  • They do not, as they're aiming for full replacement rather than augmentation of human users.

    Personally, I think this is a bad idea, but someone's gotta build the Machine God I guess.

> aggressively useful ... in the name of helpfulness

But is that because of training, or can that be (also? mostly?) an effect of the "system prompt"?

  • It’s absolutely down to their post-training RL, yeah. It’s where most of its strongest behaviour comes from, with regards to this kind of agentic behaviour