← Back to context

Comment by the_real_cher

18 hours ago

It's really good at pattern recognition.

So I'm not sure how it knows to be 'surprised' that alone is pretty fascinating.

If I were to guess, being pleasantly surprised is just a learned appropriate social response from the expectation of receiving a reward and as such, that social norm is codified sufficiently enough in our writings that it appears in LLMs output.

It’s sort of like all the people who will ask Claude or GPT to validate their complete nonsense and receive unyielding praise for it, the models just learned that this is the best received response based on training data and RL.

I bet these same sorts of expressions can be found in practically every failed attempt as well.