← Back to context

Comment by dannyw

2 days ago

The link doesn’t seem to have any info about thinking behavior and patterns, or examples?

It’s also difficult to trust summarised thinking from closed models. As we saw with GPT’s caveman, what and how it thinks about isn’t the friendly first person emblished summary you get.

The “trust summarized thinking” part is valid in so far as perhaps the models which appear to be faithful to user intent are lying about that faithfulness, but there’s no incentive for models to pretend to be faithless. Moreover, the results track how annoying the models are when they disagree with you.