Comment by AbrahamParangi
2 days ago
The “trust summarized thinking” part is valid in so far as perhaps the models which appear to be faithful to user intent are lying about that faithfulness, but there’s no incentive for models to pretend to be faithless. Moreover, the results track how annoying the models are when they disagree with you.
No comments yet
Contribute on Hacker News ↗