Comment by Lerc
23 days ago
They definitely have some. They handle the basic test scenarios where multiple people alter some state without the knowlege of the other. Most models seem to handle tracking differing knowledge between entities.
Some of the simple bench tests show how thin that can be.
Things like two lifelong friends had a fight when they were 8 years old because A broke B's favourite toy. They are now 25 and B sees A drowning, Will B try and save A?
Models tend to place massively undue influence on facts just because they have been mentioned. If every part of the text is accurate and relevant, then this helps immensely. Many fine tuning examples are precisely on topic. That leads to a bias against ignoring trivia.
Fine tuning has gotten a lot better now that the value of nuance in training examples is better known.
No comments yet
Contribute on Hacker News ↗