Comment by Jirach05
5 months ago
Can anyone explain why these models decrease in performance on this "MCRC v2 (8-needle)" long context benchmark when thinking is turned on?
5 months ago
Can anyone explain why these models decrease in performance on this "MCRC v2 (8-needle)" long context benchmark when thinking is turned on?
No comments yet
Contribute on Hacker News ↗