Comment by pimeys
2 hours ago
It is not super good in long horizon tasks and worse than 5.6 in our evals. It really failed the agentic evals where DeepSeek, Kimi and Opus are the winners.
It is great on creating summaries and content.
2 hours ago
It is not super good in long horizon tasks and worse than 5.6 in our evals. It really failed the agentic evals where DeepSeek, Kimi and Opus are the winners.
It is great on creating summaries and content.
No comments yet
Contribute on Hacker News ↗