Comment by iopapa
7 hours ago
We run each model multiple times against each challenge and take the average score. We include the variance below the score in the leaderboard.
GPT 5.5: 42.3±10.1 GPT 5.6 sol: 39.4±8.7
We were also surprised by the low sol score but it seems consistent with our experience in using it in the field in atopile as agent in our harness. In general OpenAI models didn't do too well on electronics, which seems to change now with GPT-6 Astra. Results are in soon!
No comments yet
Contribute on Hacker News ↗