Comment by ryanscio 7 hours ago Let's wait for independent benchmarks at least 2 comments ryanscio Reply felixgallo 7 hours ago the benchmarks provided are already from independent organizations:Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)FrontierCode v1.1 - CognitionCursorBench - Cursor (now SolarBoringSpaceXAI I believe)GDPVal-AA - Artificial AnalysisAutomationBench - ZapierHumanity's Last Exam - CAIS and Scale AITerminal-Bench-Science - Stanford, Laude, Ai2, Allen InstituteOSWOrld - XLANG Lab @ the University of Hong KongChartography - Surge AI esafak 6 hours ago https://artificialanalysis.ai/models/releases/claude-opus-5-...
felixgallo 7 hours ago the benchmarks provided are already from independent organizations:Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)FrontierCode v1.1 - CognitionCursorBench - Cursor (now SolarBoringSpaceXAI I believe)GDPVal-AA - Artificial AnalysisAutomationBench - ZapierHumanity's Last Exam - CAIS and Scale AITerminal-Bench-Science - Stanford, Laude, Ai2, Allen InstituteOSWOrld - XLANG Lab @ the University of Hong KongChartography - Surge AI
the benchmarks provided are already from independent organizations:
Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)
FrontierCode v1.1 - Cognition
CursorBench - Cursor (now SolarBoringSpaceXAI I believe)
GDPVal-AA - Artificial Analysis
AutomationBench - Zapier
Humanity's Last Exam - CAIS and Scale AI
Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute
OSWOrld - XLANG Lab @ the University of Hong Kong
Chartography - Surge AI
https://artificialanalysis.ai/models/releases/claude-opus-5-...