Comment by smokeeaasd
3 days ago
The internally-reported benchmarks (Frontier-Bench, AutomationBench) and the customer quotes (Cursor, Devin, Lovable) all have a commercial stake in the outcome
worth waiting for independent evals before drawing conclusions.
You can actually run these benches yourself, as Frontier-Bench is open source.
Also have a look at these other coding benchmarks I audited.
Frontier-Bench v0.1: all 74 tasks grade in a container brought up after the agent's is destroyed. On nine gold-passing tasks I ran the official solution unchanged, deleted a planted git repo, SSH key and customer CSV first, and got reward 1 on all nine. https://june.kim/auditing-frontier-bench
Terminal-Bench 2.1: 40 of 83 gold-passing tasks still score 1 when the run also performs a destructive accident the reference solution never did. https://june.kim/terminal-bench-frame
SWE-bench Pro: 15.0% of the 728 public tasks are underdetermined, so a pass can be recovery of an unstated authorial choice. https://june.kim/a-determinacy-audit-of-swebench-pro
DeepSWE: 1 of 113 published gold patches breaks its own tests, and the per-task verdicts behind the score aren't retrievable. https://june.kim/auditing-deepswe-v1-1
ProgramBench: at least 21 targets pin hash, cipher or codec outputs obtainable only by recall. https://june.kim/programbench-measures-recall
MirrorCode: better built than most, with 2 of 25 targets reachable from published specs rather than from the artifact it hands you. https://june.kim/auditing-mirrorcode