Comment by NitpickLawyer
16 hours ago
> not sure it's apples to apples comparison.
They're literally comparing the previous version of the same model with the new one. It's based on the same architecture, same pre-trained model, just different post-training. It doesn't get more apples to apples than this.
I think the commenter means the Flash vs Terra benchmarks.
Ah, my bad. Yeah that makes sense. They do say "The official V4-Flash natively supports the Responses API format and is specifically adapted for Codex.", so at some point someone will make a "same harness" comparison.