167155–178Index score#26 of 229 models
8/ 8Core tests taken6 of 6 areas · plus 19 anchor tests
−3Versus the frontier at releasebest then: GPT-5.3-CodexReplaced by Claude Opus 4.7 $5.0/ $25Price per million tokensinput / output, list price
1MContext windowtokens it reads at once
Where it sitson the frontier
Among the first models to reach this score.
Test resultsby area · tick = best by any model
Reasoning & knowledgeevidence
Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.