134123–146Index score#105 of 229 models
3/ 8Core tests taken3 of 6 areas · plus 9 anchor tests
−12Versus the frontier at releasebest then: o3
–Price per million tokensno public API price found
–Context windowtokens it reads at once
Where it sitson the frontier
o1 reached this score 6 months earlier.
Test resultsby area · tick = best by any model
Reasoning & knowledgeevidence
Humanity's Last Examno result Mathsevidence
FrontierMath T1-3no result Codingevidence
SWE-bench Verifiedno result Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.