115104–126Index score#149 of 229 models
3/ 8Core tests taken2 of 6 areas · plus 12 anchor tests
−25Versus the frontier at releasebest then: Gemini 2.5 Pro
$0.19/ $0.65Price per million tokensinput / output, list price
1.05MContext windowtokens it reads at once
Where it sitson the frontier
o1-mini reached this score 7 months earlier.
Test resultsby area · tick = best by any model
Reasoning & knowledgeevidence
Mathsevidence
FrontierMath T1-3no result Codingevidence
SWE-bench Verifiedno result Long tasksno evidence yet
METR time horizonno result Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.