131120–143Index score#114 of 229 models
4/ 8Core tests taken4 of 6 areas · plus 11 anchor tests
−15Versus the frontier at releasebest then: o3
$0.037/ $0.17Price per million tokensinput / output, list price
131KContext windowtokens it reads at once
Where it sitson the frontier
o1 reached this score 8 months earlier.
Test resultsby area · tick = best by any model
Reasoning & knowledgeevidence
Humanity's Last Examno result Mathsevidence
FrontierMath T1-3no result Codingevidence
SWE-bench Verifiedno result Novel problemsno evidence yet
Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.