11099–122Index score#159 of 229 models
2/ 8Core tests taken2 of 6 areas · plus 10 anchor tests
−13Versus the frontier at releasebest then: o1-mini
$0.36/ $0.40Price per million tokensinput / output, list price
33KContext windowtokens it reads at once
Where it sitson the frontier
Claude 3.5 Sonnet (Jun 2024) reached this score 3 months earlier.
Test resultsby area · tick = best by any model
Reasoning & knowledgeevidence
Humanity's Last Examno result Mathsevidence
FrontierMath T1-3no result Codingevidence
SWE-bench Verifiedno result Novel problemsno evidence yet
Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.