123112–134Index score#133 of 229 models
2/ 8Core tests taken2 of 6 areas · plus 9 anchor tests
−13Versus the frontier at releasebest then: o1Replaced by DeepSeek-V3.1 $0.29/ $1.14Price per million tokensinput / output, list price
164KContext windowtokens it reads at once
Where it sitson the frontier
o1 reached this score 4 months earlier.
Test resultsby area · tick = best by any model
Reasoning & knowledgeevidence
Humanity's Last Examno result Mathsevidence
FrontierMath T1-3no result Codingevidence
SWE-bench Verifiedno result Novel problemsno evidence yet
Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.
What others sayrank on each leaderboard
The Model Indexindependent tests, our method#133of 219 · 123
Further left is better. Leaderboards measure different things: LMArena is people voting blind between two answers, the others are test scores. Click any row for the source.
Best attop three among current models
Replaced by DeepSeek-V3.1; it no longer leads any category among current models.
Where this model places in the top three of today's models, on category leaderboards and on our core tests.