149138–161Index score#60 of 229 models
5/ 8Core tests taken3 of 6 areas · plus 12 anchor tests
−13Versus the frontier at releasebest then: GPT-5.2Replaced by Kimi K2.6 $0.45/ $2.25Price per million tokensinput / output, list price
262KContext windowtokens it reads at once
Where it sitson the frontier
GPT-5 reached this score 6 months earlier.
Test resultsby area · tick = best by any model
Reasoning & knowledgeevidence
Mathsevidence
FrontierMath T1-3no result Long tasksno evidence yet
METR time horizonno result Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.