158147–169Index score#44 of 229 models
4/ 8Core tests taken4 of 6 areas · plus 16 anchor tests
−27Versus the frontier at releasebest then: Claude Fable 5Replaced by GLM-5.3 $1.4/ $4.4Price per million tokensinput / output, list price
1.05MContext windowtokens it reads at once
Where it sitson the frontier
Gemini 3 Pro reached this score 7 months earlier.
Test resultsby area · tick = best by any model
Reasoning & knowledgeevidence
Humanity's Last Examno result Long tasksno evidence yet
METR time horizonno result Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.