140129–152Index score#87 of 229 models
6/ 8Core tests taken4 of 6 areas · plus 14 anchor tests
$15/ $75Price per million tokensinput / output, list price
200KContext windowtokens it reads at once
Where it sitson the frontier
Gemini 2.5 Pro reached this score 4 months earlier.
Test resultsby area · tick = best by any model
Reasoning & knowledgeevidence
Novel problemsno evidence yet
Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.