Where it sitson the frontier
o3 reached this score 5 months earlier.
Test resultsby area · tick = best by any model
Reasoning & knowledgeno evidence yet
Humanity's Last Examno result Mathsno evidence yet
FrontierMath T1-3no result Codingevidence
SWE-bench Verifiedno result Long tasksno evidence yet
METR time horizonno result Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.