o1-mini reached this score 4 months earlier.
Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.
Further left is better. Leaderboards measure different things: LMArena is people voting blind between two answers, the others are test scores. Click any row for the source.
Not in the top three among current models on any category we track.
Where this model places in the top three of today's models, on category leaderboards and on our core tests.