Claude 3.5 Sonnet (Jun 2024) reached this score 3 months earlier.
Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.
Further left is better. Leaderboards measure different things: LMArena is people voting blind between two answers, the others are test scores. Click any row for the source.
Replaced by DeepSeek-V3; it no longer leads any category among current models.
Where this model places in the top three of today's models, on category leaderboards and on our core tests.