Not plotted: no index score yet.
Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.
Further left is better. Leaderboards measure different things: LMArena is people voting blind between two answers, the others are test scores. Click any row for the source.
Replaced by Qwen1.5-7B; it no longer leads any category among current models.
Where this model places in the top three of today's models, on category leaderboards and on our core tests.