The Model Index
Up to four models

Compare

Side by side on the same independent tests, with price and context. The link in your address bar shares this exact comparison.

Next to GLM-5.3
Presets
GLM-5.3
171155–188
WeightsOpen
ReleasedAug 2026
Price in / out$1.4 / $4.4
Context1.05M
Core tests taken5 of 9
Qwen3.8-Max
170154–187
WeightsClosed
ReleasedAug 2026
Price in / out–
Context–
Core tests taken2 of 9

GLM-5.3 and Qwen3.8-Max are too close to call (1 points apart, within the uncertainty), ahead on 0 of 2 shared core tests.

Test by testindependent results

GLM-5.3Qwen3.8-Max
0%25%50%75%100%REASONING & KNOWLEDGEGPQA Diamond (Epoch AI)GPQA Diamond (Epoch AI)SimpleQA VerifiedSimpleQA VerifiedMATHSFrontierMath (tiers 1-3)FrontierMath (tiers 1-3)AIME-style maths (OTIS mock)AIME-style maths (OTIS mock)Chess puzzlesChess puzzlesFrontierMath (tier 4)FrontierMath (tier 4)CODINGDeepSWEDeepSWEFrontierSWEFrontierSWEAGENTSDTBenchDTBenchLMCALMCAMystery game puzzlesMystery game puzzles

Only tests that at least two of these models have taken. Bold rows are the core tests shown on every profile; all of these results inform the index score. Task length (METR) is listed below the chart because it is a duration, not a percentage.