The Model Index
Up to four models

Compare

Side by side on the same third-party tests, with price and context. The link in your address bar shares this exact comparison.

Next to GLM-5.3
Presets
GLM-5.3
171155–188
WeightsOpen
ReleasedAug 2026
Price in / out$0.070 / $7.0
Context1.05M
Core tests taken5 of 9
Kimi K3
170154–187
WeightsOpen
ReleasedJul 2026
Price in / out$0.99 / $14
Context1.05M
Core tests taken6 of 9

GLM-5.3 and Kimi K3 are too close to call (1 points apart, within the uncertainty), ahead on 3 of 5 shared core tests, and costs 2.4× less.

Test by testthird-party results

GLM-5.3Kimi K3
0%25%50%75%100%REASONING & KNOWLEDGEGPQA Diamond (Epoch AI)GPQA Diamond (Epoch AI)LiveBench Data AnalysisLiveBench Data AnalysisLiveBench IFLiveBench IFLiveBench LanguageLiveBench LanguageLiveBench ReasoningLiveBench ReasoningGPQA Diamond (Vals AI)GPQA Diamond (Vals AI)MMLU ProMMLU ProSimpleQA VerifiedSimpleQA VerifiedMATHSFrontierMath (tiers 1-3)FrontierMath (tiers 1-3)AIME-style maths (OTIS mock)AIME-style maths (OTIS mock)LiveBench MathematicsLiveBench MathematicsProofBench v1.1ProofBench v1.1Chess puzzlesChess puzzlesFrontierMath (tier 4)FrontierMath (tier 4)CODINGTerminal-Bench 4.0Terminal-Bench 4.0LiveBench Agentic CodingLiveBench Agentic CodingLiveBench CodingLiveBench CodingCode MigrationCode MigrationCyberBench v1.1CyberBench v1.1IOIIOILiveCodeBenchLiveCodeBenchProgramBenchProgramBenchSWE-benchSWE-benchTerminal-Bench 2.1Terminal-Bench 2.1Vibe Code Bench 1-100Vibe Code Bench 1-100Vibe Code Bench v1.1Vibe Code Bench v1.1DeepSWEDeepSWEFrontierSWEFrontierSWEWeirdMLWeirdMLAGENTSAPEX-AgentsAPEX-AgentsTerminal-Bench ScienceTerminal-Bench ScienceDTBenchDTBenchLMCALMCAMystery game puzzlesMystery game puzzlesPROFESSIONAL WORKFinance Agent (v2)Finance Agent (v2)EMBEMBHarvey's Legal Agent BenchmarkHarvey's Legal Agent BenchmarkLegalBenchLegalBenchLegal Research BenchLegal Research BenchMedCodeMedCodeMedScribeMedScribePublic Benefits Bench v1.1Public Benefits Bench v1.1Tax Agent BenchTax Agent BenchTaxEval v2TaxEval v2

Tests at least two of these models have taken. Bold rows are the core tests.