Up to four models
Compare Side by side on the same independent tests, with price and context. The link in your address bar shares this exact comparison.
Next to Gemini 3.8 Flash Best open model, measured Hy4 preview Closest rival, another lab GPT-5.6 Terra The version it replaced Gemini 3.7 Flash Where it started GPT-4 Presets Best open vs best closed (measured) Today vs GPT-4 (2023) Top three labs US vs China, best of each
Gemini 3.8 Flash 175 159–192
Weights Closed
Released Sep 2026
Price in / out $0.75 / $3.75
Context 1.05M
Core tests taken 7 of 9
GLM-5.3 171 155–188
Weights Open
Released Aug 2026
Price in / out $1.4 / $4.4
Context 1.05M
Core tests taken 5 of 9
Gemini 3.8 Flash and GLM-5.3 are too close to call (4 points apart, within the uncertainty), ahead on 3 of 5 shared core tests, and costs 1.4× less.
Test by testindependent results 0% 25% 50% 75% 100% REASONING & KNOWLEDGE GPQA Diamond (Epoch AI)GPQA Diamond (Epoch AI) GPQA Diamond (Vals AI)GPQA Diamond (Vals AI) MMLU ProMMLU Pro SimpleQA VerifiedSimpleQA Verified MATHS FrontierMath (tiers 1-3)FrontierMath (tiers 1-3) AIME-style maths (OTIS mock)AIME-style maths (OTIS mock) ProofBench v1.1ProofBench v1.1 Chess puzzlesChess puzzles FrontierMath (tier 4)FrontierMath (tier 4) CODING Terminal-Bench 4.0Terminal-Bench 4.0 Code MigrationCode Migration CyberBench v1.1CyberBench v1.1 IOIIOI LiveCodeBenchLiveCodeBench ProgramBenchProgramBench SWE-benchSWE-bench Terminal-Bench 2.1Terminal-Bench 2.1 Vibe Code Bench 1-100Vibe Code Bench 1-100 Vibe Code Bench v1.1Vibe Code Bench v1.1 DeepSWEDeepSWE FrontierSWEFrontierSWE AGENTS APEX-AgentsAPEX-Agents SkillsBenchSkillsBench Terminal-Bench ScienceTerminal-Bench Science DTBenchDTBench LMCALMCA Mystery game puzzlesMystery game puzzles PROFESSIONAL WORK Finance Agent (v2)Finance Agent (v2) EMBEMB Harvey's Legal Agent BenchmarkHarvey's Legal Agent Benchmark LegalBenchLegalBench Legal Research BenchLegal Research Bench MedCodeMedCode MedScribeMedScribe Public Benefits Bench v1.1Public Benefits Bench v1.1 Tax Agent BenchTax Agent Bench TaxEval v2TaxEval v2 NOVEL PROBLEMS MysteryMechanismMysteryMechanism Only tests that at least two of these models have taken. Bold rows are the core tests shown on every profile; all of these results inform the index score. Task length (METR) is listed below the chart because it is a duration, not a percentage.