The Model Index
Up to four models

Compare

Side by side on the same independent tests, with price and context. The link in your address bar shares this exact comparison.

Next to GPT-6 Astra
Presets
GPT-6 Astra
199183–216
WeightsClosed
ReleasedSep 2026
Price in / out$10 / $50
Context1.05M
Core tests taken7 of 9
Claude Opus 5.5
199182–215
WeightsClosed
ReleasedSep 2026
Price in / out$4.0 / $20
Context1M
Core tests taken6 of 9

GPT-6 Astra and Claude Opus 5.5 are too close to call (1 points apart, within the uncertainty), ahead on 3 of 6 shared core tests, at 2.5× the price.

Test by testindependent results

GPT-6 AstraClaude Opus 5.5
0%25%50%75%100%REASONING & KNOWLEDGEGPQA Diamond (Epoch AI)GPQA Diamond (Epoch AI)LiveBench Data AnalysisLiveBench Data AnalysisLiveBench IFLiveBench IFLiveBench LanguageLiveBench LanguageLiveBench ReasoningLiveBench ReasoningEBR-benchEBR-benchSimpleQA VerifiedSimpleQA VerifiedMATHSFrontierMath (tiers 1-3)FrontierMath (tiers 1-3)AIME-style maths (OTIS mock)AIME-style maths (OTIS mock)LiveBench MathematicsLiveBench MathematicsProofBench v1.1ProofBench v1.1FrontierMath (tier 4)FrontierMath (tier 4)CODINGTerminal-Bench 4.0Terminal-Bench 4.0LiveBench Agentic CodingLiveBench Agentic CodingLiveBench CodingLiveBench CodingCode MigrationCode MigrationCyberBench v1.1CyberBench v1.1IOIIOIProgramBenchProgramBenchTerminal-Bench 2.1Terminal-Bench 2.1Vibe Code Bench 1-100Vibe Code Bench 1-100Vibe Code Bench v1.1Vibe Code Bench v1.1FrontierSWEFrontierSWEAGENTSAPEX-AgentsAPEX-AgentsBioMysteryBenchBioMysteryBenchSRE BenchSRE BenchTerminal-Bench ScienceTerminal-Bench ScienceDTBenchDTBenchFurniture assemblyFurniture assemblyLMCALMCAMystery game puzzlesMystery game puzzlesPROFESSIONAL WORKFinance Agent (v2)Finance Agent (v2)EMBEMBHarvey's Legal Agent BenchmarkHarvey's Legal Agent BenchmarkLegal Research BenchLegal Research BenchMedCodeMedCodeMedScribeMedScribeSAGESAGETax Agent BenchTax Agent BenchNOVEL PROBLEMSARC-AGI-2ARC-AGI-2ARC-AGI-1ARC-AGI-1MysteryMechanismMysteryMechanism

Only tests that at least two of these models have taken. Bold rows are the core tests shown on every profile; all of these results inform the index score. Task length (METR) is listed below the chart because it is a duration, not a percentage.