Up to four models
Compare Side by side on the same independent tests, with price and context. The link in your address bar shares this exact comparison.
Next to Claude Fable 5.1 Cheapest at this level Claude Sonnet 5.5 Best open model, measured Hy4 preview Closest rival, another lab GPT-6.1 Sol Stronger for the same price GPT-6 Astra The version it replaced Claude Fable 5 Presets Best open vs best closed (measured) Today vs GPT-4 (2023) Top three labs US vs China, best of each
Claude Fable 5.1 191 175–208
Weights Closed
Released Sep 2026
Price in / out $10 / $50
Context 1M
Core tests taken 6 of 9
GPT-6 Luna 169 152–185
Weights Closed
Released Sep 2026
Price in / out $0.10 / $0.50
Context 1.05M
Core tests taken 6 of 9
Claude Fable 5.1 leads GPT-6 Luna by 23 points, ahead on 5 of 5 shared core tests, at 100× the price.
Test by testindependent results Claude Fable 5.1 GPT-6 Luna0% 25% 50% 75% 100% REASONING & KNOWLEDGE LiveBench Data AnalysisLiveBench Data Analysis LiveBench IFLiveBench IF LiveBench LanguageLiveBench Language LiveBench ReasoningLiveBench Reasoning SimpleQA VerifiedSimpleQA Verified MATHS FrontierMath (tiers 1-3)FrontierMath (tiers 1-3) AIME-style maths (OTIS mock)AIME-style maths (OTIS mock) LiveBench MathematicsLiveBench Mathematics ProofBench v1.1ProofBench v1.1 Chess puzzlesChess puzzles FrontierMath (tier 4)FrontierMath (tier 4) CODING Terminal-Bench 4.0Terminal-Bench 4.0 LiveBench Agentic CodingLiveBench Agentic Coding LiveBench CodingLiveBench Coding Code MigrationCode Migration CyberBench v1.1CyberBench v1.1 IOIIOI ProgramBenchProgramBench Terminal-Bench 2.1Terminal-Bench 2.1 Vibe Code Bench v1.1Vibe Code Bench v1.1 AGENTS APEX-AgentsAPEX-Agents Terminal-Bench ScienceTerminal-Bench Science DTBenchDTBench Furniture assemblyFurniture assembly LMCALMCA Mystery game puzzlesMystery game puzzles PROFESSIONAL WORK Finance Agent (v2)Finance Agent (v2) EMBEMB Harvey's Legal Agent BenchmarkHarvey's Legal Agent Benchmark Legal Research BenchLegal Research Bench MedCodeMedCode MedScribeMedScribe Public Benefits Bench v1.1Public Benefits Bench v1.1 SAGESAGE Tax Agent BenchTax Agent Bench NOVEL PROBLEMS ARC-AGI-2ARC-AGI-2 ARC-AGI-1ARC-AGI-1 MysteryMechanismMysteryMechanism Only tests that at least two of these models have taken. Bold rows are the core tests shown on every profile; all of these results inform the index score. Task length (METR) is listed below the chart because it is a duration, not a percentage.