GPT-6.1 Sol leads Qwen3.8-Max by 27 points, ahead on 2 of 2 shared core tests.
Test by testindependent results
GPT-6.1 SolQwen3.8-Max
Only tests that at least two of these models have taken. Bold rows are the core tests shown on every profile; all of these results inform the index score. Task length (METR) is listed below the chart because it is a duration, not a percentage.