Evidence. Only independent results: the Epoch AI benchmarking hub, which runs evaluations itself or collects them from independent leaderboards. Results a lab reported about its own model (technical reports, system cards) are removed: 564 of them.
Core tests. Six areas, shown on every profile: reasoning & knowledge (GPQA Diamond, Humanity's Last Exam), maths (FrontierMath), coding (SWE-bench Verified, Terminal-Bench), agents (APEX-Agents), novel problems (ARC-AGI-2) and long tasks (METR time horizon). About forty further independent tests also inform the score, so older and less-tested models can still be placed.
Model. An item response model estimates one general ability per model from every test it has taken, allowing for how hard and how discriminating each test is. It is a single latent score, not an average of the six areas. GPT-4 (March 2023) = 100; 50 points is one unit of ability.
Uncertainty. Each score has a 90% band from the fit, measured against GPT-4. It assumes the test parameters are right, so treat it as a lower bound on the real uncertainty. Rated: core results in three areas (or core results plus anchors across three). Provisional: evidence in two areas.
Estimates. Models Epoch has not tested yet get an estimate from their Artificial Analysis score, calibrated on 25 models that have both (R² 0.899). Estimates carry a 90% prediction interval, are marked ~, and are only made inside the calibrated range (AA 11–58).
World affordability. On the World page, a $20-a-month plan is set against the median household living standard per person from the latest World Bank household survey for each country, revalued from 2021 purchasing-power dollars to recent US dollars (PPP × local inflation since 2021 ÷ official exchange rate), assuming real living standards unchanged since the survey. "Over a tenth" is the share of people below $2,400 a year on the same basis. Surveys older than 2015, and countries whose official exchange rate is not what people can get, are left out. The 10% line is our threshold, not a test of who can pay. The method was reviewed by a second model (GPT-6 Astra) and its corrections applied.
Cross-check. Rankings agree almost perfectly with Epoch's own Capabilities Index, which uses the same evidence. The value here is transparency: every score shows the tests behind it, broken down by area, alongside release dates, prices and sources.