> How fast AI capability has climbed since ChatGPT: the frontier over time on one scale, the length of task a model can finish (doubling about every 129 days since 2023, per METR), and the benchmarks models have outgrown.

Capability over time

# The Curve

How long a task an AI agent can finish on its own, and how fast the tests built to measure AI are being outgrown.

4.0min

GPT-4, March 2023

16h+

Claude Mythos Preview, Apr 2026

129days

to double, since 2023

~7×

growth per year

## Longest task an agent finishes on its own · METR 50% time horizon · log scale

_Chart: Length of task AI agents can complete, log scale, since 2019_

METR times tasks on skilled people, then measures which ones an agent completes half the time. Beyond 16 hours its current tasks can no longer measure reliably, so the top reads "16 h or more".

## Tests built to last years fall in months · 30 tests · record score over time

All · 30

Still open · 14

Outgrown · 10

Replaced · 4

Gone quiet · 2

0%

100%

OSWorld 2.0

since Jun 2026

Still open

31.4%

Claude Opus 5

FrontierCode

since Jun 2026

Still open

54.6%

Claude Opus 5.5

DeepSWE

since May 2026

Still open

74.1%

GPT-6 Astra

ARC-AGI-3

since Mar 2026

Still open

62.7%

GPT-6 Astra

ProofBench

since Jan 2026

Outgrown

100%

Claude Fable 5.1

APEX-Agents

since Jan 2026

Still open

75.5%

Claude Sonnet 5.5

Chess puzzles

since Dec 2025

Still open

72%

GPT-6 Astra

Terminal-Bench 2.0

since Nov 2025

Replaced

84.7%

GPT-5.5 + NexAU-AHE

Remote Labor Index

since Oct 2025

Still open

20.8%

GPT-6 Astra

GDPval

since Sep 2025

Still open

49.7%

GPT-5.2

SWE-bench Pro (public set)

since Sep 2025

Still open

61.5%

Meta Muse Spark 1.1

SimpleQA Verified

since Sep 2025

Still open

75.6%

GPT-6 Astra

FrontierMath Tier 4

since Jul 2025

Outgrown

97.6%

GPT-6 Astra

ARC-AGI-2

since Mar 2025

Outgrown

95%

GPT-6 Astra

Humanity's Last Exam

since Jan 2025

Still open

54.8%

GPT-6 Astra

WeirdML

since Jan 2025

Replaced

93.6%

GPT-6 Astra

Aider polyglot

since Dec 2024

Gone quiet

88%

GPT-5

AIME (competition maths)

since Dec 2024

Outgrown

100%

GPT-5.5 Pro, pre-release

TheAgentCompany

since Dec 2024

Gone quiet

42.9%

TTE-MatrixAgent + Deepseek-V3.2

BALROG

since Nov 2024

Still open

68.3%

GPT-6 Astra

FrontierMath (Tiers 1-3)

since Nov 2024

Outgrown

93.7%

GPT-6 Astra

SimpleBench

since Oct 2024

Still open

81.9%

Claude Fable 5

Cybench

since Aug 2024

Outgrown

93%

Claude Opus 4.6

SWE-bench Verified

since Aug 2024

Still open

83.5%

Claude Opus 4.7

OSWorld / OSWorld-Verified

since Apr 2024

Replaced

90.2%

Intelligence-Indeed Agent [Verified]

GPQA Diamond

since Nov 2023

Outgrown

95.8%

GPT-6 Astra

GSM8K

since Oct 2021

Outgrown

94.5%

DS-Coder-V2-Instruct

MATH (level 5)

since Mar 2021

Outgrown

98.1%

GPT-5

MMLU

since Sep 2020

Outgrown

92.3%

OpenAI o1

ARC-AGI-1

since Nov 2019

Replaced

98.5%

Claude Fable 5

Each strip darkens as the best score climbs; solid ink means the field has outgrown the test. Open any row for its record history. Outgrown: the record has passed 90%. Replaced: a newer version exists. Gone quiet: no new record in a year.

## Capability firsts · 59 moments a new ability appeared

_Chart: Capability firsts by kind since 2023_

Sep 2026

59 of 59 · Medicine & biology

### AI-found target and AI-made molecule reach Phase III

Insilico announced the first patient dosed in a 52-week Phase III trial of rentosertib for idiopathic pulmonary fibrosis. It describes the drug as the first candidate with both an AI-identified target and an AI-generated molecule to reach Phase III, after positive Phase IIa results in Nature Medicine in 2025.

IM

Rentosertib (ISM001-055)

[news-medical.net ↗](https://www.news-medical.net/news/20260910/Generative-AI-driven-drug-Rentosertib-enters-Phase-III-trial-for-idiopathic-pulmonary-fibrosis.aspx)

---
Source: https://themodelindex.org/curve/ · The Model Index · data as of 2026-10-01
