The Model Index
Capability over time

The Curve

How long a task an AI agent can finish on its own, and how fast the tests built to measure AI are being outgrown.

4.0minGPT-4, March 2023
16h+Claude Mythos Preview, Apr 2026
~7×growth per year

Longest task an agent finishes on its ownMETR 50% time horizon · log scale

METR times tasks on skilled people, then measures which ones an agent completes half the time. Beyond 16 hours its current tasks can no longer measure reliably, so the top reads "16 h or more".

Tests built to last years fall in months30 tests · record score over time

0%100%

Each strip darkens as the best score climbs; solid ink means the field has outgrown the test. Open any row for its record history. Outgrown: the record has passed 90%. Replaced: a newer version exists. Gone quiet: no new record in a year.

Capability firsts59 moments a new ability appeared

Sep 202659 of 59 · Medicine & biology

AI-found target and AI-made molecule reach Phase III

Insilico announced the first patient dosed in a 52-week Phase III trial of rentosertib for idiopathic pulmonary fibrosis. It describes the drug as the first candidate with both an AI-identified target and an AI-generated molecule to reach Phase III, after positive Phase IIa results in Nature Medicine in 2025.

Rentosertib (ISM001-055)news-medical.net ↗