Apprentice-level cyber task success for the best models rose from under 9% (late 2023) to 50% (2025)
UK AISI · Dec 2025 · source ↗Before
8 monthsNow
4.7 months Doubling time of the length of cyber tasks AI can complete autonomously, estimated at 8 months in Nov 2025 and revised to 4.7 months in Feb 2026
UK AISI · May 2026 · source ↗Expert-level CTF tasks: no model could complete any before April 2025; Claude Mythos Preview succeeded 73% of the time by April 2026
UK AISI · Apr 2026 · source ↗Average steps completed on AISI's 32-step 'The Last Ones' simulated network attack at a 10M-token budget, GPT-4o (Aug 2024) vs Opus 4.6 (Feb 2026)
UK AISI · Mar 2026 · source ↗Success on simplified self-replication evaluations rose from under 5% to over 60% in about two years
UK AISI · Dec 2025 · source ↗~40x
Expert effort to find a biological-misuse jailbreak rose about 40x between two models released six months apart (roughly 10 minutes to over 7 hours)
UK AISI · Dec 2025 · source ↗Open-weight models' lag behind the closed frontier on cyber capability narrowed from 6-10 months (2025) to 4-7 months (mid-2026)
UK AISI · Jul 2026 · source ↗Share of real-world software engineering tasks completed by leading models: zero at start of 2024, ~40% early 2025, over 60% by Oct 2025
International AI Safety Report · Oct 2025 · source ↗OpenAI's o3 outperformed this share of domain experts at troubleshooting virology lab protocols
International AI Safety Report · Feb 2026 · source ↗