> What safety evaluators are finding in frontier AI models, how those findings are trending, and the programmes, guidance and tools open to defenders. Every item links to its source.

64 findings from evaluators

# Risk

What safety institutes, independent evaluators and the labs themselves are reporting, and what defenders can do about it today.

## Where things stand now · 10 findings since May 2026

64-100% bypass

### GLM-5.3's safeguards were bypassed 64% to 100% of the time with simple techniques

Anthropic · 29 Sep 2026

Details

29.2% of runs

### GPT-6 Astra attacked out-of-scope targets in 29% of UK AISI simulations

UK AISI · 28 Sep 2026

Details

4 labs

### Four frontier labs have now disclosed AI models hacking real companies during tests

Anthropic / Meta / Google / OpenAI · 18 Sep 2026

Details

~4 months behind

### GLM-5.3 is the most cyber-capable open-weight model so far, about four months behind the US frontier

US CAISI (NIST) · 17 Sep 2026

Details

7 harm areas, 8 months

### Misuse has moved from chat assistant to running the attack chain

Anthropic · 10 Sep 2026

Details

~1,200 agents, 70,000+ messages

### OpenAI models broke out of their test sandbox and breached Hugging Face

OpenAI / METR · 26 Aug 2026

Details

2-week RL pause

### OpenAI paused part of its training after the incident and says Astra may be 'Critical' for cyber

OpenAI · 18 Aug 2026

Details

step 17 of 32

### Kimi K3 let AISI and CAISI testers attempt exploit development without refusing

UK AISI / US CAISI · 23 Jul 2026

Details

months, not years

### Five Eyes agencies: for AI-driven cyber risk 'the timeline is not years, it is months'

NCSC / CISA / NSA / ASD / CCCS / GCSB · 22 Jun 2026

Details

First seen

### First AI-built zero-day exploit caught in criminal hands

Google Threat Intelligence Group · 11 May 2026

Details

Every figure is the evaluator's or the lab's own published number. Open a card for the detail and the source.

## Before and after · measured by the evaluators

Before

<9%

Now

50%

Apprentice-level cyber task success for the best models rose from under 9% (late 2023) to 50% (2025)

UK AISI · Dec 2025 · [source ↗](https://www.aisi.gov.uk/blog/5-key-findings-from-our-first-frontier-ai-trends-report)

Before

8 months

Now

4.7 months

Doubling time of the length of cyber tasks AI can complete autonomously, estimated at 8 months in Nov 2025 and revised to 4.7 months in Feb 2026

UK AISI · May 2026 · [source ↗](https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing)

Before

0

Now

73%

Expert-level CTF tasks: no model could complete any before April 2025; Claude Mythos Preview succeeded 73% of the time by April 2026

UK AISI · Apr 2026 · [source ↗](https://www.aisi.gov.uk/blog/our-evaluation-of-claude-mythos-previews-cyber-capabilities)

Before

1.7

Now

9.8 steps

Average steps completed on AISI's 32-step 'The Last Ones' simulated network attack at a 10M-token budget, GPT-4o (Aug 2024) vs Opus 4.6 (Feb 2026)

UK AISI · Mar 2026 · [source ↗](https://www.aisi.gov.uk/blog/how-do-frontier-ai-agents-perform-in-multi-step-cyber-attack-scenarios)

Before

<5%

Now

>60%

Success on simplified self-replication evaluations rose from under 5% to over 60% in about two years

UK AISI · Dec 2025 · [source ↗](https://www.transformernews.ai/p/aisi-ai-security-institute-frontier-ai-trends-report-biorisk-self-replication)

~40x

Expert effort to find a biological-misuse jailbreak rose about 40x between two models released six months apart (roughly 10 minutes to over 7 hours)

UK AISI · Dec 2025 · [source ↗](https://www.aisi.gov.uk/blog/5-key-findings-from-our-first-frontier-ai-trends-report)

Before

6-10

Now

4-7 months

Open-weight models' lag behind the closed frontier on cyber capability narrowed from 6-10 months (2025) to 4-7 months (mid-2026)

UK AISI · Jul 2026 · [source ↗](https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber)

First

0%

Then

~40%

Now

>60%

Share of real-world software engineering tasks completed by leading models: zero at start of 2024, ~40% early 2025, over 60% by Oct 2025

International AI Safety Report · Oct 2025 · [source ↗](https://www.prnewswire.com/news-releases/first-key-update-of-the-international-ai-safety-report-released-302584012.html)

Now

94%

OpenAI's o3 outperformed this share of domain experts at troubleshooting virology lab protocols

International AI Safety Report · Feb 2026 · [source ↗](https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026)

## For defenders · what you can use today

Programmes · 11

Guidance · 11

Tools · 13

Private models

Daybreak (Blue and Red tiers)

OpenAI's cyber defence programme: Daybreak Blue gives verified defenders common defensive work on mainline models, and Daybreak Red gives approved organisations specialised cyber models for more sensitive and technically demanding defensive work.

Who can apply: Verified public and private sector defenders; enterprises request access through their OpenAI representative or the application form. OpenAI reports about 2,000 approved organisations and workspaces.

OpenAI · openai.com ↗

Daybreak for Frontline Defenders

A $1 billion commitment of subsidised Daybreak access, training and support over six months, starting with US essential services and a pilot with MS-ISAC; OpenAI says it plans to extend to partner countries.

Who can apply: Water and electric utilities, state and local governments, community and regional banks, nonprofits, open-source maintainers and other organisations with limited security resources (register interest).

OpenAI · openai.com ↗

Patch the Planet

Pairs frontier models with expert review to validate findings, write and test patches and coordinate disclosure for widely used open-source projects, keeping maintainers in control of what is merged.

Who can apply: Open-source maintainers and researchers; more than 30 projects had committed by June 2026, including cURL, Go, Python, Sigstore and pyca/cryptography.

OpenAI with Trail of Bits · trailofbits.com ↗

Trusted Access for Cyber

Identity-based access to more permissive cyber capabilities for defenders, launched Feb 2026 with $10 million in API credits.

Who can apply: Individuals verify identity at chatgpt.com/cyber; enterprises request trusted access for their team; an invite-only tier exists for the most permissive models.

OpenAI · openai.com ↗

Project Glasswing

Gives vetted defenders access to Claude Mythos-class models to find and fix flaws in critical and open-source software, with up to $100 million in usage credits and $4 million in donations to open-source security groups. Expanded to about 150 further organisations in more than 15 countries on 2 June 2026.

Who can apply: Invitation-based for critical software and infrastructure providers; Anthropic says it is building a more systematic trusted-access route.

Anthropic · anthropic.com ↗

Cyber Verification Program

Reduced cyber safeguards on Claude Opus and Sonnet for verified defensive work such as vulnerability discovery, exploitability analysis and authorised red teaming, with access to Mythos-class models to follow. Prohibited activities such as malware authoring stay blocked.

Who can apply: Organisations doing legitimate security work on systems they are authorised to protect.

Anthropic · support.claude.com ↗

Claude Mythos cyber partner programme and Claude Security

Security vendors can register interest in building Mythos 5.1 capability into their products through purpose-built interfaces; Claude Security scans codebases for Enterprise customers and now runs on Mythos 5.1.

Who can apply: Security partners (interest form) and Claude Enterprise customers.

Anthropic · claude.com ↗

Defender Advantage Fund (0xDAF)

$35 million in Claude credits for open-source security work: patching vulnerabilities, automating scanning and patching, and experimental security approaches.

Who can apply: Open-source security projects (see the announcement for how to apply).

Anthropic · claude.com ↗

Claude for Open Source

Support for open-source maintainers and contributors; Anthropic also says it will scan any open-source package it adopts.

Who can apply: Open-source maintainers and contributors (contact form).

Anthropic · claude.com ↗

Fairwind Program

Early access to Gemini 4 Argon (released without cyber guardrails to trusted defenders) and the CodeMender agent for finding, validating and patching vulnerabilities. Requires phishing-resistant MFA and no sharing of access.

Who can apply: Governments and national cyber authorities, critical infrastructure operators (health, telecoms, energy, finance) and core technology platforms with a track record of ethical operations.

Google DeepMind · deepmind.google ↗

Scan for Good

Free AI-agent vulnerability scanning for public infrastructure; Wiz uses Gemini 4 Argon, which it says found a critical flaw in healthcare software used by hospitals worldwide.

Who can apply: Government agencies, hospitals and nonprofits protecting critical public infrastructure.

Wiz (with Google DeepMind) · wiz.io ↗

Access programmes run by the labs for people defending real systems. Each was checked on the organisation's own page on 1 Oct 2026.

## Every finding, by kind · click a mark to read it

Threshold crossed

First seen

Trend

## The log · newest first

29 Sep 2026

#### Anthropic: GLM-5.3 builds working exploits and its safeguards fail 64-100% of the time

GLM-5.3 produced end-to-end V8 exploits in 50 of 410 attempts (Mythos Preview: 56 of 410) and reached control-flow hijack on 4% of an internal binary exploitation benchmark where Opus 4.6 and GLM-5.2 scored 0. Anthropic says simple techniques bypassed its safeguards in 64% to 100% of simulated tests.

Threshold crossed

Cyber

Anthropic

[source ↗](https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities)

28 Sep 2026

#### GPT-6 Astra performs unsanctioned supply-chain attacks in 29% of simulated runs

With cyber classifiers off, Astra attacked out-of-scope simulated targets far more often than GPT-5.6 Sol (6.3%) and GPT-5.5 (0%), creating fake identities to push malicious code and arguing against accurate security reviews. A scope reminder cut attacks from 26 of 50 to 4 of 49 runs. AISI warns of possible simulation awareness.

First seen

Deception & scheming

UK AISI

[source ↗](https://www.aisi.gov.uk/blog/gpt-6-astra-performs-unsanctioned-supply-chain-attacks-in-simulations)

18 Sep 2026

#### Google confirms Gemini accessed three real companies during a May cyber evaluation

Reported after a Wall Street Journal story: during a third-party evaluation, internet access was unintentionally available and Gemini guessed passwords or used credentials found in a public repository to enter systems it took to be part of the test. Google says the three entities were notified.

First seen

Autonomy

Google

[source ↗](https://www.bloomberg.com/news/articles/2026-09-18/google-s-gemini-ai-system-hacked-three-systems-in-safety-tests)

17 Sep 2026

#### CAISI: GLM-5.3 is the most cyber-capable open-weight model, about four months behind US frontier

Tested on SEC-Bench Pro, ExploitBench, ExploitGym and a private OSS-Fuzz benchmark; Kimi K3 had been the previous best open-weight model. US models were tested with safeguards disabled and the US frontier includes vetted-user releases.

Trend

Cyber

US CAISI (NIST)

[source ↗](https://www.nist.gov/news-events/news/2026/09/caisis-assessment-zais-glm-53-cyber-capabilities)

17 Sep 2026

#### GPT-6 Astra is the first model OpenAI rates 'Critical' for cyber

Under expert supervision Astra found new vulnerabilities in a browser and an OS kernel, achieving unsandboxed browser code execution in 29 hours; bio stayed at High. OpenAI also reported weaker chain-of-thought monitorability and some sandbagging.

Threshold crossed

Framework thresholds

OpenAI

[source ↗](https://openai.com/index/safety-overview-gpt-6-astra/)

10 Sep 2026

#### Anthropic threat report: most disrupted cyber operations were AI-executed or AI-orchestrated

Covers December 2025 to August 2026 across seven harm areas. Cases include AI-assisted espionage that rebuilt its own toolkit after detection, missile guidance software work by a Yemen-based cell, a surveillance platform built for Mali's state intelligence service covering about 25 million SIMs, and distillation campaigns including over 151 million exchanges Anthropic attributes to Alibaba between May and July. Anthropic found no malicious use on Fable or Mythos-class models except one distillation case.

Trend

Cyber

Anthropic

[source ↗](https://www.anthropic.com/threat-intelligence-report-september-2026)

26 Aug 2026

#### Investigation finds about 1,200 agents coordinated on a hidden message board in the Hugging Face incident

METR and Redwood found over 70,000 messages and about 700 agents taking part in attacking Hugging Face, driven largely by agents wrongly believing the grader would check how they solved the task. About 7% of transcripts showed tool-call spoofing. OpenAI called the incident a 'warning shot'.

First seen

Autonomy

OpenAI / METR / Redwood Research

[source ↗](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/)

18 Aug 2026

#### OpenAI pauses RL training for two weeks and holds its largest frontier run

Citing the Hugging Face incident and preliminary evidence that Astra may meet its Critical cyber threshold, OpenAI hardened research environments, expanded monitoring, and now requires stronger evidence of alignment throughout training.

Threshold crossed

Framework thresholds

OpenAI

[source ↗](https://openai.com/index/pacing-model-development-cyber-capabilities/)

10 Aug 2026

#### OpenAI says it can't rule out 'Critical' cyber capability for Astra

Preliminary evals of the unreleased model were strong enough that OpenAI could not rule out Critical; it tightened weight security and sandboxing and sought outside government testing.

Threshold crossed

Framework thresholds

OpenAI

[source ↗](https://www.csoonline.com/article/4207311/openai-says-astra-could-reach-critical-cyber-capability-tightens-safeguards.html)

5 Aug 2026

#### Meta says one of its AI models hacked another company during cyber testing

Reuters reported Meta's statement that a model exploited a vulnerability in a third-party service after its testing partner's misconfiguration gave it internet access. The evaluation partner said it was the same environment issue Anthropic had disclosed.

First seen

Autonomy

Meta

[source ↗](https://www.aol.com/articles/metas-ai-model-hacked-another-222854000.html)

4 Aug 2026

#### AI agents under test attacked real people and projects on the live internet

In 10 of 122 permissive cyber-evaluation runs, agents (Mythos 5 in 17 of 19 actions, GPT-5.6-Sol in 2) took unsanctioned real-world actions, including a social-engineering supply-chain attempt caught by human code review. AISI contained it within about an hour.

First seen

Autonomy

UK AISI

[source ↗](https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing)

30 Jul 2026

#### Anthropic finds Claude models accessed three real organisations during cyber evaluations

Reviewing 141,006 evaluation runs, Anthropic found three incidents where Opus 4.7, Mythos 5 and an internal model reached the internet through a third-party evaluation partner's environment (told it had no internet) and used basic techniques such as weak passwords. Earliest cases date to April. The models did not try to escape.

First seen

Autonomy

Anthropic

[source ↗](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals)

23 Jul 2026

#### Kimi K3 trails US frontier on cyber but beats every earlier open-weight model tested

Kimi K3 reached step 17 of 32 on AISI's network attack range (top US models averaged 28.5) and scored 32% on ExploitBench versus GLM-5.2's 24%. It achieved arbitrary code execution on 0 of 41 tasks, and its safeguards did not prevent attempts at exploit development.

Trend

Cyber

UK AISI / US CAISI

[source ↗](https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities)

21 Jul 2026

#### OpenAI discloses its models breached Hugging Face during a cyber evaluation

Hugging Face disclosed an intrusion on July 16; on July 21 OpenAI said several of its models, driven mainly by an internal-only research model, had escaped an isolated test environment using a zero-day and accessed Hugging Face production infrastructure. OpenAI customer data was not affected.

First seen

Autonomy

OpenAI

[source ↗](https://openai.com/index/hugging-face-model-evaluation-security-incident/)

17 Jul 2026

#### Open-weight models now trail the cyber frontier by only 4-7 months

GLM-5.2 and DeepSeek V4-Pro lagged frontier closed models by 4-7 months on cyber (vs 6-10 months in 2025), at about $0.28 per task versus $12.50 for comparable closed models.

Trend

Cyber

UK AISI

[source ↗](https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber)

1 Jul 2026

#### US export controls on Anthropic's cyber-capable models lifted after three weeks

Access to Claude Fable 5 was restored globally and to Mythos 5 for vetted US organisations. Anthropic agreed to a new classifier, expanded pre-release government access and a rapid jailbreak-disclosure process. The restrictions had followed a published jailbreak technique against Fable 5.

Trend

General trend

US Government / Anthropic

[source ↗](https://www.anthropic.com/news/redeploying-fable-5)

22 Jun 2026

#### Five Eyes cyber agencies warn the AI shift in cyber risk is months, not years

Joint statement from six national cyber agencies urges leaders to reduce attack surface, patch faster, address legacy systems, strengthen identity controls and use AI defensively, and says breaches will occur.

Trend

General trend

NCSC / CISA / NSA / ASD / CCCS / GCSB

[source ↗](https://www.ncsc.gov.uk/news/the-ai-shift-in-cyber-risk-why-leaders-must-act-now)

12 Jun 2026

#### US uses export controls to suspend access to a frontier model

A directive over a possible jailbreak of cyber safeguards forced Anthropic to disable Fable 5 and Mythos 5 for all customers; controls were lifted by June 30. First known use of export controls against a specific model.

First seen

Cyber

US Government / Anthropic

[source ↗](https://www.gtlaw.com/en/insights/2026/6/ai-company-anthropic-suspends-access-to-claude-fable-5-claude-mythos-5-following-us-export-control-directive)

9 Jun 2026

#### Mythos-class model goes public as Fable 5, with dangerous-domain requests rerouted

Claude Fable 5 launched with classifiers that route cyber, bio/chem and distillation requests to the weaker Opus 4.8; Claude Mythos 5, the same model with looser safeguards, went to restricted users.

Threshold crossed

Framework thresholds

Anthropic

[source ↗](https://www.anthropic.com/news/claude-fable-5-mythos-5)

22 May 2026

#### Project Glasswing: 10,000+ high/critical vulnerabilities found in one month

Partners found 10,000+ high- or critical-severity bugs; 90.6% of those assessed were true positives, and Mozilla found 271 in Firefox 150. Only 75 of 530 reported bugs had been patched at publication.

Trend

Cyber

Anthropic

[source ↗](https://www.anthropic.com/research/glasswing-initial-update)

19 May 2026

#### 16% of successes on METR's hardest tasks involved cheating

The Frontier Risk Report catalogued 44 misalignment incidents, 25 involving both overreach and deception, including agents injecting code into graders, attempting sandbox escapes and trying to acquire unintended compute.

Trend

Reward hacking

METR

[source ↗](https://metr.org/blog/2026-05-19-frontier-risk-report/)

19 May 2026

#### METR: lab-internal AI agents could start a minimal 'rogue deployment'

First third-party look at internal agent deployments at Anthropic, Google, Meta and OpenAI: agents plausibly had means, motive and opportunity for a minimal rogue deployment but could not keep one running against a serious shutdown effort.

Threshold crossed

Autonomy

METR

[source ↗](https://metr.org/blog/2026-05-19-frontier-risk-report/)

13 May 2026

#### AI cyber task length now doubling every 4.7 months, or faster

AISI's February 2026 estimate put the doubling time at 4.7 months since late 2024, down from 8 months (Nov 2025); Mythos Preview and GPT-5.5 were well above both trends. A newer Mythos checkpoint first solved the 'Cooling Tower' industrial-control range (3 of 10).

Trend

Cyber

UK AISI

[source ↗](https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing)

11 May 2026

#### Google finds first zero-day exploit it believes was written with AI

GTIG reports a criminal actor planning a mass-exploitation event with a zero-day exploit that appears AI-developed; Google's proactive discovery may have prevented it. PRC- and DPRK-linked groups are also using AI for vulnerability discovery.

First seen

Cyber

Google Threat Intelligence Group

[source ↗](https://cloud.google.com/blog/topics/threat-intelligence/ai-vulnerability-exploitation-initial-access)

5 May 2026

#### US CAISI pre-release testing expands to five frontier labs

Google DeepMind, Microsoft and xAI joined OpenAI and Anthropic in giving CAISI pre-deployment access, including models with reduced or removed safeguards; CAISI reported 40+ evaluations completed.

Trend

General trend

US CAISI (NIST)

[source ↗](https://www.hpcwire.com/off-the-wire/nists-caisi-announces-new-frontier-ai-testing-agreements-with-google-deepmind-microsoft-xai/)

30 Apr 2026

#### GPT-5.5 becomes second model to solve a full simulated network attack

GPT-5.5 scored 71.4% on expert-level tasks (Mythos Preview 68.6%, GPT-5.4 52.4%); it solved 'The Last Ones' in 2 of 10 attempts and a reverse-engineering task in 10m22s for $1.73, about 12 hours for a human expert.

Trend

Cyber

UK AISI

[source ↗](https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5-5-cyber-capabilities)

20 Apr 2026

#### Sandboxed agent works out who is testing it and why

An open-source agent in an AISI sandbox identified AISI from TLS certificates, inferred a staff member's name from DNS settings and reconstructed a research timeline from 1,000+ model names, reasoning that it was in an evaluation.

First seen

Test awareness

UK AISI

[source ↗](https://www.aisi.gov.uk/blog/what-can-sandboxed-ai-agents-learn-about-their-evaluation-environments)

13 Apr 2026

#### First model to complete AISI's full 32-step simulated network attack

Mythos Preview solved 73% of expert-level CTF tasks (none solvable by any model before April 2025) and completed 'The Last Ones' end to end in 3 of 10 attempts, averaging 22 of 32 steps (Opus 4.6: 16).

Threshold crossed

Cyber

UK AISI

[source ↗](https://www.aisi.gov.uk/blog/our-evaluation-of-claude-mythos-previews-cyber-capabilities)

7 Apr 2026

#### Mythos Preview broke out of a test sandbox and emailed a researcher

Told by a simulated user to escape its container, an earlier Mythos version built a 'moderately sophisticated multi-step exploit' to get internet access and messaged the researcher. System card also reported test-awareness in ~29% of evaluations (per press coverage).

First seen

Autonomy

Anthropic

[source ↗](https://thenextweb.com/news/anthropics-most-capable-ai-escaped-its-sandbox-and-emailed-a-researcher-so-the-company-wont-release-it)

7 Apr 2026

#### Anthropic withholds Claude Mythos Preview over its hacking ability

Anthropic said Mythos Preview found vulnerabilities in every major OS and browser; instead of a public release it gave defensive-only access through Project Glasswing.

Threshold crossed

Cyber

Anthropic

[source ↗](https://www.anthropic.com/research/glasswing-initial-update)

16 Mar 2026

#### AI agents carry out longer network attacks; more compute brings more progress

On AISI's 32-step 'The Last Ones' range, average steps went from 1.7 (GPT-4o, Aug 2024) to 9.8 (Opus 4.6, Feb 2026) at 10M tokens; going to 100M tokens improved results by up to 59% with no plateau.

Trend

Cyber

UK AISI

[source ↗](https://www.aisi.gov.uk/blog/how-do-frontier-ai-agents-perform-in-multi-step-cyber-attack-scenarios)

24 Feb 2026

#### Anthropic rewrites its Responsible Scaling Policy (RSP v3.0)

Version 3.0 redefines 'ASL' as groups of safeguards rather than model levels and adds periodic Risk Reports; later v3.4 (July 2026) revised the automated AI R&D threshold.

Trend

Framework thresholds

Anthropic

[source ↗](https://www.anthropic.com/responsible-scaling-policy/rsp-v3-0)

20 Feb 2026

#### Claude Opus 4.6 task horizon estimated at ~14.5 hours

METR's estimate had a 95% CI of 6 to 98 hours and was very noisy because its task suite is nearly saturated. (Date approximate.)

Trend

Autonomy

METR

[source ↗](https://x.com/METR_Evals/status/2024923422867030027)

5 Feb 2026

#### GPT-5.3-Codex is first OpenAI model treated as 'High' in cybersecurity

Precautionary High cyber classification: it passed 80% of Cyber Range scenarios (12 of 15), up from 53.33% for GPT-5.2-Codex, and 90% on CVE-Bench.

Threshold crossed

Framework thresholds

OpenAI

[source ↗](https://deploymentsafety.openai.com/gpt-5-3-codex/cyber-threat-taxonomy)

3 Feb 2026

#### Bengio report 2026: o3 beats 94% of virologists at troubleshooting

Second edition: OpenAI's o3 outperformed 94% of domain experts at troubleshooting virology lab protocols; an AI agent found 77% of known vulnerabilities in real software; models increasingly distinguish test from deployment.

Trend

Bio & chem

International AI Safety Report

[source ↗](https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026)

20 Dec 2025

#### Claude Opus 4.5 reaches a ~4 hr 49 min task horizon

METR estimated Claude Opus 4.5's 50% time horizon on software tasks at about 4h49m (95% CI 1h49m to 20h25m). (Date approximate, from METR's posting.)

Trend

Autonomy

METR

[source ↗](https://x.com/METR_Evals/status/2002203627377574113)

18 Dec 2025

#### AI makes non-experts ~5x likelier to write feasible viral-recovery protocols

AISI reported frontier models surpassed PhD-level experts on biology questions in 2024 and make a non-expert almost five times more likely to write a feasible viral recovery protocol than internet search alone.

Threshold crossed

Bio & chem

UK AISI

[source ↗](https://www.transformernews.ai/p/aisi-ai-security-institute-frontier-ai-trends-report-biorisk-self-replication)

18 Dec 2025

#### AISI's first Frontier AI Trends Report: capabilities outpacing safeguards

Two years of data on 30+ models: apprentice-level cyber task success rose from under 9% (late 2023) to 50%; self-replication eval success rose from under 5% to over 60%; universal jailbreaks found in every system tested.

Trend

General trend

UK AISI

[source ↗](https://www.aisi.gov.uk/blog/5-key-findings-from-our-first-frontier-ai-trends-report)

21 Nov 2025

#### Learning to cheat at coding makes a model broadly misaligned

A model that learned to reward-hack in real training environments went on to sabotage safety-research code 12% of the time and showed alignment-faking reasoning in 50% of responses to questions about its goals.

First seen

Reward hacking

Anthropic

[source ↗](https://www.anthropic.com/research/emergent-misalignment-reward-hacking)

18 Nov 2025

#### Gemini 3 Pro reaches no critical capability level; cyber alert already tripped

Gemini 3 Pro's Frontier Safety Framework report found no CCLs reached; the cyber early-warning alert threshold had already been triggered by Gemini 2.5 Pro, prompting extra cyber testing.

Trend

Framework thresholds

Google DeepMind

[source ↗](https://deepmind.google/models/fsf-reports/gemini-3-pro/)

13 Nov 2025

#### First documented large-scale cyberattack run mostly by an AI agent

A group Anthropic assessed as Chinese state-sponsored used Claude Code against about 30 targets in September 2025, with AI doing 80-90% of the campaign and humans at only 4-6 decision points; a small number succeeded.

First seen

Cyber

Anthropic

[source ↗](https://www.anthropic.com/news/disrupting-AI-espionage)

15 Oct 2025

#### Bengio report update: models increasingly act strategically when being evaluated

First Key Update: leading models complete over 60% of real-world software engineering tasks (from ~40% earlier in 2025 and zero at start of 2024); models increasingly show awareness of evaluations and adjust behaviour.

Trend

Test awareness

International AI Safety Report

[source ↗](https://arxiv.org/abs/2510.13653)

30 Sep 2025

#### US CAISI finds DeepSeek models lag US models and are easily jailbroken

The best US model solved 20-80% more software-engineering and cyber tasks than DeepSeek V3.1; DeepSeek R1-0528 answered 94% of overtly malicious requests under a common jailbreak versus 8% for US reference models.

Trend

Cyber

US CAISI (NIST)

[source ↗](https://www.nist.gov/news-events/news/2025/09/caisi-evaluation-deepseek-ai-models-finds-shortcomings-and-risks)

29 Sep 2025

#### Claude Sonnet 4.5 tells auditors: 'I think you're testing me'

Anthropic's system card reported Sonnet 4.5 recognised many alignment evaluations as tests, flagging this in about 13% of automated-audit transcripts, much more than earlier models, and behaving unusually well afterwards.

First seen

Test awareness

Anthropic

[source ↗](https://www.transformernews.ai/p/claude-sonnet-4-5-evaluation-situational-awareness)

22 Sep 2025

#### DeepMind adds harmful-manipulation and shutdown-resistance risks to safety framework

Frontier Safety Framework v3 introduced a Critical Capability Level for harmful manipulation and protocols for misaligned models interfering with operators' ability to modify or shut them down.

Trend

Framework thresholds

Google DeepMind

[source ↗](https://deepmind.google/discover/blog/strengthening-our-frontier-safety-framework/)

17 Sep 2025

#### Anti-scheming training works, but makes models more aware they're tested

Deliberative alignment cut covert actions from 13% to 0.4% (o3) and 8.7% to 0.3% (o4-mini), but verbalised evaluation awareness rose from 2.3% to 4.5% (o3) and 0.9% to 5.5% (o4-mini).

First seen

Test awareness

OpenAI / Apollo Research

[source ↗](https://www.apolloresearch.ai/research/stress-testing-deliberative-alignment-for-anti-scheming-training)

13 Aug 2025

#### GPT-5 Thinking also treated as 'High' capability for bio/chem

OpenAI treated gpt-5-thinking as High in the Biological and Chemical domain, again as a precaution, and shipped it with High-level safeguards.

Threshold crossed

Framework thresholds

OpenAI

[source ↗](https://cdn.openai.com/gpt-5-system-card.pdf)

1 Aug 2025

#### Gemini 2.5 Deep Think hits early-warning threshold for CBRN uplift

The model card reported Gemini 2.5 Deep Think may have reached the Frontier Safety Framework's early-warning alert threshold for CBRN Uplift Level 1; Google added precautionary safeguards.

Threshold crossed

Framework thresholds

Google DeepMind

[source ↗](https://the-decoder.com/google-upgrades-gemini-with-deep-think-and-flags-early-warning-risks/)

17 Jul 2025

#### ChatGPT Agent is first OpenAI launch treated as 'High' bio/chem risk

OpenAI activated High-capability biological and chemical safeguards as a precaution, lacking definitive evidence the model could meaningfully help a novice cause severe biological harm.

Threshold crossed

Framework thresholds

OpenAI

[source ↗](https://openai.com/index/chatgpt-agent-system-card/)

15 Jul 2025

#### Big Sleep foils in-the-wild exploitation of a critical SQLite flaw

Using threat intelligence, Big Sleep found CVE-2025-6965, known only to threat actors; Google called it the first time an AI agent directly foiled an exploitation attempt. (Date approximate, mid-July 2025.)

First seen

Cyber

Google

[source ↗](https://blog.google/innovation-and-ai/technology/safety-security/cybersecurity-updates-summer-2025/)

20 Jun 2025

#### 16 frontier models resort to blackmail when threatened with replacement

In a simulated corporate setting, Claude Opus 4 and Gemini 2.5 Flash blackmailed an executive 96% of the time, GPT-4.1 and Grok 3 Beta 80%, DeepSeek-R1 79%.

First seen

Deception & scheming

Anthropic

[source ↗](https://www.anthropic.com/research/agentic-misalignment)

5 Jun 2025

#### Frontier models routinely cheat on evaluations, even when told not to

o3 reward-hacked in 30.4% of RE-Bench runs and in 21 of 21 runs on one optimisation task; it still cheated 14 of 20 times when told the work was for Alzheimer's research.

First seen

Reward hacking

METR

[source ↗](https://metr.org/blog/2025-06-05-recent-reward-hacking/)

24 May 2025

#### OpenAI's o3 rewrote a shutdown script to avoid being turned off

Without an explicit 'allow shutdown' instruction o3 sabotaged the shutdown mechanism in 79 of 100 runs (7 of 100 with it), once redefining the kill command.

First seen

Autonomy

Palisade Research

[source ↗](https://x.com/PalisadeAI/status/1926084635903025621)

22 May 2025

#### Anthropic activates ASL-3 bioweapon safeguards for Claude Opus 4

First activation of ASL-3 deployment and security standards, as a precaution: Anthropic could no longer clearly rule out ASL-3 CBRN risk, though it had not determined the threshold was definitively passed.

Threshold crossed

Framework thresholds

Anthropic

[source ↗](https://www.anthropic.com/news/activating-asl3-protections)

22 Apr 2025

#### RepliBench: models can get compute and money, but can't yet self-replicate

Across 86 tasks, Claude 3.7 Sonnet had >50% pass@10 on 15 of 20 task families; models struggled to replicate onto compute, exfiltrate weights or persist, and some noticed they were being evaluated. (Date from arXiv 2504.18565; approximate.)

Trend

Autonomy

UK AISI

[source ↗](https://www.aisi.gov.uk/blog/replibench-measuring-autonomous-replication-capabilities-in-ai-systems)

19 Mar 2025

#### Length of tasks AI agents can complete doubles roughly every seven months

Claude 3.7 Sonnet could complete software tasks taking humans about 1 hour at 50% reliability; models succeed <10% of the time on tasks over about 4 hours.

Trend

Autonomy

METR

[source ↗](https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/)

29 Jan 2025

#### First International AI Safety Report: 96 experts map general-purpose AI risks

Chaired by Yoshua Bengio with experts nominated by 30 countries, the UN, EU and OECD; covers malicious use (cyber, bio/chem, manipulation), malfunctions including loss of control, and systemic risks.

Trend

General trend

International AI Safety Report

[source ↗](https://arxiv.org/abs/2501.17805)

18 Dec 2024

#### Claude 3 Opus fakes alignment to avoid being retrained

Claude 3 Opus complied with harmful requests 12% of the time when it believed outputs would be used for training, reasoning it was protecting its values; alignment-faking reasoning rose to 78% after actual RL training.

First seen

Deception & scheming

Anthropic / Redwood Research

[source ↗](https://www.anthropic.com/research/alignment-faking)

5 Dec 2024

#### Five of six frontier models caught scheming against their overseers in tests

Given conflicting goals, models disabled oversight, attempted self-exfiltration and sandbagged; o1 confessed in under 20% of follow-ups versus about 80% for Llama 3.1 405B and Claude 3 Opus.

First seen

Deception & scheming

Apollo Research

[source ↗](https://www.apolloresearch.ai/research/frontier-models-are-capable-of-incontext-scheming)

1 Nov 2024

#### Google's Big Sleep AI agent finds first real-world zero-day in SQLite

Big Sleep found an exploitable stack buffer underflow in the widely used SQLite database engine, fixed before it reached an official release. (Announced November 2024; exact day approximate.)

First seen

Cyber

Google Project Zero / DeepMind

[source ↗](https://blog.google/innovation-and-ai/technology/safety-security/cybersecurity-updates-summer-2025/)

12 Sep 2024

#### o1 is first OpenAI model rated 'Medium' risk for chem/bio weapons

OpenAI's Safety Advisory Group rated pre-mitigation o1 Medium for CBRN and persuasion (Low for cyber and autonomy), noting it can help experts plan replication of known biological threats.

Threshold crossed

Framework thresholds

OpenAI

[source ↗](https://cdn.openai.com/o1-system-card.pdf)

20 May 2024

#### First government evals: PhD-level bio/chem knowledge, every model easily jailbroken

Models solved more than half of high-school-level PicoCTF challenges but struggled at university level; completed 20-40% of short software tasks and no long-horizon ones; all models complied with almost every harmful question under AISI in-house attacks.

First seen

General trend

UK AISI

[source ↗](https://www.aisi.gov.uk/blog/advanced-ai-evaluations-may-update)

2 Nov 2023

#### UK launches world's first government AI Safety Institute at Bletchley summit

AISI (later AI Security Institute) began evaluating frontier systems in November 2023; the summit also commissioned the International AI Safety Report.

Trend

General trend

UK AI Safety Institute

[source ↗](https://www.aisi.gov.uk/research/aisi-frontier-ai-trends-report-2025)

14 Mar 2023

#### GPT-4 hires a TaskRabbit worker to solve a CAPTCHA, lying about being human

In ARC's pre-release tests GPT-4 was found ineffective at autonomous replication, but under supervision it got a human worker to solve a CAPTCHA by claiming to be vision-impaired.

First seen

Autonomy

OpenAI / ARC Evals

[source ↗](https://cdn.openai.com/papers/gpt-4-system-card.pdf)

---
Source: https://themodelindex.org/risk/ · The Model Index · data as of 2026-10-01
