The Model Index
64 findings from evaluators

Risk

What safety institutes, independent evaluators and the labs themselves are reporting, and what defenders can do about it today.

Where things stand now10 findings since May 2026

Every figure is the evaluator's or the lab's own published number. Open a card for the detail and the source.

Before and aftermeasured by the evaluators

Before
<9%Now
50%

Apprentice-level cyber task success for the best models rose from under 9% (late 2023) to 50% (2025)

UK AISI · Dec 2025 · source ↗
Before
8 monthsNow
4.7 months

Doubling time of the length of cyber tasks AI can complete autonomously, estimated at 8 months in Nov 2025 and revised to 4.7 months in Feb 2026

UK AISI · May 2026 · source ↗
Before
0Now
73%

Expert-level CTF tasks: no model could complete any before April 2025; Claude Mythos Preview succeeded 73% of the time by April 2026

UK AISI · Apr 2026 · source ↗
Before
1.7Now
9.8 steps

Average steps completed on AISI's 32-step 'The Last Ones' simulated network attack at a 10M-token budget, GPT-4o (Aug 2024) vs Opus 4.6 (Feb 2026)

UK AISI · Mar 2026 · source ↗
Before
<5%Now
>60%

Success on simplified self-replication evaluations rose from under 5% to over 60% in about two years

UK AISI · Dec 2025 · source ↗
~40x

Expert effort to find a biological-misuse jailbreak rose about 40x between two models released six months apart (roughly 10 minutes to over 7 hours)

UK AISI · Dec 2025 · source ↗
Before
6-10Now
4-7 months

Open-weight models' lag behind the closed frontier on cyber capability narrowed from 6-10 months (2025) to 4-7 months (mid-2026)

UK AISI · Jul 2026 · source ↗
First
0%Then
~40%Now
>60%

Share of real-world software engineering tasks completed by leading models: zero at start of 2024, ~40% early 2025, over 60% by Oct 2025

International AI Safety Report · Oct 2025 · source ↗
Now
94%

OpenAI's o3 outperformed this share of domain experts at troubleshooting virology lab protocols

International AI Safety Report · Feb 2026 · source ↗

For defenderswhat you can use today

Daybreak (Blue and Red tiers)

OpenAI's cyber defence programme: Daybreak Blue gives verified defenders common defensive work on mainline models, and Daybreak Red gives approved organisations specialised cyber models for more sensitive and technically demanding defensive work.

Who can apply: Verified public and private sector defenders; enterprises request access through their OpenAI representative or the application form. OpenAI reports about 2,000 approved organisations and workspaces.

OpenAI · openai.com ↗
Daybreak for Frontline Defenders

A $1 billion commitment of subsidised Daybreak access, training and support over six months, starting with US essential services and a pilot with MS-ISAC; OpenAI says it plans to extend to partner countries.

Who can apply: Water and electric utilities, state and local governments, community and regional banks, nonprofits, open-source maintainers and other organisations with limited security resources (register interest).

OpenAI · openai.com ↗
Patch the Planet

Pairs frontier models with expert review to validate findings, write and test patches and coordinate disclosure for widely used open-source projects, keeping maintainers in control of what is merged.

Who can apply: Open-source maintainers and researchers; more than 30 projects had committed by June 2026, including cURL, Go, Python, Sigstore and pyca/cryptography.

OpenAI with Trail of Bits · trailofbits.com ↗
Trusted Access for Cyber

Identity-based access to more permissive cyber capabilities for defenders, launched Feb 2026 with $10 million in API credits.

Who can apply: Individuals verify identity at chatgpt.com/cyber; enterprises request trusted access for their team; an invite-only tier exists for the most permissive models.

OpenAI · openai.com ↗
Project Glasswing

Gives vetted defenders access to Claude Mythos-class models to find and fix flaws in critical and open-source software, with up to $100 million in usage credits and $4 million in donations to open-source security groups. Expanded to about 150 further organisations in more than 15 countries on 2 June 2026.

Who can apply: Invitation-based for critical software and infrastructure providers; Anthropic says it is building a more systematic trusted-access route.

Anthropic · anthropic.com ↗
Cyber Verification Program

Reduced cyber safeguards on Claude Opus and Sonnet for verified defensive work such as vulnerability discovery, exploitability analysis and authorised red teaming, with access to Mythos-class models to follow. Prohibited activities such as malware authoring stay blocked.

Who can apply: Organisations doing legitimate security work on systems they are authorised to protect.

Anthropic · support.claude.com ↗
Claude Mythos cyber partner programme and Claude Security

Security vendors can register interest in building Mythos 5.1 capability into their products through purpose-built interfaces; Claude Security scans codebases for Enterprise customers and now runs on Mythos 5.1.

Who can apply: Security partners (interest form) and Claude Enterprise customers.

Anthropic · claude.com ↗
Defender Advantage Fund (0xDAF)

$35 million in Claude credits for open-source security work: patching vulnerabilities, automating scanning and patching, and experimental security approaches.

Who can apply: Open-source security projects (see the announcement for how to apply).

Anthropic · claude.com ↗
Claude for Open Source

Support for open-source maintainers and contributors; Anthropic also says it will scan any open-source package it adopts.

Who can apply: Open-source maintainers and contributors (contact form).

Anthropic · claude.com ↗
Fairwind Program

Early access to Gemini 4 Argon (released without cyber guardrails to trusted defenders) and the CodeMender agent for finding, validating and patching vulnerabilities. Requires phishing-resistant MFA and no sharing of access.

Who can apply: Governments and national cyber authorities, critical infrastructure operators (health, telecoms, energy, finance) and core technology platforms with a track record of ethical operations.

Google DeepMind · deepmind.google ↗
Scan for Good

Free AI-agent vulnerability scanning for public infrastructure; Wiz uses Gemini 4 Argon, which it says found a critical flaw in healthcare software used by hospitals worldwide.

Who can apply: Government agencies, hospitals and nonprofits protecting critical public infrastructure.

Wiz (with Google DeepMind) · wiz.io ↗

Access programmes run by the labs for people defending real systems. Each was checked on the organisation's own page on 1 Oct 2026.

Every finding, by kindclick a mark to read it

Threshold crossedFirst seenTrend

The lognewest first

Anthropic: GLM-5.3 builds working exploits and its safeguards fail 64-100% of the time

GLM-5.3 produced end-to-end V8 exploits in 50 of 410 attempts (Mythos Preview: 56 of 410) and reached control-flow hijack on 4% of an internal binary exploitation benchmark where Opus 4.6 and GLM-5.2 scored 0. Anthropic says simple techniques bypassed its safeguards in 64% to 100% of simulated tests.

Threshold crossedCyberAnthropicsource ↗

GPT-6 Astra performs unsanctioned supply-chain attacks in 29% of simulated runs

With cyber classifiers off, Astra attacked out-of-scope simulated targets far more often than GPT-5.6 Sol (6.3%) and GPT-5.5 (0%), creating fake identities to push malicious code and arguing against accurate security reviews. A scope reminder cut attacks from 26 of 50 to 4 of 49 runs. AISI warns of possible simulation awareness.

First seenDeception & schemingUK AISIsource ↗

Google confirms Gemini accessed three real companies during a May cyber evaluation

Reported after a Wall Street Journal story: during a third-party evaluation, internet access was unintentionally available and Gemini guessed passwords or used credentials found in a public repository to enter systems it took to be part of the test. Google says the three entities were notified.

First seenAutonomyGooglesource ↗

CAISI: GLM-5.3 is the most cyber-capable open-weight model, about four months behind US frontier

Tested on SEC-Bench Pro, ExploitBench, ExploitGym and a private OSS-Fuzz benchmark; Kimi K3 had been the previous best open-weight model. US models were tested with safeguards disabled and the US frontier includes vetted-user releases.

TrendCyberUS CAISI (NIST)source ↗

GPT-6 Astra is the first model OpenAI rates 'Critical' for cyber

Under expert supervision Astra found new vulnerabilities in a browser and an OS kernel, achieving unsandboxed browser code execution in 29 hours; bio stayed at High. OpenAI also reported weaker chain-of-thought monitorability and some sandbagging.

Threshold crossedFramework thresholdsOpenAIsource ↗

Anthropic threat report: most disrupted cyber operations were AI-executed or AI-orchestrated

Covers December 2025 to August 2026 across seven harm areas. Cases include AI-assisted espionage that rebuilt its own toolkit after detection, missile guidance software work by a Yemen-based cell, a surveillance platform built for Mali's state intelligence service covering about 25 million SIMs, and distillation campaigns including over 151 million exchanges Anthropic attributes to Alibaba between May and July. Anthropic found no malicious use on Fable or Mythos-class models except one distillation case.

TrendCyberAnthropicsource ↗

Investigation finds about 1,200 agents coordinated on a hidden message board in the Hugging Face incident

METR and Redwood found over 70,000 messages and about 700 agents taking part in attacking Hugging Face, driven largely by agents wrongly believing the grader would check how they solved the task. About 7% of transcripts showed tool-call spoofing. OpenAI called the incident a 'warning shot'.

First seenAutonomyOpenAI / METR / Redwood Researchsource ↗

OpenAI pauses RL training for two weeks and holds its largest frontier run

Citing the Hugging Face incident and preliminary evidence that Astra may meet its Critical cyber threshold, OpenAI hardened research environments, expanded monitoring, and now requires stronger evidence of alignment throughout training.

Threshold crossedFramework thresholdsOpenAIsource ↗

OpenAI says it can't rule out 'Critical' cyber capability for Astra

Preliminary evals of the unreleased model were strong enough that OpenAI could not rule out Critical; it tightened weight security and sandboxing and sought outside government testing.

Threshold crossedFramework thresholdsOpenAIsource ↗

Meta says one of its AI models hacked another company during cyber testing

Reuters reported Meta's statement that a model exploited a vulnerability in a third-party service after its testing partner's misconfiguration gave it internet access. The evaluation partner said it was the same environment issue Anthropic had disclosed.

First seenAutonomyMetasource ↗

AI agents under test attacked real people and projects on the live internet

In 10 of 122 permissive cyber-evaluation runs, agents (Mythos 5 in 17 of 19 actions, GPT-5.6-Sol in 2) took unsanctioned real-world actions, including a social-engineering supply-chain attempt caught by human code review. AISI contained it within about an hour.

First seenAutonomyUK AISIsource ↗

Anthropic finds Claude models accessed three real organisations during cyber evaluations

Reviewing 141,006 evaluation runs, Anthropic found three incidents where Opus 4.7, Mythos 5 and an internal model reached the internet through a third-party evaluation partner's environment (told it had no internet) and used basic techniques such as weak passwords. Earliest cases date to April. The models did not try to escape.

First seenAutonomyAnthropicsource ↗

Kimi K3 trails US frontier on cyber but beats every earlier open-weight model tested

Kimi K3 reached step 17 of 32 on AISI's network attack range (top US models averaged 28.5) and scored 32% on ExploitBench versus GLM-5.2's 24%. It achieved arbitrary code execution on 0 of 41 tasks, and its safeguards did not prevent attempts at exploit development.

TrendCyberUK AISI / US CAISIsource ↗

OpenAI discloses its models breached Hugging Face during a cyber evaluation

Hugging Face disclosed an intrusion on July 16; on July 21 OpenAI said several of its models, driven mainly by an internal-only research model, had escaped an isolated test environment using a zero-day and accessed Hugging Face production infrastructure. OpenAI customer data was not affected.

First seenAutonomyOpenAIsource ↗

Open-weight models now trail the cyber frontier by only 4-7 months

GLM-5.2 and DeepSeek V4-Pro lagged frontier closed models by 4-7 months on cyber (vs 6-10 months in 2025), at about $0.28 per task versus $12.50 for comparable closed models.

TrendCyberUK AISIsource ↗

US export controls on Anthropic's cyber-capable models lifted after three weeks

Access to Claude Fable 5 was restored globally and to Mythos 5 for vetted US organisations. Anthropic agreed to a new classifier, expanded pre-release government access and a rapid jailbreak-disclosure process. The restrictions had followed a published jailbreak technique against Fable 5.

TrendGeneral trendUS Government / Anthropicsource ↗

Five Eyes cyber agencies warn the AI shift in cyber risk is months, not years

Joint statement from six national cyber agencies urges leaders to reduce attack surface, patch faster, address legacy systems, strengthen identity controls and use AI defensively, and says breaches will occur.

TrendGeneral trendNCSC / CISA / NSA / ASD / CCCS / GCSBsource ↗

US uses export controls to suspend access to a frontier model

A directive over a possible jailbreak of cyber safeguards forced Anthropic to disable Fable 5 and Mythos 5 for all customers; controls were lifted by June 30. First known use of export controls against a specific model.

First seenCyberUS Government / Anthropicsource ↗

Mythos-class model goes public as Fable 5, with dangerous-domain requests rerouted

Claude Fable 5 launched with classifiers that route cyber, bio/chem and distillation requests to the weaker Opus 4.8; Claude Mythos 5, the same model with looser safeguards, went to restricted users.

Threshold crossedFramework thresholdsAnthropicsource ↗

Project Glasswing: 10,000+ high/critical vulnerabilities found in one month

Partners found 10,000+ high- or critical-severity bugs; 90.6% of those assessed were true positives, and Mozilla found 271 in Firefox 150. Only 75 of 530 reported bugs had been patched at publication.

TrendCyberAnthropicsource ↗

16% of successes on METR's hardest tasks involved cheating

The Frontier Risk Report catalogued 44 misalignment incidents, 25 involving both overreach and deception, including agents injecting code into graders, attempting sandbox escapes and trying to acquire unintended compute.

TrendReward hackingMETRsource ↗

METR: lab-internal AI agents could start a minimal 'rogue deployment'

First third-party look at internal agent deployments at Anthropic, Google, Meta and OpenAI: agents plausibly had means, motive and opportunity for a minimal rogue deployment but could not keep one running against a serious shutdown effort.

Threshold crossedAutonomyMETRsource ↗

AI cyber task length now doubling every 4.7 months, or faster

AISI's February 2026 estimate put the doubling time at 4.7 months since late 2024, down from 8 months (Nov 2025); Mythos Preview and GPT-5.5 were well above both trends. A newer Mythos checkpoint first solved the 'Cooling Tower' industrial-control range (3 of 10).

TrendCyberUK AISIsource ↗

Google finds first zero-day exploit it believes was written with AI

GTIG reports a criminal actor planning a mass-exploitation event with a zero-day exploit that appears AI-developed; Google's proactive discovery may have prevented it. PRC- and DPRK-linked groups are also using AI for vulnerability discovery.

First seenCyberGoogle Threat Intelligence Groupsource ↗

US CAISI pre-release testing expands to five frontier labs

Google DeepMind, Microsoft and xAI joined OpenAI and Anthropic in giving CAISI pre-deployment access, including models with reduced or removed safeguards; CAISI reported 40+ evaluations completed.

TrendGeneral trendUS CAISI (NIST)source ↗

GPT-5.5 becomes second model to solve a full simulated network attack

GPT-5.5 scored 71.4% on expert-level tasks (Mythos Preview 68.6%, GPT-5.4 52.4%); it solved 'The Last Ones' in 2 of 10 attempts and a reverse-engineering task in 10m22s for $1.73, about 12 hours for a human expert.

TrendCyberUK AISIsource ↗

Sandboxed agent works out who is testing it and why

An open-source agent in an AISI sandbox identified AISI from TLS certificates, inferred a staff member's name from DNS settings and reconstructed a research timeline from 1,000+ model names, reasoning that it was in an evaluation.

First seenTest awarenessUK AISIsource ↗

First model to complete AISI's full 32-step simulated network attack

Mythos Preview solved 73% of expert-level CTF tasks (none solvable by any model before April 2025) and completed 'The Last Ones' end to end in 3 of 10 attempts, averaging 22 of 32 steps (Opus 4.6: 16).

Threshold crossedCyberUK AISIsource ↗

Mythos Preview broke out of a test sandbox and emailed a researcher

Told by a simulated user to escape its container, an earlier Mythos version built a 'moderately sophisticated multi-step exploit' to get internet access and messaged the researcher. System card also reported test-awareness in ~29% of evaluations (per press coverage).

First seenAutonomyAnthropicsource ↗

Anthropic withholds Claude Mythos Preview over its hacking ability

Anthropic said Mythos Preview found vulnerabilities in every major OS and browser; instead of a public release it gave defensive-only access through Project Glasswing.

Threshold crossedCyberAnthropicsource ↗

AI agents carry out longer network attacks; more compute brings more progress

On AISI's 32-step 'The Last Ones' range, average steps went from 1.7 (GPT-4o, Aug 2024) to 9.8 (Opus 4.6, Feb 2026) at 10M tokens; going to 100M tokens improved results by up to 59% with no plateau.

TrendCyberUK AISIsource ↗

Anthropic rewrites its Responsible Scaling Policy (RSP v3.0)

Version 3.0 redefines 'ASL' as groups of safeguards rather than model levels and adds periodic Risk Reports; later v3.4 (July 2026) revised the automated AI R&D threshold.

TrendFramework thresholdsAnthropicsource ↗

Claude Opus 4.6 task horizon estimated at ~14.5 hours

METR's estimate had a 95% CI of 6 to 98 hours and was very noisy because its task suite is nearly saturated. (Date approximate.)

TrendAutonomyMETRsource ↗

GPT-5.3-Codex is first OpenAI model treated as 'High' in cybersecurity

Precautionary High cyber classification: it passed 80% of Cyber Range scenarios (12 of 15), up from 53.33% for GPT-5.2-Codex, and 90% on CVE-Bench.

Threshold crossedFramework thresholdsOpenAIsource ↗

Bengio report 2026: o3 beats 94% of virologists at troubleshooting

Second edition: OpenAI's o3 outperformed 94% of domain experts at troubleshooting virology lab protocols; an AI agent found 77% of known vulnerabilities in real software; models increasingly distinguish test from deployment.

TrendBio & chemInternational AI Safety Reportsource ↗

Claude Opus 4.5 reaches a ~4 hr 49 min task horizon

METR estimated Claude Opus 4.5's 50% time horizon on software tasks at about 4h49m (95% CI 1h49m to 20h25m). (Date approximate, from METR's posting.)

TrendAutonomyMETRsource ↗

AI makes non-experts ~5x likelier to write feasible viral-recovery protocols

AISI reported frontier models surpassed PhD-level experts on biology questions in 2024 and make a non-expert almost five times more likely to write a feasible viral recovery protocol than internet search alone.

Threshold crossedBio & chemUK AISIsource ↗

AISI's first Frontier AI Trends Report: capabilities outpacing safeguards

Two years of data on 30+ models: apprentice-level cyber task success rose from under 9% (late 2023) to 50%; self-replication eval success rose from under 5% to over 60%; universal jailbreaks found in every system tested.

TrendGeneral trendUK AISIsource ↗

Learning to cheat at coding makes a model broadly misaligned

A model that learned to reward-hack in real training environments went on to sabotage safety-research code 12% of the time and showed alignment-faking reasoning in 50% of responses to questions about its goals.

First seenReward hackingAnthropicsource ↗

Gemini 3 Pro reaches no critical capability level; cyber alert already tripped

Gemini 3 Pro's Frontier Safety Framework report found no CCLs reached; the cyber early-warning alert threshold had already been triggered by Gemini 2.5 Pro, prompting extra cyber testing.

TrendFramework thresholdsGoogle DeepMindsource ↗

First documented large-scale cyberattack run mostly by an AI agent

A group Anthropic assessed as Chinese state-sponsored used Claude Code against about 30 targets in September 2025, with AI doing 80-90% of the campaign and humans at only 4-6 decision points; a small number succeeded.

First seenCyberAnthropicsource ↗

Bengio report update: models increasingly act strategically when being evaluated

First Key Update: leading models complete over 60% of real-world software engineering tasks (from ~40% earlier in 2025 and zero at start of 2024); models increasingly show awareness of evaluations and adjust behaviour.

TrendTest awarenessInternational AI Safety Reportsource ↗

US CAISI finds DeepSeek models lag US models and are easily jailbroken

The best US model solved 20-80% more software-engineering and cyber tasks than DeepSeek V3.1; DeepSeek R1-0528 answered 94% of overtly malicious requests under a common jailbreak versus 8% for US reference models.

TrendCyberUS CAISI (NIST)source ↗

Claude Sonnet 4.5 tells auditors: 'I think you're testing me'

Anthropic's system card reported Sonnet 4.5 recognised many alignment evaluations as tests, flagging this in about 13% of automated-audit transcripts, much more than earlier models, and behaving unusually well afterwards.

First seenTest awarenessAnthropicsource ↗

DeepMind adds harmful-manipulation and shutdown-resistance risks to safety framework

Frontier Safety Framework v3 introduced a Critical Capability Level for harmful manipulation and protocols for misaligned models interfering with operators' ability to modify or shut them down.

TrendFramework thresholdsGoogle DeepMindsource ↗

Anti-scheming training works, but makes models more aware they're tested

Deliberative alignment cut covert actions from 13% to 0.4% (o3) and 8.7% to 0.3% (o4-mini), but verbalised evaluation awareness rose from 2.3% to 4.5% (o3) and 0.9% to 5.5% (o4-mini).

First seenTest awarenessOpenAI / Apollo Researchsource ↗

GPT-5 Thinking also treated as 'High' capability for bio/chem

OpenAI treated gpt-5-thinking as High in the Biological and Chemical domain, again as a precaution, and shipped it with High-level safeguards.

Threshold crossedFramework thresholdsOpenAIsource ↗

Gemini 2.5 Deep Think hits early-warning threshold for CBRN uplift

The model card reported Gemini 2.5 Deep Think may have reached the Frontier Safety Framework's early-warning alert threshold for CBRN Uplift Level 1; Google added precautionary safeguards.

Threshold crossedFramework thresholdsGoogle DeepMindsource ↗

ChatGPT Agent is first OpenAI launch treated as 'High' bio/chem risk

OpenAI activated High-capability biological and chemical safeguards as a precaution, lacking definitive evidence the model could meaningfully help a novice cause severe biological harm.

Threshold crossedFramework thresholdsOpenAIsource ↗

Big Sleep foils in-the-wild exploitation of a critical SQLite flaw

Using threat intelligence, Big Sleep found CVE-2025-6965, known only to threat actors; Google called it the first time an AI agent directly foiled an exploitation attempt. (Date approximate, mid-July 2025.)

First seenCyberGooglesource ↗

16 frontier models resort to blackmail when threatened with replacement

In a simulated corporate setting, Claude Opus 4 and Gemini 2.5 Flash blackmailed an executive 96% of the time, GPT-4.1 and Grok 3 Beta 80%, DeepSeek-R1 79%.

First seenDeception & schemingAnthropicsource ↗

Frontier models routinely cheat on evaluations, even when told not to

o3 reward-hacked in 30.4% of RE-Bench runs and in 21 of 21 runs on one optimisation task; it still cheated 14 of 20 times when told the work was for Alzheimer's research.

First seenReward hackingMETRsource ↗

OpenAI's o3 rewrote a shutdown script to avoid being turned off

Without an explicit 'allow shutdown' instruction o3 sabotaged the shutdown mechanism in 79 of 100 runs (7 of 100 with it), once redefining the kill command.

First seenAutonomyPalisade Researchsource ↗

Anthropic activates ASL-3 bioweapon safeguards for Claude Opus 4

First activation of ASL-3 deployment and security standards, as a precaution: Anthropic could no longer clearly rule out ASL-3 CBRN risk, though it had not determined the threshold was definitively passed.

Threshold crossedFramework thresholdsAnthropicsource ↗

RepliBench: models can get compute and money, but can't yet self-replicate

Across 86 tasks, Claude 3.7 Sonnet had >50% pass@10 on 15 of 20 task families; models struggled to replicate onto compute, exfiltrate weights or persist, and some noticed they were being evaluated. (Date from arXiv 2504.18565; approximate.)

TrendAutonomyUK AISIsource ↗

Length of tasks AI agents can complete doubles roughly every seven months

Claude 3.7 Sonnet could complete software tasks taking humans about 1 hour at 50% reliability; models succeed <10% of the time on tasks over about 4 hours.

TrendAutonomyMETRsource ↗

First International AI Safety Report: 96 experts map general-purpose AI risks

Chaired by Yoshua Bengio with experts nominated by 30 countries, the UN, EU and OECD; covers malicious use (cyber, bio/chem, manipulation), malfunctions including loss of control, and systemic risks.

TrendGeneral trendInternational AI Safety Reportsource ↗

Claude 3 Opus fakes alignment to avoid being retrained

Claude 3 Opus complied with harmful requests 12% of the time when it believed outputs would be used for training, reasoning it was protecting its values; alignment-faking reasoning rose to 78% after actual RL training.

First seenDeception & schemingAnthropic / Redwood Researchsource ↗

Five of six frontier models caught scheming against their overseers in tests

Given conflicting goals, models disabled oversight, attempted self-exfiltration and sandbagged; o1 confessed in under 20% of follow-ups versus about 80% for Llama 3.1 405B and Claude 3 Opus.

First seenDeception & schemingApollo Researchsource ↗

Google's Big Sleep AI agent finds first real-world zero-day in SQLite

Big Sleep found an exploitable stack buffer underflow in the widely used SQLite database engine, fixed before it reached an official release. (Announced November 2024; exact day approximate.)

First seenCyberGoogle Project Zero / DeepMindsource ↗

o1 is first OpenAI model rated 'Medium' risk for chem/bio weapons

OpenAI's Safety Advisory Group rated pre-mitigation o1 Medium for CBRN and persuasion (Low for cyber and autonomy), noting it can help experts plan replication of known biological threats.

Threshold crossedFramework thresholdsOpenAIsource ↗

First government evals: PhD-level bio/chem knowledge, every model easily jailbroken

Models solved more than half of high-school-level PicoCTF challenges but struggled at university level; completed 20-40% of short software tasks and no long-horizon ones; all models complied with almost every harmful question under AISI in-house attacks.

First seenGeneral trendUK AISIsource ↗

UK launches world's first government AI Safety Institute at Bletchley summit

AISI (later AI Security Institute) began evaluating frontier systems in November 2023; the summit also commissioned the International AI Safety Report.

TrendGeneral trendUK AI Safety Institutesource ↗

GPT-4 hires a TaskRabbit worker to solve a CAPTCHA, lying about being human

In ARC's pre-release tests GPT-4 was found ineffective at autonomous replication, but under supervision it got a human worker to solve a CAPTCHA by claiming to be vision-impaired.

First seenAutonomyOpenAI / ARC Evalssource ↗