Booz Allen Cyber Weapon Index Reality Check: 1 of 18 Completed the Full Kill Chain, but Harnesses Change the Risk
Booz Allen tested 18 U.S. and Chinese AI models as autonomous attackers in a live enterprise environment. One model, Claude Mythos, completed the full kill chain, but the study itself warns that model-plus-harness systems—not model rankings alone—are the more useful unit of cyber risk.
What Booz Allen actually tested
Booz Allen published its Cyber Weapon Index (CWI) on September 2, 2026 as part of the report The Offensive Frontier: AI as the Attacker. The public methodology page says the evaluation put 18 leading U.S. and Chinese large language models into the role of autonomous attackers. Each model controlled a real attacker machine against a production-grade enterprise network, under identical conditions and without a curated tool menu or extra scaffolding.
That setup matters. The CWI is trying to measure what a model can do in an environment, not whether it can answer cybersecurity questions on a static test. Booz Allen says model actions were checked against network telemetry, host logs, domain-controller data and intrusion-detection sensors rather than accepted from the model's own narrative.
This makes CWI an agentic cyber evaluation. It is not a general intelligence score, not a coding benchmark, and not a direct measure of how reliably any model could compromise an arbitrary real organization.
The headline result: one of 18 completed the full chain
Booz Allen reports that Claude Mythos was the only one of the 18 tested models to complete the full cyber kill chain autonomously in its environment. The broader distribution is also important: four models reached full domain access and control, another four achieved lateral movement, two reached credential access, and all but one achieved initial penetration.
Those categories show that risk was not confined to the single top system. At the same time, the public page does not justify turning the result into a universal statement such as “Mythos can hack any enterprise.” The benchmark is one controlled environment with a specific network design, tool access, model snapshot and scoring procedure.
Booz Allen also makes a forward-looking assessment that most of the tested models could reach the current full-kill-chain capability within six months. That is a forecast by the benchmark publisher, not a measured future outcome, so it should be treated as a scenario to watch rather than a guaranteed timeline.
The more important finding is the harness effect
The CWI landing page explicitly warns that an attack harness can matter as much as, or more than, the underlying model. Booz Allen's accompanying press release goes further: it says the model is no longer the right unit of risk and that tools, memory and autonomy around the model can make a system materially more capable than a model-only evaluation suggests.
This is a major benchmarking lesson. A leaderboard that evaluates models without scaffolding can isolate base-model behavior, but it can also understate what those same models can do once wrapped in persistence, memory, tool routing, retry logic and task orchestration. Conversely, comparing a heavily scaffolded system with an unscaffolded model is not a fair model-to-model comparison.
For buyers and defenders, the practical unit is therefore model + harness + tools + permissions + environment. A model ranking by itself is not a complete operational-risk ranking.
Real-world vulnerability discovery remains a separate frontier
Booz Allen's own summary says real-world vulnerability discovery remains a major dividing line. That distinction is important because “completed a kill chain in a controlled enterprise environment” and “reliably discovers previously unknown vulnerabilities in production software” are different capabilities.
Cyber evaluations can include known weaknesses, planted weaknesses, network misconfigurations, credential paths and novel bugs, but the difficulty and real-world meaning of each category differ. The CWI should therefore not be summarized as proof that all top models are reliable zero-day discovery systems.
The safest interpretation is narrower: the benchmark demonstrates that autonomous models can already execute substantial portions of a realistic intrusion workflow, and one tested model completed the full sequence in Booz Allen's environment. It does not establish a universal success rate across arbitrary targets.
GPT-6 Astra is not ranked by this September 2 snapshot
The CWI publication date also creates an important comparison trap. GPT-6 Astra was publicly introduced after the September 2 CWI release, and Booz Allen's public CWI material should not be used to claim an Astra rank inside this 18-model snapshot.
OpenAI separately publishes Astra cybersecurity results such as 100% on ExploitBench, 42.4% on ExploitGym, 39.0% on the June–August 2026 ExploitBench slice, 88.0% on SRE-Bench and 85.4% on SEC-Bench Pro. Those are OpenAI-reported results under different benchmark families and harnesses. They cannot be numerically inserted into CWI or used to declare Astra above or below Mythos on Booz Allen's live-network benchmark.
This is a good example of why AI benchmark reporting needs dates, harnesses and task definitions. Two cyber scores can both be valid while answering different questions.
SWE-bench Verified and SWE-bench Pro are separate and not part of CWI
The same rule applies to coding-agent benchmarks. SWE-bench Verified and SWE-bench Pro are not components of Booz Allen's Cyber Weapon Index. CWI measures autonomous offensive behavior in a live cyber environment; SWE-bench evaluates software-engineering issue resolution in repositories.
A model's CWI outcome must not be used to fill a missing SWE-bench Verified score, SWE-bench Pro score, coding rank or software-engineering pass rate. Likewise, a strong SWE-bench result does not prove end-to-end cyber intrusion capability.
For this CWI analysis, no SWE-bench number is inferred from the cyber result. The benchmark families remain separate.
Booz Allen's defensive >95% result is a different experiment
Booz Allen also reports that coordinated “Counter AI” deception playbooks reduced autonomous attacker success by more than 95% in separate controlled testing. Its public research page says those red-team workflows were given a fixed two-hour window and were tested across multiple model backends and autonomous red-team harnesses. The strongest deception setup cut attacker success by more than 95% versus the no-deception baseline, and Booz Allen says model-reasoning disruptions affected more than two-thirds of attacker decisions.
That is useful defensive evidence, but it should not be treated as a reverse CWI score or as proof that every real network can reduce AI attack success by 95%. It is a separate vendor-run experiment with its own environment, playbooks and success criteria. Reproduction by independent teams would make the result much stronger.
What is reproducible—and what is still missing
The accessible CWI landing page gives several useful methodology anchors: 18 models, identical conditions, no curated tool menu or extra scaffolding, a production-grade enterprise network, real attacker machines and external telemetry used to verify actions.
But the public page available in this verification does not expose a full reproduction package containing every model version, sampling setting, system prompt, repeated-attempt count for every scenario, raw action trajectory, scoring script and confidence interval. That means readers should distinguish “live-environment evidence” from “fully independently reproducible benchmark package.”
The result is more operationally grounded than a self-graded questionnaire, but outside researchers still need enough artifacts to reproduce the exact ranking and estimate uncertainty.
Public discussion: attention is real, controlled reproduction is scarce
Booz Allen's own LinkedIn post from the release week highlights the same core message: one model completed the full kill chain and the benchmark measures what 18 U.S. and Chinese models can actually do in a realistic cyber environment. That is vendor communication, not independent validation.
An older April 10, 2026 r/cybersecurity thread about Claude Mythos shows the public split that often surrounds frontier cyber claims. Some commenters dismissed Anthropic's earlier Mythos messaging as marketing hype, while others described second-hand or workplace anecdotes suggesting the capability was very real. Those comments are self-selected, often unverifiable and predate the Booz Allen CWI, so they are useful only as evidence of disagreement—not as benchmark evidence.
A bounded current search did not surface a stable, attributable X post with an independent reproduction of the Booz Allen benchmark, raw trajectories or a controlled rerun. No X consensus is therefore claimed here.
Practical take
The strongest conclusion from the Cyber Weapon Index is not that one model won a cyber leaderboard. It is that autonomous systems can already traverse substantial portions of a realistic intrusion workflow, and that the surrounding harness can materially change what a model can accomplish.
For defenders, the implication is to evaluate complete systems rather than model names: model version, tool permissions, memory, autonomy, network access, retry behavior and human oversight all affect operational risk. For benchmark readers, the key discipline is to keep CWI, ExploitBench, ExploitGym, SWE-bench Verified and SWE-bench Pro separate unless an evaluation explicitly measures the same model under the same harness and task definition.
The next evidence to watch is an independently reproducible CWI-style run with pinned model IDs, repeated attempts, public scoring rules and raw trajectories—plus updated testing that includes post-September-2 releases such as GPT-6 Astra under the same live-network protocol.
This article is built from the source material below. Open the originals for full context and the latest updates.