One of the largest cybersecurity contractors in the United States has put a number on what security researchers have been warning about all year: the frontier AI models built for chatbots and coding assistants are now capable of running cyberattacks largely by themselves. For ongoing coverage of AI security developments, see AI Buzz Wire.
A new report from Booz Allen Hamilton, covered by SC Media and The Register and making the rounds across the industrial security world over the past week, evaluated 18 AI models — nine from the United States and nine from China — and found that one of them autonomously completed a full cyber kill chain in testing: identifying vulnerabilities, gaining access to a target network and escalating to administrator-level control without human intervention.
One model finished the kill chain
The report, known as the Cyber Weapon Index, scored Anthropic's Claude Mythos as the only model to complete the entire attack sequence autonomously. Several others came uncomfortably close. According to SC Media's coverage of the findings, xAI's Grok-4.5 and OpenAI's GPT-5.6 Sol demonstrated significant autonomous attack progress, reaching full domain access or lateral movement within compromised networks — the intermediate steps an operator performs between initial entry and full control.
A kill chain, in cybersecurity terms, spans reconnaissance through exploitation, escalation and persistence. A model completing it unassisted has effectively functioned as an offensive cyber operator, not a tool that merely writes exploit code when asked.
The harness matters more than the model
Perhaps the report's most consequential finding concerns software, not models. The attack harness — the scaffolding that connects an AI model to real hacking tools, scanners and exploit frameworks — dramatically amplifies whatever offensive capability a model already has. A mid-tier model wired into a capable harness can outperform a stronger model operating raw, which means the barrier to autonomous attack capability is dropping for reasons that have nothing to do with the next model release.
Booz Allen's analysts warn that mainstream AI-driven attacks from financially motivated criminal groups and state-sponsored actors are now imminent. Their projection: most of the models tested are expected to reach weaponization levels comparable to today's leaders within six months. The firm also urges the US government to set deadlines for critical infrastructure operators to demonstrate resilience against AI-enabled attacks, and to develop superior offensive and defensive AI capabilities of its own.
A warning that matches the industry's own
The assessment lands amid a string of similar alerts from inside the AI industry itself. Anthropic has previously flagged AI-driven cyberattacks as a critical inflection point, and OpenAI's own rollout of its frontier model was accompanied by unusual public disclosures about advanced cyber capabilities — including a decision to restrict access while the company assessed what its model could do to real networks. Separate reporting on Booz Allen's findings, carried by Industrial Cyber, framed the same report as a warning that response windows for critical infrastructure are narrowing as AI autonomy increases.
The policy debate this feeds is no longer hypothetical. Lawmakers have already opened inquiries into the cybersecurity risks of foreign-origin AI models running inside US critical infrastructure, and grid operators have reported a surge in AI-driven electricity demand coinciding with growing state-sponsored interest in energy systems. A documented test in which a commercial AI model completes a kill chain on its own converts those abstractions into a benchmark.
What defenders should take from it
For security teams, the report's practical implications are narrower than its headlines. Autonomous offense presumes autonomous defense: the same agentic capabilities that let a model chain exploits are already being deployed on the blue side, triaging alerts, correlating telemetry and patching at machine speed. The gap Booz Allen identifies is organizational — most critical infrastructure operators still measure incident response in hours, while AI-assisted attacks compress the equivalent timeline to minutes.
The firm's recommendation of hard deadlines is a bet that regulation, not market pressure, will close that gap. Whether the US moves on it — and whether AI vendors treat offensive capability evaluations as a pre-release gate the way pharmacology treats clinical trials — will shape how the next six months of the report's projection actually play out.
The offensive market is already forming
If any confirmation were needed that this capability is becoming a product category, it arrived a day after the report circulated. As SC Media separately reported on September 4, a startup called Abliteration.ai launched a service explicitly aimed at enabling offensive cyber operations, red-teaming and agent testing with AI models that other providers refuse. The timing underscores the report's core point about harnesses: the bottleneck for autonomous attacks is no longer model capability but the surrounding tooling, and tooling is something the market is now supplying on its own.
Defensive teams do not need to wait for the next benchmark cycle to act on the findings. The report's six-month weaponization window is, functionally, a deadline for asset inventory, segmentation and response-time measurement — the unglamorous controls that determine whether an autonomous attacker meets a hardened network or an open one.
Stay Ahead of AI
Track the AI security stories that matter. Read more AI news before it hits your network.
Read more AI news →