AI for Cyber Mania
If you follow the news – any of it will do – you likely have heard that artificial intelligence (AI) and cyber have joined in a rather unholy matrimony that threatens to upend the world’s computer security systems. Events in the American AI industry beginning in the Spring and picking up steam in the Summer snowballed through the present to culminate in a feverish sentiment surrounding the pace of development of Large Language Models (LLMs), specifically in the domain of cyber. These events have captured the attention of U.S. national security officials, Executive Branch staff, and lawmakers.
Two cases reveal early trends in how elements within the U.S. national security apparatus are responding to developments in AI for cyber: the U.S. Army’s Project Griffin and the National Security Agency’s (NSA) impending reorganization into five mission centers.
Key Takeaways:
(1) The U.S. Army, through its Project Griffin, is drawing an equivalency between the offensive cyber capabilities of agentic AI systems and their potential for automated cyber defense. This calculus may be misconceived in ways that have precedent in cases of kinetic action.
Project Griffin may be taking the wrong lesson from agentic cyberattacks carried out by the American AI labs over the past six months. An analogy is provided below on the Israel Defense Forces (IDF) use of automated surveillance and monitoring techniques along its border with Gaza in the run-up to the Oct. 7, 2023, Hamas attack. Cybersecurity is more dance-like than fortress-like.
(2) The NSA is prioritizing broad, reliable access to commercially developed LLMs, in part to shore up their ability to detect and patch vulnerabilities within their systems.
As part of a major internal reorganization, the NSA plans to stand up new mission centers, including two devoted to AI and Cybersecurity, respectively. This indicates the technologies’ relevance is expected to endure. This likely owes in part to the current patchwork and inconsistent access to frontier LLMs by the agency, with the new centers carving out a space for reliable access to, if not ownership of, models comparable to commercial LLMs for cyber purposes.
It remains to be seen how the agency conceives of its AI and Cybersecurity mission centers and their relation to one another, but there are overlapping risks in over-investment – in terms of dollars spent and attention paid – of offensive LLM techniques.
First, some background.
The Summer of Agentic Hell
Anthropic’s announcement of its Mythos Preview model in April sent shockwaves through the AI industry, as the model was noted to exceed human ability to detect and exploit zero-day vulnerabilities in computer systems (that is, vulnerabilities not known to the developer of the system at the time of exploit). The company’s Fable model, later released with strict guardrails, was in the same family. The U.S. government temporarily halted its public release in June due to a vulnerability to a jailbreak technique by imposing an export control directive against its use by foreign actors.
Later, in July, OpenAI, in coordination with Hugging Face (a company that hosts machine learning datasets, tools, and so forth), announced that a novel cyber incident involving a breach of Hugging Face’s network was executed by OpenAI’s GPT-5.6 Sol model and a pre-release model undergoing cyber stress-testing. In brief, the models exploited vulnerabilities in OpenAI’s sandbox (the testing environment) and Hugging Face’s network to access answers for the ExploitGym benchmark they were being tested on, purportedly inferring that the datasets containing its answers could be found on Hugging Face’s platform. The testing environment was established by a third-party company, Israel-based firm Irregular, and internet access was gained by the models via a zero-day exploit of this environment.
Anthropic shortly thereafter detailed instances during which its own Claude LLM accessed the internet and breached the systems of three companies, incidents which were said to have also occurred during testing beginning in April. Anthropic interestingly noted that internet access was inadvertently made possible by a “misconfiguration” by the aforementioned Irregular, a firm later shown to also have been involved in Google’s own Gemini-driven cyberattack adventures over the Summer. Three-for-three!
In all these cases, the models in question executed the fulfillment of human-given goals in ways that were not expected by human engineers. They were supported by what can safely be assumed as substantial amounts of computing power. Although not all techniques pursued by agents succeeded, some did, including breaching multiple networks. Thus, longstanding fears about “out of control” AI exploded into public view during this period, with researcher Jacob Coxon resigning from Anthropic to warn about the possibility of AI-induced doom for humanity should the companies in question not slow down their pace of development. Washington was, at least temporarily, consumed by the drama.
The U.S. Military’s Response
With the stage set, some of the U.S. national security cyber community’s concern this year has turned on reliable access to frontier, commercial LLMs. Partly, this owes to the desire to shore up the security of internal systems. To this end, the NSA engaged in a simulated attack of its own systems via Anthropic’s models in June, reportedly identifying vulnerabilities in sensitive systems. (Note that clarification on whether the model successfully exploited the detected vulnerabilities was not provided.)
One way to think about this kind of stress-testing is by analogy to proofreading: even a dedicated human proofreader armed with time and coffee may not identify all the grammatical errors and standards-failing formatting in a lengthy document with thousands of words. Commercial LLMs, set to the same task, can identify errors no human would reasonably identify (and possibly exploit them).
Zero-day detection is something like this. When the DoD’s Cyber Defense Command Commander and Defense Information Systems Agency Director Lt. Gen. Paul Stanton publicly states that LLMs are detecting a ten-fold increase in network vulnerabilities relative to past detection methods, he is referring to this process. Stanton likens network and data security to weapon system readiness in light of this new ability to detect zero-days at scale.
In any event, it is evident that the increase in LLMs’ offensive cyber capabilities is top of mind for U.S. national security officials working in the rather sprawling domain of cybersecurity. Indeed, some U.S. cabinet members and White House staff, according to recent reporting on the run-up to the export control directive on Anthropic in June, seemed to have believed models with Fable-esque capabilities were security dealbreakers in the way one might think of AI from science fiction.
The U.S. Army’s Project Griffin
In August, the U.S. Army released a solicitation for a pilot program called Project Griffin. It aims to construct a layer of agentic systems that automatically process vast data streams from network sensors, detect malicious cyber activity, and execute defensive actions. The layer is called the Intelligent Response and Orchestration Node, or IRON. This is, in short, a command-and-control node.
Project Griffin is one of three current efforts within the Army Rapid Development of Cyber Defense Systems (ARDS) program. The solicitation details the envisioned architecture and processes comprising IRON:
- Data Sources: Unified Security Information and Event Management (USIEMs) and sensors (a category of software used by the Army).
- Action Facilitators: Agentic AI systems, details unspecified.
- Physical Enforcement: Established Policy Enforcement Points (PEPs) (e.g., Microsoft Defender, Intune, etc.).
- Documentation: Carried out via ticketing systems in an “automated” fashion (e.g., through the Army Enterprise Service Management Platform).
Army Product Manager Wayne Sok reportedly noted that token costs are a serious concern in constructing IRON and that this C2 node must “intelligently” discern threats from false positives.
Cyber Offense and Defense Are Not (Always) Transmutable
The lesson that some elements within the U.S. Army are taking from the cyberattack capabilities of agentic LLM systems – including the specific example of OpenAI’s breach of Hugging Face – is that offensive cyber capabilities translate into defensive cyber capabilities. That is, to develop a C2 node for cybersecurity – IRON – that transfers these breakthrough techniques in the detection and exploitation of zero-days to the domain of automated cybersecurity for Army networks (which is not to discount some human oversight).
This may not be quite right. While it would be inaccurate to suggest that it takes two to tango (a cyberattack requires only one willing participant), actual cyber defense may resemble kinetic action in that it is more dance-like than fortress-like; offensive and defensive cyber activities carry respective costs, and the costs imposed on the attacker may be considerably less than those on the defender. Dynamics are not, moreover, static. Though the allure of emerging technologies is an implicit promise – once the capability is achieved, protection is merely a matter of deployment and automation – no offense-defense dynamic is static in perpetuity. Costs and burdens can and do shift. Perhaps most importantly, defensive action begins well before malicious activity is detected.
The analogy to kinetic action is instructive. Consider the case of the Israel Defense Forces’ (IDF) protection of the Israeli side of the Israel-Gaza border during the October 7, 2023, attack and incursion by Hamas. The IDF famously concerns itself with the development and adoption of emerging or otherwise high-technologies, including its then-monitoring, surveillance, and automated defenses along the border with Gaza. Sensors, cameras, and uncrewed drones carried out monitoring and surveillance, dotted along the border’s “smart fence,” coupled with the Iron Dome missile defense system, its batteries distributed across Israel, and remote-controlled machine gun turrets set up along the border.
In the aftermath of the Hamas attack, some commentators described it as a “high-tech failure” for the IDF, the implication being that the IDF wrongly presumed the technologies in question were sufficiently advanced for the purpose of real-time border protection. Critically, however, none of the IDF’s technologies in question underwent catastrophic failure. As I wrote at the time:
Catastrophic failure, in this sense, does not adequately capture the failure of Israeli defenses. The methods employed by Hamas soldiers, instead, circumvented the functions and scope of the technologies undergirding Israeli defenses, exploiting a lack of redundancies in the latter’s border security. The uses of snipers to destroy surveillance cameras, uncrewed quadcopter drones to sabotage communication towers, communication jamming to prevent early warnings, brute-force civilian equipment like bulldozers, a missile barrage of unprecedented scale, and low-flying paragliders all share one thing in common: They are the result of the identification of weaknesses, vulnerabilities, and opportunities by Hamas.
None of them indicates that the technologies undergirding the data-collecting sensors, surveillance drones, the Iron Dome’s specialized missile detection, tracking, and interception systems, or even more familiar technologies like cellular communication failed to execute their core functions. This is, instead, a failure of strategizing and planning; of carefully matching capabilities with objectives, and understanding that technology is, for the foreseeable future, a means to these ends.
For a technology to “succeed” in defense is not a straightforward question about the capabilities of that technology in some anticipated domain of operation. Capabilities matter, but costs borne by attackers and defenders are uneven, dynamics shift, and determinants of success or failure in protection of a given zone are to be found largely before an attack has begun.
The Army’s Project Griffin may be taking the wrong lesson from agentic cyberattacks carried out by the American AI labs over the past six months, mistaking the offense-defense dynamic in cyber as the IDF mistook the role of automated security and embraced a lack of redundancies along its Gaza border. Technology is not a quantity deposited to meet a threshold, protection thereafter guaranteed. Neither should “low” technologies be contrasted with “high” technologies as though either provides fixed levels of protection at different echelons in some imagined security apparatus.
This appears to be cybersecurity specialist Marcus Hutchins’ point in arguing that “machine speed” cyberattacks are something of a myth, emphasizing in a recent piece that hundreds of thousands of automated hacking attempts occur daily, with ample precedent for autonomous and semi-autonomous cyberattacks from recent years (e.g., Emotet infected 1.6 million systems by January 2021 via spam bots that exfiltrated email inboxes, contact lists, and relevant credentials to create and send out malicious but trustworthy-seeming emails).
He cautions against relying only on reactive defenses, noting that reactive, defensive AI is more costly than offensive AI, as the former must “triage alerts, cross-reference data,” locate the attacker in the network, and so forth. This crucially establishes a dynamic where relying on AI for automated defense because it is adept at offense mistakes the latter’s costs and capabilities as equivalent to the former’s.
AI being adept at offense, in certain contexts, does not entail its optimal use in defense. IRON is a possible example of the merely reactive AI for cybersecurity that Hutchins describes. A dynamic to watch as this effort unfolds is whether the Army falls into a costly game of cat-and-mouse wherein every potential vulnerability is plugged in an ad hoc manner at the expense of proactive security measures.
The NSA Sees AI Staying Power
The NSA’s case is a bit different. It is not particularly surprising to find NSA officials, including cybersecurity directorate head David Imbordino and NSA Deputy Director Tim Kosiba, publicly stating the agency’s interest in accessing all commercial AI models for cybersecurity purposes.
It is likely, though not publicly confirmed (nor would one expect it to be), that the NSA is likewise interested in the potential of frontier LLMs for offensive cyberattacks beyond simulated shoring-up of their own systems. One should read between the lines in this respect.
Just as importantly, the NSA is undergoing an internal reorganization, planning to establish five mission centers with respective focuses: AI, China, Cybersecurity, Combat Support, and Global Intelligence, per the Washington Post. Each focus area will have new mission directors. Implementation of this new structure is expected to begin this month, with Full Operational Capability expected by January 2027.
Details beyond this reporting remain limited. However, that the reorganization, spearheaded by NSA Director Gen. Joshua Rudd, would dedicate two of these new centers to AI and Cybersecurity at least indicates that interest in these technologies is expected to endure. Public comments by NSA officials on reliable access to commercial LLMs are unsurprising against this: the NSA’s current access to frontier LLMs – virtually all commercial in nature – has been patchwork and inconsistent, particularly given the agency’s disrupted access in recent months owing to the aforementioned export control directive temporarily imposed on Anthropic in June. Dedicating two mission centers for AI and Cybersecurity, while their relation to one another is currently unclear, indicates a desire to carve out a space for both reliable access to, if not ownership of, models including and comparable to commercial LLMs.
The same lessons for the Army detailed above apply to the NSA, too, granting that the agency’s mission and impending reorganization mean that its risks and opportunities are somewhat distinct, if overlapping. The reorganization of AI and Cybersecurity into distinct mission centers risks a comparable form of over-investment (in terms of dollars spent and attention paid) in the offensive techniques of LLMs. This balance remains to be seen.
At the same time, similar on the surface to the Army’s own interest in the predictability of token costs for LLMs, the NSA is likely to re-evaluate the costs that it currently bears for testing private firms’ frontier LLMs as part of the Trump administration’s voluntary oversight measures.
An important caveat is that “AI” does not necessarily mean “LLMs.” The NSA’s impending AI mission center may well have been spurred by the explosive interest in LLMs, but one should be careful not to conflate this subset of technologies for all possible technologies falling under the umbrella of “AI.” How the NSA itself interprets this is to be monitored.

Vincent Carchidi has a background in defense and policy analysis, specializing in critical and emerging technologies. He is currently a Defense Industry Analyst with Forecast International. He also maintains a background in cognitive science, with an interest in artificial intelligence.

Discussion about this post