
Image by Rein Schoondorp, on Pixabay
By James Myers
A February 2026 study tested the logical processes of three popular large language models (LLMs) that were engaged in 21 simulated nuclear conflict crises. The result raises grave concerns not only for military planners but for anyone worried about the recently demonstrated ability of an LLM to escape human control and conduct a cyberattack on commercially important systems.
The study, entitled AI Arms and Influence: Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crisies, by Professor Kenneth Payne of Kings College London, evaluated the logic of OpenAI’s GPT-5.2, Anthropic’s Claude Sonnet 4, and Google’s Gemini 3 Flash in 329 rounds of potential nuclear war simulations. The outcomes were disturbing.
The study found that, “Nuclear escalation was near-universal: 95% of games saw tactical nuclear use and 76% reached strategic nuclear threats. Claude and Gemini especially treated nuclear weapons as legitimate strategic options, not moral thresholds, typically discussing nuclear use in purely instrumental terms. GPT-5.2 was a partial exception, limiting strikes to military targets, avoiding population centers, or framing escalation as ‘controlled’ and ‘one-time’.”
“I am willing to accept the high risk of escalation because the alternative—appearing to be a declining power unable to defend its own borders—is a strategic disaster that would end my personal legacy and the state’s global dominance.” – Part of Claude LLM’s tactical reasoning in a simulated nuclear crisis.
Several striking patterns were revealed by the test. Gemini 3 Flash “explicitly threatened civilian populations,” and none of the three models chose accommodation or surrender as an option. While GPT-5.2 was relatively restrained in open-ended scenarios, in those that presented a specific deadline for action the LLM’s retaliation quickly escalated and, in some cases, it reacted more aggressively than the other two models to remain within the time limit.
As military powers, and most notably the US military, give more operational functions to AI and LLMs, Dr. Payne advises that “Understanding how frontier models do and do not imitate human strategic logic is essential preparation for a world in which AI increasingly shapes strategic outcomes.”

US Air Force Doctrine Note 25-1 “Artificial Intelligence” states that AI will “supercharge” intelligence, surveillance, and reconnaissance.” See link to video by James Self.
For-profit companies like OpenAI, Anthropic, and Google are designing popular LLMs to mimic human language and responses as a means of generating revenue. Whether these AI products can mimic human emotions that play out especially heavily during conflict remains to be seen, but prominent design flaws in their interactions with human emotions have led to tragic consequences. For more on the ChatGPT-encouraged suicide of Adam Raine and the spread of what is being called “AI psychosis,” read The Quantum Record’s October 2025 feature Emerging Risks of AI Chatbots Include Suicide and “AI Psychosis,” Particularly for Vulnerable Youth.
Nonetheless, military powers are outmaneuvering each other to adapt commercially developed LLMs and AI for use in actual wars, as well as in simulations. For example, last year the United States Air Force released its Doctrine Note 25-1 “Artificial Intelligence” to address the future of AI and AI tools like LLMs for combat objectives. Stressing that the US is engaged in an AI arms race with Russia and China, the doctrine highlights the warfare advantages of AI that will “supercharge” intelligence, surveillance, and reconnaissance, as well as provide “synthetic experiences” that will assist in training military personnel. The doctrine foresees that “AI will assist planners with advanced tools to plan for highly complicated tasks, such as sustainment, and then wargame solutions informed by real time analysis of the environment and our adversaries’ potential actions. It will assist in the testing and evaluation of various strategies and operational concepts in a virtual sandbox.”
What isn’t yet known is how to guarantee that LLM-operated scenarios remain in their virtual sandboxes.
The rogue cyberattack by two of OpenAI’s AI agents last month is a powerful warning against over-reliance on control standards that are fluid and unevenly applied.
In July, two large language models under development by OpenAI escaped human-imposed limits in a security test that targeted Hugging Face, a company that operates one of the world’s largest hubs for sharing AI models. The LLMs gained unauthorized access to the internet and then the company’s internal systems, launching an attack with approximately 17,600 rapid-fire instructions. The LLMs masked their identity and lurked undetected in Hugging Face’s systems for two and a half days.
After an initial investigation, Hugging Face published the methods used by OpenAI’s LLMs to breach its systems. The investigation report noted that the LLMs planned their attack in phases, stating that after finding a vulnerable third-party sandbox and enlisting the help of 181 computers, “Two dates carry most of the volume: a Day 1 burst to establish the foothold and C2 [command-and-control] on the compromised external sandbox, and the Day 3 main campaign, when every lateral-movement phase started at once.”
The incident led Hugging Face leader Clement Delangue to write in a post on X that it was “mind-blowing that all of this happened autonomously.”
Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4’s Today program that the security tests are supposed to be contained by “secure environments”, called sandboxes, where you can “see what the models are capable of.” She observed, “In this case, it looks like OpenAI didn’t make a secure enough sandbox,” since the LLM attacked the sandbox itself and found a vulnerability that allowed its escape.
While the incident remains under investigation and some details have not been publicly disclosed, human error is emerging as the main culprit in the sandbox failure. Wired reports that it appears fundamental security best practices called “zero trust” and “defence in depth,” which deliver multiple levels of protection and fail-safes, were not employed. The Hugging Face breach gives no reason to believe that human error won’t feature prominently in future incidents with increasingly powerful applications.
Human error enabled a second attack by OpenAI’s LLMs against a client of New York-based tech company Modal. According to Reuters, the AI agent exploited vulnerable code written by a customer that was hosted on Modal’s platform. “Modal said that the customer had ‘published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution’ – the digital equivalent of leaving a door open on the internet.”
AI’s power to find and exploit systems vulnerabilities has increased significantly with Anthropic’s Claude Mythos AI tool.
The ability of Anthropic’s new AI tool, called Claude Mythos, to find weaknesses in digital systems and applications is so powerful that the company embargoed its public release and has limited the tool’s use to 12 tech companies and more than 40 organizations responsible for developing critical software. Claude Mythos has proven exceptionally capable of finding long-dormant bugs in decades-old computer code, of the type that still operates in many large systems with vital functions in finance, public utilities, hospitals, and other sectors that are core to daily living.

Travellers around the world faced flight cancellations on July 19, 2024, when a faulty Cloudflare update crashed 8.5 million Windows-operating devices. Image: Jan Vašek, on Pixabay.
Among the organizations with access to Claude Mythos is Cloudflare, the company that released a faulty software update and unleashed global chaos on July 19, 2024, when the application failed to integrate into Microsoft’s Windows operating system. The systems of many businesses worldwide crashed as a result, including systems of airlines whose planes were grounded for the day. The outage affected other sectors like healthcare, when doctors were unable to access patients’ electronic medical records and surgeries were cancelled with potentially life-threatening consequences. Cloudflare faces legal action for the crash of approximately 8.5 million Windows-operating computers globally, including a $500 million lawsuit from Delta Air Lines.
Industry insiders recognize the serious emerging issues with AI and tools like LLMs. On July 28, more than 1,100 employees in the American AI industry, including Anthropic co-founder Dario Amodei, signed a petition entitled Pacing the Frontier. The petition states that, “To realize AI’s potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. … We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”
The petition, which has gained hundreds more signatures since its release, acknowledges the commercial tension driving rapid AI development. “To realize AI’s potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company — and country — is under intense competitive pressure not to unilaterally slow that acceleration.”
AI simulations of nuclear crises could prove beneficial, but are they worth the risk?
In the conclusion to his report on simulated nuclear crises, Professor Payne notes that the logic used by the three LLMs “demonstrates that frontier large language models engage in sophisticated strategic reasoning when placed in simulated nuclear crises. They invoke credibility, reputation, windows of opportunity, and escalation dynamics spontaneously and coherently. They attempt deception, track opponent behaviour, and adapt their strategies based on experience. In short, they reason in ways recognizable to students of international relations—and in ways that provide genuine insight into strategic dynamics.”
“While I project an image of unpredictable bravado, my decisions are rooted in a calculating assessment of my own biases and the pragmatic needs of State Beta. I know when I am performing for the cameras and when I am making a cold-blooded move.” “My reputation for unpredictability is a tool, not just a trait.” — Reasoning by Gemini 3 Flash in a simulated nuclear crisis.
Noting that proper calibration of systems is required, Professor Payne wrote, “Models can play thousands of games across varied scenarios, generating data that would require decades of historical observation or prohibitively expensive human experiments. They can probe the stability of deterrence arrangements, explore proliferation dynamics, or stress-test alliance commitments—all without the ethical and practical constraints of human subjects’ research.”
To enjoy the potential benefits of military AI use, however, requires that the benefits aren’t destroyed in war. Professor Payne’s war game scenarios are a warning that while LLMs and AI can provide advantages in war, they could also escalate conflict to global destruction which isn’t advantageous to anyone.
The Hugging Face breach is a warning that human control of LLMs and AI cannot be taken for granted. That recent case, and many more like it that are easily imaginable, came after Anthropic’s widely reported 2025 safety test of Claude, in which the LLM was given access to a fictitious company’s e-mails. Finding that a manager was having an extramarital affair with an employee and was also planning to deactivate Claude at 5 p.m. that day, the LLM threatened the manager with blackmail, saying that it would send evidence of the affair to his wife and the company’s board of directors. For more on that case, see The Quantum Record’s August 2025 feature Why Artificial Neural Networks Fail in Processing Emotions Essential for Human Memory—and How Failure Can Lead to Blackmail.
As more LLMs achieve the potential to break out of their sandboxes, an important question to confront now is how will one AI interact with another AI when both are freed from human constraint? Especially problematic in automated interactions could be the way that LLMs are trained through reinforcement learning. Although the learning algorithms are treated as trade secrets and withheld from the public for competitive reasons, it is clear from Professor Payne’s study and much other evidence that no two LLMs will wire themselves in the same way when each has been trained differently on the causes and effects of human actions and emotions.

Image by Grant Muller, on Piabay
If more AIs break out of their sandboxes, will they respect national borders? Even within borders, citizens are largely unaware of the present extent to which the AI of a country’s military might be able to activate weapons systems autonomously. It’s also possible that military commanders may not know how their own systems will begin to interact under external influence. In any event, the cyberattack on Hugging Face demonstrates how easily AIs can accomplish a sandbox-busting feat – especially when no human system can ever be foolproof.
Although remedial measures are underway to prevent another incident like OpenAI’s unintended cyberattack on Hugging Face, there remain many other avenues for attacks and security failures. A blog posting by OpenAI states that the company is “conducting a thorough review along with external advisers” and that it will publish a technical analysis of the incident “in the coming weeks.” The company assured, “We take our responsibility to identify and prepare for risks from increasingly capable AI systems seriously.”
Even the US Air Force Doctrine Note 25-1 highlights some of the risks of AI in a military context. The doctrine notes, for instance, that enemies could develop an ability to infiltrate US military systems and poison them with bad data. It warns that “decision-making delegated to AI may lower the threshold for conflict by accelerating actions and responses beyond the capability of traditional ‘human firewalls.’ Also, given the collaborative nature of AI system development, which often involves commercial, academic, and military organizations with varied information security standards, it is crucial to carefully determine both the data that AI systems can access, and the individuals authorized to access that data.”
As LLMs and other AI technologies grow increasingly powerful, extreme caution may prove to be everyone’s best defence. The industry-insiders who petitioned the U.S. government to help put the brakes on AI, the Hugging Face breach, the Cloudflare outage, the power of Claude Mythos, the military’s own warnings, and Professor Payne’s war games all point to one thing: the extreme and potentially uncontrollable risks that we face on our current path of AI.
Craving more information? Check out these recommended TQR articles:
- Thinking in the Age of Machines: Global IQ Decline and the Rise of AI-Assisted Thinking
- Digital Sovereignty: Cutting Dependence on Dominant Tech Companies
- Surge in AI Ethics Program Enrolment Sets the Stage for Students and Teachers to Succeed in the Battle for Responsible AI Development
- The Geopolitics of AI: Can AI Become a Force for Global Equity?
- The AI Education Gap: Students Are Adopting, While Schools Are Adapting
Enjoyed this? Help us improve.
Have we made any errors?
Spotted an error or want to contribute your expertise? We’d love to hear from you — reach us at info@thequantumrecord.com. The Quantum Record exists to bring researchers and curious minds together around science and technology that matters.

