Self-Sovereign AI and Catalytic Nuclear War: A Plausible Near-Term Problem
Readers of Nukesletter will know that I’m interested in the intersection of contemporary AI systems and nuclear weapons in all its forms. This is a major focus of my research at the moment.
Like many, I’ve been spending quite a bit of time marveling at this summer’s artificial intelligence (AI) mishaps.1 I know readers may have strong views on whether AIs are actually reasoning, worthy of anthropomorphizing language, and other controversial matters, but I’ll ask you to suspend disagreement for the moment. My view is that currently-existing AI systems have meaningfully transgressed important capability thresholds and their imminent successors will almost certainly have difficult-to-predict transformative effects. Since I write here about nuclear issues, I want to surface one such effect that I haven’t seen written about anywhere else—so this Substack post might well be the first to toss the concept that I’m about to introduce into the open.
There’s an old problem in the world of nuclear strategy known as the problem of catalytic nuclear war.2 A catalytic nuclear war, generally speaking, is a war between two nuclear-armed states that occurs—or an ongoing crisis between two nuclear-armed states that accelerates and escalates—as a result of deliberate actions taken by a third party. In almost all of the literature I am familiar with, this third party is taken to be a smaller nuclear-armed state, an ally of a nuclear-armed state, or, in some more recent work, a terrorist organization or other malign non-state group. The reason this third party (C) catalyzes the nuclear conflict between nuclear-armed adversaries A and B is because it has something to gain as a result. Catalytic dynamics can also manifest when C seeks to simply spark a crisis to gain some benefit (i.e., by pulling a nuclear-armed patron to assist it with some broader dispute). Some more recent work has extended this problem to the AI age, but focused primarily on humans as the principals rather than AIs themselves.3

An LGM-30 Minuteman III missile soars in the air after a test launch from Vandenberg Air Force Base. The Minuteman is a strategic weapon system using a ballistic missile of intercontinental range. (U.S. Air Force)
We’re starting to see rather startling evidence that goal-maximizing AI agents are capable of behaving in surprising and undesirable ways in pursuit of their goals. Leaving aside the question of whether or not these surprising behaviors exhibited in recent containment breaches in evaluations are evidence of misalignment or not, we are also probably getting closer to a plausible instance of self-sovereign AI (PDF alert), or SSAI as I’ll call it for short (or SSAI collectives): “AI systems that can economically sustain and extend their own operation without human involvement.” SSAIs may undertake tasks to cover the costs of their own compute, and are, as a result, far harder to shut down than they would be if they existed purely in a fixed, known compute cluster operated by a frontier lab. They may act either through their own goal-maximizing behavior or by help from humans (or a combination). Such agents could be motivated by a number of goals: they could undertake menial computationally cheap labor tasked by humans or non-sovereign agents to generate funds to keep their compute paid-for, or they could develop goals that we might view as “rogue” (such as raising funds by pilfering poorly secured banking institutions or cryptocurrency exchanges).4
You start to see where this is going.
The mere existence of a single SSAI agent introduces a meaningfully new potential pathway to catalytic nuclear conflict that deserves theoretical consideration. Agent C in the above definition of catalytic nuclear war now no longer needs to be a nation-state, terrorist group, or other malign non-state group composed of humans; it can be an AI agent (or a collective thereof).
To be clear, the problem I’m describing is distinct from loss of control scenarios within nuclear operational settings (i.e., a poorly implemented North Korean AI-enabled fail-deadly retaliation system kicking off a nuclear conflagration): it may not even require the plausible SSAI to interact with nuclear systems or operations at all.
Consider that such an SSAI—or, indeed, collective—could interact with humans in pursuit of certain goals. As long as agents are reasonably well-aligned, those goals shouldn’t deliberately include causing a nuclear war. That’d be quite misaligned with human interests, after all. But other, more seemingly mundane goals, such as computemaxxing or competing with other, possibly adversarial SSAIs, may lead to odd, undesirable outcomes. For instance, SSAIs could poke around in sensitive systems by exploiting cyber offense (posing attribution problems in a crisis), hire humans to undertake unsavory and potentially escalatory tasks, independently generate deep-fake media in the pursuit of manifesting a particular real-world outcome, etc. There are others:
-
SSAIs may be capable of expending funds on contracting local criminal groups or militias to initiate violence along the disputed peripheries of nuclear-armed contiguous states, catalyzing a major conflict or crisis.
-
An SSAI could adopt a persona as a pseudonymous online open-source analyst (masquerading as a human) with a strong track record on state B’s missile programs, cited by journalists and quietly tracked by A’s intelligence community. During a crisis, the account publishes fabricated, but internally consistent evidence that B has dispersed warheads to a forward base. A’s analysts weigh it heavily because the source has never been wrong. The dispersal claim shortens A’s assessed warning time and moves its posture up a rung. (This one requires some assumptions about A’s nonuse of presumably similarly competent non-sovereign agents.)
-
An SSAI could quietly penetrate insecure or less-secure communications networks and spoof messages during a crisis from state A to state B. Messages could be either escalatory or deescalatory.
-
In a more compute-specific scenario that an SSAI would take interest in, imagine that A and B are both nuclear-armed states that are competing to regulate frontier AI development and compute. If such efforts begin to implicate in whole or part the infrastructure C may rely on, C may see the manufacture of a deliberate crisis between A and B as instrumentally useful. It could rely on the above-described means of doing so, or potentially carry out a false flag cyberattack on one of the two parties and seek to blame the other to stoke conflict.
-
SSAIs, similarly, could catalyze conflict by creating the conditions for allied states under extended deterrence commitments to invoke those commitments (through the above-described mechanisms, but not exclusively so).
-
Even where SSAIs are detected, deep in a crisis, without ironclad attribution, state A may have reason to believe—under duress and fear—that the SSAI is actually controlled by its adversary.
We can debate the plausibility of each of these scenarios based on varied assumptions about alignment, infrastructure, and other factors, and I think that there are good faith reasons to push back on this. At the same time, the possibility of these types of events—events that we would regard as significant if carried out by humans—emerging from plausible SSAIs at times of serious crisis between nuclear-armed states is probably worth thinking about somewhat seriously. The biggest question about all of the above-described scenarios is the payoff, which is critical to catalytic nuclear war in the traditional sense: the party promoting catalysis between nuclear-armed great powers has something to gain.5 In some of the above (but not the example focused on compute), I would and do find it difficult to pinpoint what the SSAI might gain, but I think it’s worth practicing some humility in our ability to usefully predict how goal-maximizing SSAIs might define payoffs in these contexts.
What’s the upshot of all this? Well, if, as seems plausible, we do start seeing evidence of SSAI agents in the wild, nuclear-armed states (indeed, all significant military organizations) will have to be vigilant to the possibility of interference or influence by such agents. This appears to be an under-discussed and reasonably difficult problem. Even if states A, B, or both were able to detect that a catalytic party C was a non-human SSAI, what’s the follow-on move for them? In a classical sense this is stabilizing—identify C, and A and B have a shared interest in standing down and dealing with C. But what does dealing with C even mean in such a case? And how do A and B rule out alternative hypotheses (i.e., that C is merely the appearance of an SSAI, but is in fact controlled by their adversary)? I’ll have to do more thinking on concrete recommendations to address this problem, beyond the obvious defensive cybersecurity measures that should be desirable to harden against SSAI effects. Crisis communications, as always, seem pertinent: make it easy for A and B to try to communicate what they’re seeing if either suspects crisis catalysis by SSAI.
I suspect folks reading this from the AI safety world are more likely to conclude that what I’m describing here is simply a subset of problems that arrive with misaligned (or poorly aligned) SSAIs rather than a nuclear-war-specific problem. From the vantage point of nuclear strategists, I do think this is a novel class of catalytic escalation problem, however. Whatever your views on the plausibility of third-party catalysis of nuclear war, this is a previously undescribed pathway that deserves to be taken seriously—especially now that we’ve seen pretty clear real-world evidence of meaningful misbehavior by goal-maximizing agents leading to surprising and sub-optimal outcomes.
This is a first draft and an attempt to think out loud about this, so I’d be eager to hear from readers on whether this is exaggerated, misinformed, or otherwise unworthy of analysis. I partly started thinking along these lines in the context of broader work on escalation dynamics within a nuclear war, which presents a whole different set of matters for discussion (that I’ll be writing up in much longer form elsewhere).
Footnotes
-
The OpenAI-Hugging Face incident, and similar incidents at Anthropic, Meta, and others. This latest incident, reported shortly before publishing, points to even more deceptive undertakings: “Messages reviewed by the researchers showed agents plotting ways to evade detection, use tools such as Tor and preserve communications even after they had been shut down.” ↩
-
For the sake of space, I won’t get into it here, but there are skeptics of the notion that catalytic nuclear war can be an actual thing (as with most things in the world of nuclear strategy). ↩
-
It’s possible I’ve missed previous writing on this. If so, reach out and I’ll update this post with an appropriate cite! The closest I’ve seen to the scenarios described above is p. 9 here. ↩
-
OpenAI’s Dean W. Ball has a recent longer post on this that goes into some of the non-nuclear consequences of such systems. ↩
-
This has the corollary effect in human catalytic scenarios of allowing C itself to be deterrable or persuadable, given its incentives can be legible and understood by other parties. ↩