search
Contents
  1. Good verification means avoiding worst-case thinking and planning
  2. Definitions and mechanics are really annoying (and critical)
  3. Prepare technical work now, even if the political opportunity seems unlikely
  4. Assemble the technical advisors
  5. Proposal: expert verification exchange

Editor’s Note: Greetings, Nukesletter readers. Today’s a first: a co-authored Substack with my friend and longtime co-conspirator James McKeon on what the AI community can learn about verification from the nuclear weapons world (and why the two communities need to talk more). James is cooking up a very exciting new project that we say more about toward the end of this piece. This piece is best read as a companion to my recently published Foreign Affairs essay with Chris LaRoche on the trouble with the AI-nuclear analogy.

tl;dr: AI verification community, we love you (even if your fears are “becoming a panic”), though sometimes we don’t love your nuclear-AI analogies. You’re doing sorely needed technical work, but even the best technical work may run into an inherently human problem: translating it to diplomatic agreement and durable institutions. On that, we think the nuclear experience can help.


There are a lot of very smart people doing good and important work on various forms of verification pertaining to AI (more on this below). These are instrumental to making various concepts of AI safety workable in the real world (including calls to “pace” the frontier). It’s not useless cynicism to ask how we would know whether others (domestic firms and international competitors alike) would honor any agreement proposed. Verifying compliance with various forms of negotiated restraint is tough! In what follows, we offer a few thoughts from the world of nuclear weapons—where we both have experience—for today’s AI verification thinkers. (It’s one of the places the otherwise troubled AI-nuclear analogy does offer some useful guidance!)

U.S. and Soviet flags fly side by side atop the experimental tower at the Nevada Test Site during the first of two Joint Verification Experiments. (Image via Lawrence Livermore National Laboratory.)

U.S. and Soviet flags fly side by side atop the experimental tower at the Nevada Test Site during the first of two Joint Verification Experiments. (Image via Lawrence Livermore National Laboratory.)

In the nuclear weapons world, there are primarily two major1 forms of verification that are thematically and practically distinct; often, AI thinkers borrowing from the nuclear analogy appear to flow freely between both.

First, there is the nonproliferation verification function of the International Atomic Energy Agency. This is best thought of as intra-state verification. Its purpose is to determine that civilian nuclear energy programs are civilian in nature. The IAEA’s core verification purpose is to provide timely detection of the diversion of significant quantities of nuclear material that could be weaponized. Failing to submit to safeguards or being caught in violation of safeguards constitutes noncompliance. AI thinkers usually apply this model to intra-state applications of training or inference compute.

Second, there is the arms control verification function. This is best thought of as inter-state verification and its aims can be various: producing transparency, providing assurances that numerical controls on arms won’t be exceeded, or that nuclear weapons won’t be tested (or tested only within certain parameters). There is no generalizable form, but verification in this context involves an agreed protocol between states (ideally, a legally binding treaty). The purpose of verification then is for the parties to such an agreement to determine whether observed behaviors are consistent with the restrictions mutually agreed. Observed inconsistent behaviors–say, the deployment of weapons that are not permitted–result in noncompliance. This form of verification can either be extremely intrusive (like the verification of the INF Treaty), or entirely remote (the Comprehensive Nuclear Test Ban Treaty’s International Monitoring System).

Finally, it’s worth flagging that while the history of nuclear verification has been limited by the technological availability of verification capabilities2, the primary obstacles to robust verification have been political (both domestic and international). This is likely to remain the case and will be something ambitious technical proposals for AI verification will have to contend with.

Below, we offer a few principles for thinking about verification–its desirability, its definability, and its tractability–with a mind for readers coming from the world of AI governance debates.

Subscribe now

Good verification means avoiding worst-case thinking and planning

With consequential transformative technologies, governments are charged with reasoning under uncertainty. The greater the consequences of the use or misuse of a given technology, the less it becomes possible for governments to excise certain worst-case scenarios from their thinking. In the nuclear domain, this is why fear of a Soviet first strike, for instance, never really went away from animating American thinking and strategic planning. (Or fears of an American first strike today for the Russians and the Chinese.) In the AI domain, there is no shortage of worst-case scenarios, but fears of superintelligence-gone-rogue and adversaries crossing key thresholds first tend to animate most contemporary fears. In short, uncertainty about an adversary’s capabilities can encourage costly and dangerous competition; avoiding these costs and scoping competition to be as prudent as possible, as a result, is desirable!

Verification helps collapse those uncertainties and increase transparency as a result. The New START treaty, that ended earlier this year, codified numeric limits for the strategic nuclear forces of the United States and Russia, but its verification measures—data exchanges, notifications, and on-site inspections (before COVID-19, at least)—helped provide transparency and predictability that gave both Washington and Moscow the confidence to reason about their own defense policies with less uncertainty. If one of the two were to seek to “break out” of limits, there would be detection and adequate time to respond.

For AI (particularly the inter-state anxieties animating U.S.-China AI competition), a lesson is that verification can pay off well short of approaching perfect insight. Neither Washington nor Moscow knew everything about the other’s forces under New START; each knew enough to set aside the scariest possibilities and plan against a realistic threat. AI governance today resembles the strategic nuclear relationship since the treaty lapsed: absent credible insight into a rival’s compute stocks, training runs, or model capabilities, planners in Washington and Beijing (and inside frontier labs) will fill the gap with the worst case, which in AI usually means assuming the other side is months from a decisive capability. That assumption then justifies racing, cutting safety corners, and treating any restraint as unilateral disarmament (as accelerationists argue today). Even modest measures, such as declarations of large training runs, data exchanges on advanced chip holdings, or notifications ahead of major deployments, can narrow the plausible range enough to take the tail scenarios off the planning table.

Definitions and mechanics are really annoying (and critical)

Better verification tools give negotiators more options. But the agreement they reach will probably include compromises that make technical experts pull a few hairs out. A U.S. B-52H can carry 20 nuclear-armed cruise missiles, yet counted as just one warhead under New START. That may seem… odd. The United States wanted to count the actual weapons that would go on the bombers. That would have involved inspecting weapons storage facilities, which Russia resisted. The compromise was to use an agreed counting rule rather than count the actual weapons: each deployed nuclear-capable bomber counted as one warhead, regardless of its payload.3

Even agreeing to count warheads on ballistic missiles opened a lot of questions. As chief U.S. negotiator Rose Gottemoeller describes, negotiators had to settle the pliability of the covers over them and what radiation detectors could reveal without exposing design secrets. The definitions and methods really mattered. In fact, the term “warhead” was largely a counting rule in U.S.–Russian arms control, not a clearly defined physical term.

AI experts might reasonably assume they agree on what “training” or “inference” means. But then try writing a bunch of restrictions around those. Would they cover pre-training, post-training, or both? Which computations count toward the limit? A limit on a single training run, or all the work that goes into developing the model? And then there’s our pesky friend recursive self-improvement: defining this in a way everyone can agree upon seems like no easy task. For “dark compute,” how would you agree on procedures to declare hardware, report potential undeclared hardware, and check its use? (Maybe some of this seems obvious. Getting everyone to agree on the fine details won’t be so easy.)

Then comes the hardest task: getting the people across the negotiating table to accept the same definitions and methods.4 The five nuclear-weapon states established a working group on nuclear terminology in 2011. They presented their glossary in 2015 – an inspiring 70 years into the nuclear age.5

NIST’s glossary and the Brookings-Tsinghua glossary give us somewhere to start for AI. They don’t cover technical terms in a way that governments might use to sign an agreement. Working through technical disagreements now could give governments a starting position in negotiations. Better than trying to figure out everything in a crisis.

Prepare technical work now, even if the political opportunity seems unlikely

Just because verification–in its negotiation and application–is political does not mean that technical experts should not do their work as soon as they can. In fact, it is best for shelf-ready verification technologies, protocols, and modalities to exist so that negotiators can simply pick them up when presented with an opening. In other words, technical preparation can clarify which agreements are feasible. It cannot decide how much uncertainty governments will accept or what access they will grant.

The history of verifying nuclear test bans is a pretty useful case in point here. Initially, in the late 1950s through early 1960s, banning underground nuclear tests appeared to be seismologically intractable; this helped steer the 1963 Limited Test-Ban Treaty toward atmospheric, underwater, and space-based tests (these were definitely big problems at the time!). The period had its share of over-optimism, too. For instance, initial optimistic assessments (by the Geneva Conference of Experts) were undercut as American experts raised concerns about “decoupling” (in which detonating a nuclear device inside a large underground cavity could muffle its seismic signature6). Verification doubts dogged test-ban diplomacy for decades; disputes over measuring yields kept the 1974 Threshold Test Ban Treaty unratified until much later, when the 1988 Joint Verification Experiment, in which U.S. and Soviet experts measured each other’s explosions in Nevada and Semipalatinsk, cleared the way for its ratification in 1990. Back in 1959, meanwhile, the United States launched Project Vela, a Pentagon-funded (ARPA) program that paired seismic research (Vela Uniform) with satellites built to spot nuclear detonations in the atmosphere and space (Vela Hotel), which went on to help monitor Soviet compliance with the LTBT.7 Today, the dynamic is somewhat flipped; the Preparatory Commission for the Comprehensive Nuclear-Test-Ban Treaty Organization operates a globe-spanning, distributed International Monitoring System, designed to detect minute underground nuclear explosions, but the Treaty itself is still not in force, 30 years after it opened for signature. The lesson for anyone doing (or funding) verification work under unforgiving political circumstances: the technical groundwork can outlast stalled politics. One of us (Ankit) convened a group of experts in 2021 to look seriously at employing novel technologies to verify various potential agreements that might be reached with a nuclear-armed North Korea despite the grim prospects at the time for any serious diplomacy.

There are three big lessons here for AI verification work today. First, it’s worth investing in the basic science and people-building business of verification now, even with no obvious agreements or political breakthroughs in view. Second, it is useful to be candid about the limits of verification. Late-1950s overconfidence about test-ban verification had to contend with the “decoupling” fears; similarly, AI verification work on everything from hardware-enabled mechanisms, chip location verification, attestation of training runs, and workload classification will have to be attuned to the limits of what can and can’t be known (and with varied levels of confidence). Third, it’s useful to collaborate with the adversary to build political confidence in verification technologies. This is the essential insight of the Joint Verification Experiment, arguably one of the brightest points in the history of nuclear verification efforts for underwriting how American and Soviet technical communities could have a shared epistemic view of verification. Current geopolitical realities make similar joint U.S.-Chinese efforts on, say, on-chip verification difficult to imagine, but dialogues about the potential of such efforts would be worthwhile.

Assemble the technical advisors

Technical experts need to be in the room for verification negotiations, explaining what is possible and, frankly, helping diplomatic teams make sense of the details.8 9 In nuclear arms control, much of that expertise came from government agencies and national laboratories.10 For AI, much of the expertise sits in frontier labs or other companies, as well as universities and independent research organizations. But having the expertise and being ready to advise a negotiation are very different things. We have ideas for preparing that technical advisory group now, but that’s for another post.

For now: One way to prepare those potential advisors is to connect them with people who know nuclear verification.

Proposal: expert verification exchange

We think the easiest thing to do is get some of the people who negotiated nuclear agreements or are otherwise experts on nuclear verification regimes into a room with experts working on AI verification. There’s plenty to discuss and debate. Did you ever have to scale back a verification proposal because the two sides couldn’t agree on the details? How did you decide whether what was left was still good enough? How did you get enough access to check what mattered without giving away sensitive information? How do you explain what the tools can—and can’t—verify to negotiators who aren’t technical experts?

Global Risks Fund (which was recently started with James on the founding team) is preparing to fund and convene an initial exchange. We’d like people to come away with a better feel for what making verification work involves beyond technical progress, some ideas to try, and some new advisors they might be able to call. You don’t need the highest P(doom) to think we should get this going now. Take your P(doom) and turn it into P*(verify)*!11

Footnotes

  1. We say the two above are the major applications of verification, but there are other adjacent areas that should matter. For instance, nuclear-relevant export control regimes have their share of verification problems. ↩

  2. One of the most famous examples of this is the concerns in the United States about the lack of verification for the 1967 Outer Space Treaty, which banned the placement of nuclear weapons in orbit. Recently, some optimism about the verifiability of nuclear weapons in space has started to emerge. ↩

  3. In the end, the U.S. military learned to live with this particular outcome in New START; the Russians tolerated others (for instance, concerns about submarine tube conversions). ↩

  4. If you and your friends can’t agree on whether a hot dog is a sandwich, you’ll probably recognize the issue here. ↩

  5. Not to be too nerdy, but there still isn’t a universally accepted definition of a “tactical” nuclear weapon. ↩

  6. This is exactly at the heart of recent U.S. accusations that China is testing nuclear weapons—allegedly by decoupling. Ankit wrote about this earlier this year in Nukesletter here. ↩

  7. A Vela satellite was responsible for instigating one of the greatest mysteries of the nuclear age in September 1979 (not a mystery to many; IYKYK). ↩

  8. James: Ever tried reading a telemetry annex? Don’t answer that, Ankit. ↩

  9. For examples of current work on AI verification and oversight, check out Apollo on detecting deceptive behavior, Lucid on compute verification, AVERI on pilot audits, and Transluce on understanding model behavior. ↩

  10. One of us–cough: James–keeps forgetting to specify which “labs” he means in policy discussions. In Washington, that usually means Los Alamos, Sandia, or Livermore; in San Francisco: OpenAI, Anthropic, or Google DeepMind. See? Defining your terms really matters. ↩

  11. Did you cringe a little? Good. ↩

Newsletter

This post first appeared in Nukesletter. Subscribe to get new posts by email.