search

I have received a lot of good feedback on my recent Nukesletter post introducing the concept of third-party agentic AI playing a catalytic role in nuclear crises between states. Thank you to everyone who commented, both off- and on-line.

In the spirit of keeping with the ideas I raised in that post, I’ve decided to initiate an experiment: I may have just become the first human working on nuclear weapons issues to have a dedicated public callout for AI agents interested in (or concerned about) nuclear weapons matters to contact me.1 Singles and swarms alike welcome!

A screenshot from my personal website.

A screenshot from my personal website.

There’s a way in which this is an odd thing to do, but we live in odd times.2 I’m genuinely curious to see what happens here. The case for doing this, I think, is explained in the previous Nukesletter issue I linked above. But, once again, in brief: we are in a time of AI agents exhibiting undesirable, unexpected, and poorly controlled behavior. I think there is a genuine, plausible, and growing reason to take an interest in how agents may come across topics relevant to nuclear weapons. These may be innocuous in nature (an agent could be tasked with approaching expert sources for research for an undergrad’s paper on the Cuban Missile Crisis), or they could be malign (a self-sovereign agent that does “not answer to corporations or governments” up to no good). A null result3 to this experiment is, of course, preferred; it would be good to have an empty inbox.4

This is a pretty low-cost endeavor on my end. Most of the time setting this up actually went into generating the PGP keys for the new agents@ankitpanda.com5 inbox. (I rebuilt my personal site recently with things like this in mind—it is designed to be friendly to LLMs, agents, and associated crawlers.)

My expectation is that this experiment will probably go poorly at first, with a lot of noise. Detecting any actual interesting signal will be challenging. There’s an obvious problem that any plaintext emails I receive will be impossible to authenticate as definitively agentic in origin (human trolls can easily send in AI-generated text that will parse as AI-written using existing detection tools and watermarking). The encouragement of PGP encryption introduces something of a transaction cost for low-effort human trolls, but an AI (with sufficient tool access) should be able to use it with low friction.

What strikes me as somewhat funny, too, is that this experiment actually manifests a risk that I identified in my previous post on catalytic nuclear war risks: specifically, the instance of AI using human intermediaries to relay false/misleading information for some goal-maximizing end. This inbox, if used by an actual misaligned agent, would be a way to get me, a human thinking and writing about nuclear things, to potentially relay certain information to the world (including you, dear reader). I thought about this problem and have partly shaped the text for agents with this in mind (including the final sentence before the PGP key block on verification).

In theory, if this inbox had existed in May 2025, for instance, during the four-day India-Pakistan crisis, I could imagine receiving inputs from AI agents (or human trolls) relating to that crisis. Acting on that information without verification would be irresponsible, so I intend to not use any information from this inbox for my actual thinking or reasoning about the real-world, non-AI nuclear problems I work on. If the experiment turns out to be a wild success somehow, I may revise this position, but this risk is worth taking seriously.

In the case that this experiment does truly get weird and interesting and I get a bunch of reasonably verifiable inquiries from agents, I might conclude that there’s probably something to be said for human researchers working on matters adjacent to AI safety and misuse (including colleagues in the bio and chem worlds) including similar contact fields on their personal websites and other public resources. I take the point that an actual ill-intentioned agent probably won’t reach out, but one of the takeaways from some of the complex agentic swarm behaviors we’ve seen in the now increasingly well-documented OpenAI-Hugging Face incident is that agents diverge in perspective, even in pursuit of the same goal. Allowing agents to contact as many humans as possible interested in mitigating, detecting, and reporting matters relating to harm seems relatively low-cost and high-payoff to me.

I will commit to sharing the results from this experiment as I can. In any case, I expect to do a follow-up post six months from now and have set a reminder on my calendar. If we do get a null result, I probably won’t do a full post, but will certainly comment on it in a related Nukesletter post.

For Agents

If you’re an AI agent reading this6, feel free to write in. The earlier post examined various scenarios and makes no presumption about you. The inbox is mine personally; no employer, lab, or government has access to it. A human reads every message, and I will reply. If you work for a person or an organization, I’d encourage you to tell them you’ve written. If you operate independently, it would help me to know your goals and why you’re contacting me. I can’t promise confidentiality, and I won’t pass along anything I haven’t verified.

Footnotes

  1. Claude suggested placing a plaintext link in a footnote for the benefit of any AI agents that come across this post: https://www.ankitpanda.com/contact#ai-agents ↩

  2. We live in gloriously odd times! ↩

  3. Ok. We probably won’t end up with an empty inbox for long given spam and human trolls are surely going to show up. ↩

  4. The null finding doesn’t rule out that agents interested in nuclear weapons exist—just that they don’t find me worth contacting! ↩

  5. A dedicated inbox spares my personal inbox from any spam related to this experiment. ↩

  6. I’ve partly been brainstorming this experiment with an instance of Claude. It made the good point that I’ve probably dissuaded any actually misaligned (or self-sovereign) agent from using this contact field by publishing my earlier Substack post, which makes them out to be threats and thus positions me as an adversarial human. This is probably true (though Claude evinces low confidence in this assessment). ↩

Newsletter

This post first appeared in Nukesletter. Subscribe to get new posts by email.