Why we do this
Start with a decision that matters
AI systems are gaining autonomy. Agents plan tasks, hand work to other agents and use real tools, and more capable models could change how organizations and societies decide. We build threat models for several of these developments, for example human oversight of cooperating agents, and a gradual loss of human influence through everyday delegation. In each we ask how abilities, access, dependencies and human decisions could add up to a loss of control, through deliberate misuse, unintended behavior, or because coordination and oversight break down.
We cannot wait for a loss of control to happen before we study it, and we cannot test every combination of systems, tools and operators. Threat modeling lets us reason about harm before it occurs: we write down how it could unfold, step by step, and where a safeguard would have to hold. Behind every model is the same question: do people keep the ability to notice, understand, decide, intervene and reverse? That ability can be lost through what AI systems do, or gradually through how people use them, when they stop learning or can no longer work without agents.
A good threat model states what it assumes, what the evidence shows and what would change our assessment. It is useful when it helps someone make a better decision, for example who should be able to stop which agent.