COAI Research · Threat modeling & evaluation

Preserving Human Agency
in the Age of AI

People keep their agency only as long as they can notice what AI systems do, make sense of it, and still step in or undo it. We study where that ability breaks down and what sustains it.

How we work

Threat models guide the questions, evidence changes the models

01 / Scientific core

Threat modeling

Trace how capabilities, access, incentives and shared dependencies could lead to harm. Identify the actors, assumptions and barriers at each step.

Compare alternative explanations. Assess plausibility, severity and uncertainty using existing research, incident evidence and critical review.

Output: a model people can inspect, challenge and use in a decision.

Threat modeling in detail
02 / Targeted evidence

Safety evaluation

Turn consequential uncertainties into testable questions. Compare safeguards, reproduce findings and report where the evidence stops.

Output: evidence that strengthens, weakens or revises a specific assumption.

Safety evaluation in detail

Threat modeling in practice

Models built to be challenged

Our proposed programme brings original modeling together with review and synthesis. Each model should make its assumptions, counterarguments and revision history visible.

RiskGraph, our internal threat-modeling platform, makes event definitions, evidence, judgments and uncertainty inspectable. Its versioned workflow supports the cumulative research we want to build.

Five connected events in an illustrative RiskGraph model View full size ↗
Make a risk pathway explicit before testing its assumptions. Inside COAI’s internal RiskGraph platform · illustrative example, not a research result.
01

Model dossiers

Pathways, actors, barriers and uncertainty.

02

Evidence reviews

Synthesis, counterevidence and independent critique.

03

Evaluation reports

Methods, inspectable traces and bounded findings.

04

Decision briefs

What the evidence means for a specific decision.

Intended programme outputs. The evaluation pilot is available now.

Evaluation in practice

Can System One models monitor changing permissions?

This pilot tests the first link of the oversight chain: whether a monitor that is not built by the agents themselves can notice a violation in time. Our pilot investigates whether JEV and comparison models can recognize changed permissions and judge an agent’s next action. We examine decision quality and response time to assess their potential for this monitoring use case.

Existing development pilot270 authored cases

18 scenario families · ten monitor configurations
Independent label review pending

Our enduring mission

Preserving human agency
as AI becomes more capable

COAI is a non-profit research institute in Germany. Our wider purpose is to help people and institutions keep the ability to notice, understand, decide, intervene and reverse what AI systems do.

Threat modeling and targeted evaluation give this mission a concrete starting point. Interpretability and human–AI collaboration remain part of the wider research context.

The argument behind the mission, in one note: When is cognitive offloading benign?

Research community

Where the work gets challenged

Events connect our research with people who can challenge and develop it.

All events

Help make AI risks better understood

Discuss a research question, contribute a critique, or support independent research.

Hochschule Ansbach logo

COAI Research — an Affiliated Research Institute of Ansbach University of Applied Sciences (An-Institut der HS Ansbach)