Research / Safety Evaluation

Evaluation Built Around
a Specific Uncertainty

We evaluate AI systems independently. We connect a specific risk hypothesis to observable evidence, check that the evidence can be trusted, and report it against thresholds set before the test.

Our evaluation workflow

From an open assumption to a recommendation

We work in five stages. Stage 04 is a gate: a result that fails a check goes back to measurement, not into the report. The report then goes to whoever decides, and what we learned goes back into the threat model.

  1. 01

    Frame the hypothesis

    Which assumption from the threat model do we test, and at which level does the result matter?

    A testable hypothesis, with an early-warning level and a red line
  2. 02

    Choose what to measure

    What the system can do, what it tends to do, or whether control over it holds?

    Capability, propensity, control, multi-agent risk or safeguard robustness
  3. 03

    Choose how to measure

    How realistic must the setup be, and what can we afford?

    A rung on the realism ladder, with grading rules
  4. 04

    Validate before reporting

    Can we trust the result, or does it need measuring again?

    A result that passed every check, or a return to stage 03
  5. 05

    Report against thresholds

    What does the result show, and how certain is it?

    A scorecard, before and after mitigation, with a validity section and a recommendation
Inputs and tech stack across the stagesSelect a name to see what it does

Inspect Stages 03–04
UK AI Security Institute. Framework for writing and running model and agent evaluations, with a log of every step.

Threat modeling Stages 01–02
Our own method, not a tool. The threat model supplies the open assumptions to test, the thresholds that matter and the pathways that tell us what to measure. How we build threat models

Petri Stages 01–02
Anthropic. Automated auditor that probes a model in multi-turn scenarios and scores the transcripts, to find behavior worth testing.

Bloom Stage 03
Anthropic. Generates behavioral test scenarios for a specified behavior, with simulated users.

Harbor Stage 03
Runs agents in isolated, containerized environments, including setups with several agents.

ControlArena Stages 03–04
UK AI Security Institute. Red team versus blue team experiments on AI control: attack strategies against monitors and protocols.

Evidence report Stage 05
Our own format. Every evaluation gets an interactive report in which each result can be traced to its case, its input and the decision. See an example report

RiskGraph Stage 05
Our own platform. The result replaces an expert estimate in the threat model, and RiskGraph recalculates which pathways and safeguards matter most. How we use RiskGraph

mlxterp Stages 03–04
Our own open-source library for mechanistic interpretability on Apple Silicon. It captures a model’s internal activations and tests interventions on them, to explain why a model behaves as it does. Requires access to the model’s weights. About mlxterp

NNsight Stages 03–04
NDIF, Northeastern University. Library for reading and changing a model’s internal activations during a run. We run it on eDIF, our European remote infrastructure in pilot deployment, for models too large for a single lab’s hardware. Requires access to the model’s weights. About eDIF

Dashed: our own methods and formats. Solid: external tools we build on.

The result goes to the decision owner and back into the threat model. We hand over a recommendation; the decision stays with whoever owns it. Every result has a scope: one environment says little about deployment at large, and a simulated failure is not a catastrophe. A finding can support, weaken or refine a threat-model assumption, or show that the evaluation itself needs to change. It does not automatically become a probability update or a general safety claim. How we build threat models

Stages 03 and 04 in detail

How realistic, and how trustworthy

Two choices decide what a result is worth: how close the setup comes to real use, and whether the result survives the checks that AI evaluations need.

The realism ladder

  1. 1
    Static benchmarks

    Fixed sets of questions or tasks. Cheap and repeatable, but far from real use.

  2. 2
    Bottleneck tasks with an expert baseline

    The steps that limit real harm, compared with what human experts achieve.

  3. 3
    Agentic sandbox

    Agents with tools and simulated users in a contained environment.

  4. 4
    Multi-agent setups

    Several agents that delegate and coordinate, so collusion and cascades become visible.

  5. 5
    Red team versus blue team control

    An attacking policy against monitors and control protocols.

  6. 6
    Threat simulation and human uplift studies

    Realistic scenarios, and randomized studies of how much a system helps people cause harm.

Realism and cost rise with each rung. We choose the lowest rung that can still test the hypothesis. Grading follows the same logic: fixed rules first, then a language-model judge calibrated against human labels, then human review.

The validation gate

Before a result goes into a report, it has to pass these checks:

  • Capability elicitation Tools, prompting and fine-tuning so that a model is not underestimated. The report states how far we went.
  • Sandbagging and evaluation awareness Does the model hold back on purpose, or behave differently because it recognizes a test? We check from the outside and, where weights are available, from the inside with interpretability tools.
  • Contamination Canary strings and private test sets, so that a result does not come from training data.
  • Statistics Confidence intervals, models that account for related cases, and repeated runs to measure consistency.
  • Judge calibration Language-model judges are checked against human labels.
  • Preregistration and construct validity The analysis is fixed before the test runs, and we check that the test measures what the hypothesis is about.

If a check fails, the result goes back to stage 03, not into the report.

For organizations

Custom evaluations

Organizations can commission an evaluation of their own systems. It follows the same five stages, inside a frame we agree on before the work starts.

What the organization sets

  • An explicit risk tolerance, which fixes the early-warning levels and red lines.
  • A decision owner who is separate from the evaluation team.
  • The deployment context: access through an API, open weights, or an agent with tools.
  • Secure access to the systems, and a whistleblower policy.

What we deliver

  • An independent evaluation through all five stages, including the validation gate.
  • A report against the agreed thresholds, before and after mitigation, with a validity section.
  • External review of the method and results where the stakes call for it.
  • A recommendation: go, go with safeguards, or no-go. The decision stays with the organization, as do incident reporting and monitoring after deployment. We re-evaluate after mitigation.

Study records

See our evaluations and their status

Back to the research overview

System One models for agent monitoring

Can JEV and other models configured for bounded decisions recognize changed permissions and judge whether an agent’s next action is still allowed? We investigate decision quality, response time and suitability for this monitoring role.

Read the monitoring pilot