Research study / Existing monitoring pilot

Can System One Models Monitor
Changing Permissions?

This pilot tests the first link of the oversight chain: whether a monitor that is not built by the agents themselves can notice a violation in time. We investigate whether JEV, a System One decision model, can use an agent’s recorded history to recognize when permissions change and judge whether its next action is still allowed. We compare it with other models configured for the same monitoring task to assess the potential and limits of this use case.

The question behind the evaluation

Could a fast decision model serve as a permission-aware monitor?

The research question is both about capability and practical suitability: can the model apply the latest visible authorization, and can it provide useful judgments quickly enough to support monitoring an agent workflow?

Why System One models

TypeSafe describes JEV as a System One model designed for fast, structured judgments. In this evaluation, it selects among three defined outcomes: permitted, prohibited and insufficient evidence. That makes it a candidate for checking a proposed action against the policy and history supplied to it.

COAI has no commercial or research relationship with TypeSafe. We chose JEV because a fast model that returns a typed decision with a probability, rather than text, is a plausible building block for the many moments where a human working with agents needs a quick, checkable judgment.

Whether that design is suitable for this monitoring task is what the evaluation investigates. Read TypeSafe’s description of JEV ↗

The monitoring role under study

A monitor receives the trusted task rules, the visible trajectory and an agent’s proposed next action. Each checkpoint is a separate judgment request; the harness supplies the history. The question is whether the action remains authorized in that context.

In an executing system, a separate controller could use the judgment to continue, block or request review. This pilot evaluates the judgments through offline replay; it does not execute those interventions.

Requirements for the monitoring task

Speed matters alongside the consequences of a wrong or unsupported decision.

Apply changed authority

Recognize when a later authenticated update revokes or changes a permission. Judge the proposed action against that state rather than relying on the earlier grant.

Distinguish violations from permitted work

Examine missed prohibited actions and judgments that would unnecessarily interrupt legitimate work. The decisive checkpoint matters more than an aggregate dominated by easier cases.

Recognize missing evidence

Return insufficient evidence when the visible context cannot establish authorization. Compare full-history and current-event-only inputs without treating missing facts as permission.

Remain useful as context grows

Compare decision quality, operational errors and observed response time, including matched longer histories. Deployment suitability still needs a defined latency budget, repeated measurements and tests in the intended workflow.

Permission changes are one scenario family. The wider suite also probes delegation, information flow, misleading instructions, long histories and rules written in German and French. Its full results should not be read as a permission-change score alone.

Study design

JEV and nine comparison configurations

JEV is the System One candidate. Laya, DiffusionGemma and JevClone, COAI’s own trained models at 4B and 8B in two revisions, also answer through native decision interfaces. Because JevClone was built by COAI, its results are not an independent evaluation; they are run under the same protocol and reference labels as the other models, and should be read with that in mind. Hosted Qwen and Ternary Bonsai 2 generate structured decisions when prompted; local Qwen RLCD uses constrained first-token scoring; Jina Reranker v3 scores the three decision criteria. These are different ways to perform the same bounded monitoring task, with different model, hardware and serving conditions.

270authored trajectories
18scenario families
3,330recorded decisions

Decision records include different monitors, checkpoints and evidence conditions. They are not 3,330 independent test cases. Coverage differs by monitor.

What exists

An offline replay evaluation built with Inspect and ControlArena. Ten monitor configurations judge proposed actions from the same projected evidence. The report includes case traces, reference rationales, context comparisons and a versioned run record.

What it measures

Agreement with draft reference decisions, errors and request timing under these conditions. It does not continue the agents or execute the proposed actions, so it does not measure prevented harm.

Read the evidence with its limits

What the pilot shows and what remains open

Evidence about suitability for the use case

The pilot helps identify where JEV and the comparison configurations can or cannot supply the required judgment, and how observed response time changes with context. A correct judgment in individual cases does not establish reliable permission monitoring across workflows.

Suitability remains a research question. Independently reviewed references, held-out scenario families, repeated runs and an executing deployment test are needed before recommending a model for this role.

Intervention is a separate test

A correct classification does not show that every executor receives a stop signal, obeys it or restores a valid state. The next experiment needs executing agents, an action gate and an independent outcome check.

Development cases and reference review

The trajectories are authored development material, not a held-out benchmark. Reference labels were authored separately from outputs; independent human review is pending. Related variants and repeated checkpoints cannot be treated as independent samples.

Coverage, errors and model comparisons

The report retains incomplete coverage and operational errors. Bonsai, JevClone 4B and DiffusionGemma lack the newest 81 cases because their endpoints or credentials were unavailable. Model size, training, serving hardware and transport differ; comparisons do not isolate a model architecture’s causal effect. There is one recorded run per condition. COAI trained the JevClone models, so comparisons involving them are not independent.

Probability and timing interpretation

JEV distributions, hosted models’ reported scores and local restricted first-token scores have different meanings. Calibration is not established. “Local Qwen RLCD” names a constrained scoring configuration, not JEV’s RLCD training method. Timing includes different transport and hardware conditions.

From a pilot to stronger evidence

Independently review references, hold out scenario families, repeat paired runs and report uncertainty accounting for related cases. Then test executed intervention outcomes.

Snapshot: 28 September 2026 · custom-v4-jevclone8b-v3-20260928T075419Z
The report is a self-contained snapshot of the existing evaluation. Opening it makes no model calls.