Calibration Overfit in Trusted AI Sabotage Monitors

Arxiv pdf 2026-07-01T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

Trusted monitoring is a central defense in AI control: a cheaper trusted model scores an untrusted models actions for sabotage, and the most suspicious are audited or deferred. Such monitors are evaluated against one or two untrusted models, and the resulting accuracy is reported as a property of the monitor. We ask whether it is partly a property of the _pairing_ . We make the untrusted policy family the controlled axis: we fit a monitor on family As transcripts, apply it to family B, and decompose the cross-family AUROC into how obvious each familys sabotage is, how capable each monitor is, and the residual own-family advantage after both are removed: the interaction. A positive interaction means a monitor ranks its own lineages sabotage above an attack from a different lineage after those row and column main effects are removed; in our matched controls, this remains true when attacks are matched on two obviousness proxies. On code-backdoor transcripts the interaction is positive and survives the dominant confounds: **+0.172** (95% CI [+0.158, +0.185]) on four open-weight families on a strict leak-free basis, holding across an eight-family matrix spanning four frontier vendors, blatancy-matched controls, a trained openweight transcript-reading monitor up to 34B parameters, and a second code benchmark. The gap is geometric: each lineages sabotage occupies a different scoring direction, and rotating the monitors scoring direction toward the target lineage recovers most of the lost detection while a matched random rotation recovers nothing. Translated to a defer-to-trusted protocol at a 1% audit budget, an off-lineage monitor catches **19%** of attack attempts where an in-lineage monitor catches **41%** at the same audit cost, a +21.7-point gap on the strict leak-free basis (42% versus 24% on the transductive basis). We measure attack-attempt detection as the primary axis; realized sabotage is frontier-only on this substrate, too rare among the open families to power a full matrix, though on the frontier lineages that do realize it the interaction survives and is, by effect size, if anything sharper. We flag the absence of an all-family realized-sabotage matrix as a limitation. The takeaway for practice is that control evaluations should report cross-family transfer matrices, because a single-pairing accuracy overstates the safety a monitor delivers against a lineage it was not calibrated on; the four-step protocol that acts on the gap (report the full matrix, detect off-lineage drift, recalibrate with labeled sabotage, treat pooling as a bounded partial defense) is in Section 5.

Loading executive summary...

LINK COPIED TO CLIPBOARD