Detecting Marginal-Preserving AI Sabotage
Abstract
AI-control monitors score individual agent actions to detect misbehavior, but real harm can be _distributed_ across many benign-looking steps, each individually below any per-step alarm. We construct a **marginal-preserving, correlation-encoded distributed-sabotage attack** using a Gaussian-copula AR(1) construction: the per-step monitorscore _marginal_ is held exactly equal to benign, so mean, max, top- _k_ tail, and threshold monitors (Monitor A) are defeated _by construction_ , while harm is encoded in the _temporal correlation_ structure. We sequence the paper around three reviewer-mandated gates. (1) **Realizability gate:** the stealthy attack achieves KS-distance to benign of 0 _._ 013 (effectively zero) at all tested harm levels up to 3 _._ 0, confirming that harm is fully decoupled from the per-step marginal and realizability is not harm-limited. (2) **Monitor-A-vs-B reconciliation:** we show formally that the attack, built against Monitor As score marginal, remains marginal-preserving under a _different-score_ Monitor B (the correlation/sequence family: CUSUM, SPRT, HMM-LR, runs test, autocorrelation, windowed logistic), and scope worst-case claims to score functions that admit a temporal signature. (3) **Non-empty detectability band:** Monitor A achieves AUC __ 0 _._ 52 (chance); Monitor B spans AUC 0 _._ 790 _._ 97 at the same 1% FPR target, and as harm is amortized over more steps Monitor A collapses to chance while Monitor B holds at AUC __ 0 _._ 95. These results demonstrate a non-empty detectability band and characterize the sub-threshold sabotage frontier: _distribution-shape monitors fail by construction; temporal-correlation monitors can detect but are not trivially optimal_ .