LLM Safety Classifier Shift Detection

Arxiv pdf 2025-12-10T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

We present an online monitoring system for distributional shift in deployed safety classifiers, using calibrated sequential statistics on KS test statistics to detect when a classifier has moved out of distribution. Upon detection, a conformal abstention layer adapts decision thresholds to recover a target error rate __ = 0 _._ 1 when density-ratio estimation is effective. In a pre-registered factorial evaluation across 4 classifiers __ 5 shift conditions __ 20 seeds __ 2 window sizes (800 cells), the system achieves an 86.6% valid detection rate (693/800 cells, 95% CI [84.1%, 88.8%]), with mean detection latency of 39.5 steps at window size 100 and empirical false alarm rates of 210% across classifiers. Detection holds across three ground-truth regimes: synthetic onset (86.6%), real temporal jailbreaks from public red-team databases (85%, 17/20), and a GCG adversarial demonstration showing that successful attacks produce insufficient distributional signal for detection in the target classifier while appearing anomalous to non-target classifiers. In a 4-classifier __ 3-shift conformal evaluation, weighted conformal prediction recovers up to 39 pp of lost coverage for DeBERTa (effective sample size 46/300) but eliminates data-driven reweighting for all other classifiers (ESS __ 300; residual recoveries __ +0.10 are a mechanical artifact of the test-point contribution): logistic density ratio estimation achieves perfect source/target separability in their embedding spaces, clipping all importance weights to the floor. This collapse is observed consistently across three of four classifiers and all three shift types; De= BERTa shows a gradient from effective correction (paraphrase, ESS 46) to neartotal collapse (adversarial suffix, ESS=206) depending on shift type. PCA to 32 dimensions before density ratio estimation breaks the collapse, recovering 33 pp for Llama Guard and 21 pp for ShieldGemma on temporal shift. Variance decomposition reveals that classifier ( __[2] = 0 _._ 243), shift type ( __[2] = 0 _._ 237), and their interaction ( __[2] = 0 _._ 185) all contribute substantially to detection latency variance (all _p <_ 0 _._ 001), indicating that neither factor alone determines detection difficulty and that per-classifier monitoring profiles are necessary.

Loading executive summary...

LINK COPIED TO CLIPBOARD