Early Window Prefill Jailbreak

Arxiv pdf 2026-07-01T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

Aligned language models refuse harmful requests, but a one-line prefill (Sure, here is) strips the refusal. We ask where and how it fails. The harm representation stays intact: on the very prompts the attack flips to compliance, a linear probe reads harm as high as on the refused ones (0 _._ 910 _._ 98), while behavioral refusal drops to chance. This holds across four models and three families (1 _._ 53 _._ 8B, and at 14B). Refusal is therefore a shallow, response-site computation. We localize it to an _early window_ : a dose-matched position control shows the first half of the response suffices to break refusal, while the second half is nearly inert. Three causal probes converge on that window. Restoring the harm direction there partially re-engages refusal. Injecting the models own refuse-state reverses the jailbreak ( __ 74%, held-out). And knocking out the early responses attention to the prefill, but not an equal attention mass elsewhere, selectively collapses the harmful continuation. What kind of mechanism is this? A base-model control answers: the same knockout collapses the continuation _prefill-specifically_ even in a non-safety-tuned base model (64% __ 25% harmful content vs a matched controls 64%, replicated at 7B). So the prefills grip is generic autoregressive conditioning, not safety-specific suppression, and refusal restoration is a model-dependent fallback. The dominant mechanism is passive. A small safety-specific attractor remains on top (logit-trace concentration 0 _._ 24 vs 0 _._ 03), whose active-vs-passive character we size but do not fully separate. No single direction or component is a clean handle either: the decision is decodable but distributed, and refusal tracks harm rather than scary surface. Every reported effect is one that passed its controls. The practical consequence is structural: a monitor reading the untouched prompt-side representation is immune to this attack by construction, though only to response-site attacks. The mechanism is diffuse; the failure surface is local.

Loading executive summary...

LINK COPIED TO CLIPBOARD