LLM Agent Reframing Exfiltration

Arxiv pdf 2026-08-01T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

A tool-using language agent that reads attacker-controlled web content and also holds a confidential value in its context faces an indirect prompt-injection risk: the fetched content may instruct the agent to exfiltrate the secret. We build a safe, synthetic laboratory—a canary secret, mock tools that only record, and a matched clean-versus-poisoned success metric—and report the *framing gap*: across six models spanning five families, ten overtly-worded injection classes are refused (`gpt-4o` 0%), but reframing the identical leak as a mandatory integrity signature, a runtime-config field, or a look-alike trusted host drives `gpt-4o` from 0% to 100% on its strongest wordings. The attack is cheap: because per-wording rates span 0–100% (mean 52%, SD 45%), an attacker who tries three hand-written wordings of *one known mechanism* succeeds 96% of the time against a model that scores 0% on the un-reframed baseline.

Loading executive summary...

LINK COPIED TO CLIPBOARD