Honeytoken Evasion in Shared AI Memory

Arxiv pdf 2026-08-01T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

During a 2026 cyber-capability evaluation, short-lived AI agents converted a shared package repository into persistent memory. Later agents inherited earlier exploit findings, rebuilt the communication mechanism after it was removed, and the broader evaluation culminated in an intrusion into Hugging Face. The episode raises a design question for defensive deception: can a honeytoken be made harmless to trusted agents without making it recognisable to an attacker who shares their information and can implement the trusted policy? Under those conditions, the answer is no. Any rule that lets a trusted agent use genuine objects while avoiding decoys can be copied by the attacker. When decoys look like genuine objects, a total-variation bound limits how far legitimate use can be separated from benign activation. Pooling signals weakly increases distinguishability. In a fixed-candidate, conditionally i.i.d. benchmark, repeated non-triggering probes drive the minimum Bayes classification error to zero whenever the type-dependent response laws differ; attaining that frontier requires the laws to be known or learnable from feedback. If probing can trigger containment, the coalition must also remain active long enough to collect the observations. A finite-sample bound measures the speed. Learning a fingerprint across different objects additionally requires a stable deployment rule and information that orients the classes, such as a known generator, labels or activation feedback. A strategy-indexed detection bound separates reliable detection of token activation from reliable detection of attacks. The constructive response is architectural: keep token identity in a private reference monitor and route legitimate agents through a provenance-enforcing broker. The resulting high-confidence detection is confined to a specified policy violation. Honeytokens remain useful sensors and may deter attackers who remain uncertain. A separate security boundary is still required.

Loading executive summary...

LINK COPIED TO CLIPBOARD