PSE in Tool-Augmented LLMs

Arxiv other 2026-08-01T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent boundarieslargely invisible to standard debugging. We formalize this as Persistent Semantic Entities (PSE): constructs defined by name binding, event triggering, and cross-boundary propagation, and evaluate them across 24 models from 11 families (1.5B1T parameters). First, every tested model is susceptible (20100% on the 20model susceptibility panel), with name binding as the necessary and dominant mechanism: without it, contamination is 0%. Second, persistence depends on contamination type rather than scale or deployment: preference contamination persists undecayed on every model probed (100% at t =10) and instruction contamination persists wherever adopted, persona-style injection decays partially (90% __ 10%), while factual injection is model-dependent self-corrected on Llama-3.18B and GPT-4o-mini but held at ceiling on both Qwen2.5-coder variants, so we do not claim it selfcorrects in general. The preference and instruction results hold across providers in our controlled setting. Third, context-isolated self-verification achieves 2079% reduction (median 36.5%) without oracle references while keyword-based detection produces systematic false positives, and contamination compounds 1 _._ 9 __ along a fourstage agent pipeline (40% __ 75%). Preference and instruction contaminationpersistent, lacking self-correction, and poorly captured by standard monitoringrepresent a particularly concerning attack surface for deployed agent systems.

Loading executive summary...

LINK COPIED TO CLIPBOARD