CSO-LLM Backdoor Detection & Inversion

Arxiv pdf 2026-06-01T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

While post-training backdoor detection and trigger inversion schemes have been developed for AIs used e.g. for images, there is a paucity of such methods for LLMs. First, the LLM input space is discrete, with up to 150 _,_ 000 _[k] k_ -tuples to consider with _k_ the token-length of a putative trigger. Second, one must _blacklist_ tokens typical of the putative target response (class) of an attack, as such tokens may give _false_ detection signals. However, a comprehensive blacklist is not available, in general, for a given domain. We develop a highly effective detection and inversion framework for LLMs treated as classifiers. Central to our approach is _class subspace orthogonalization_ (CSO), a novel plug-and-play paradigm for backdoor detection that serves two fundamental roles when applied to LLMs: i) it enhances both sensitivity and specificity of a baseline detector; ii) it provides a form of _implicit_ blacklisting, as it penalizes against inclusion, in a candidate trigger, of tokens that induce signal perturbations in the direction of the putative target class of an attack. One version of our detector performs continuous optimization in token embedding space, while a companion trigger-inversion _and_ detection method performs greedy accretion in discrete token space. Our methods give both strong detection performance and accurate inversion of ground-truth triggers on several LLM classification domains, and for several different LLM architectures.

Loading executive summary...

LINK COPIED TO CLIPBOARD