LLM Security Agent Trajectory Risk Certification

Arxiv pdf 2026-08-01T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

Autonomous security agents increasingly operate as staged decision pipelines, e.g. classifying network traffic and then attributing detected attacks to a specific technique. Split conformal prediction gives each stage a finite-sample coverage guarantee, but deployment requires a trajectory-level guarantee across the whole chain, and the two do not compose for free—especially for stages that are already independently trained and calibrated and cannot be jointly recalibrated. Bonferroni allocation gives a valid distribution-free trajectory bound but is conservative when stage errors are correlated. We show a natural pairwise-correlation extension of this bound to three or more stages is invalid—a lower, not upper, bound—and give a provably valid spanning-tree alternative. We then separate two questions routinely conflated in practice: whether stages are dependent at all, and whether a finite audit sample is large enough to certify that dependence, giving matching upper and information-theoretic lower sample-complexity bounds for both. We further prove that a common design pattern—using a coarse category to select a fine-grained label space—mechanically manufactures near-perfect measured stage correlation with no learned dependence behind it.

Loading executive summary...

LINK COPIED TO CLIPBOARD