FILTERING BY: CLEAR FILTER

HARD Framework: Towards Self-Evolving Defense for LLM Agents

This research introduces the HARD (Harness-based Autonomous Runtime Defense Evolution) framework to mitigate the vulnerabilities inherent in autonomous LLM agents. Current defensive postures rely on manual, "handcrafted" rules that fail to intercept multi-step execution exploits and complex agentic workflows. HARD moves security into the runtime execution loop via a harness-level formulation, integrating defense mechanisms directly into the agent's operation. By utilizing failure trace analysis engines, the system automatically identifies defense gaps and evolves security artifacts, such as dynamic policies and filters. This approach aims to reduce the Attack Success Rate (ASR) while maintaining utility through a continuous, self-improving cycle of autonomous intervention.


LINK COPIED TO CLIPBOARD