Verifiable AI Agent Guardrails

Arxiv pdf 2026-03-01T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

As AI agents become widely deployed as online services, users often rely on an agent developers claim about how safety is enforced, which introduces a threat where safety measures are falsely advertised. To address the threat, we propose proof-of-guardrail, a system that enables developers to provide cryptographic proof that a response is generated after a specific open-source guardrail. To generate proof, the developer runs the agent and guardrail inside a Trusted Execution Environment (TEE), which produces a TEE-signed attestation of guardrail code execution verifiable by any user offline. We implement proof-of-guardrail for OpenClaw agents and evaluate latency overhead and deployment cost. Proof-of-guardrail ensures integrity of guardrail execution while keeping the developers agent private, but we also highlight a risk of deception about safety, for example, when malicious developers actively jailbreak the guardrail.

Loading executive summary...

LINK COPIED TO CLIPBOARD