CLIPGuard: Black-Box Defense Against Embedding-Space Backdoors
Abstract
Contrastive LanguageImage Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities. However, recent studies reveal a critical vulnerability: _embedding-space backdoor attacks_ . By poisoning only a tiny fraction of imagetext pairs, adversaries can implant stealthy triggers that induce targeted shifts in CLIPs joint embedding space. Unlike conventional backdoors that manipulate classifier logits, these attacks corrupt representations directly, making them highly effective under extremely low poisoning ratios and difficult to detect. Existing defenses require access to model parameters, gradients, logits, or clean validation dataassumptions that rarely hold in realistic black-box deployments. Moreover, current black-box methods struggle to accurately localize small or out-of-distribution triggers. We propose CLIPGuard, a lightweight and fully black-box defense specifically designed to mitigate embeddingspace backdoors in CLIP encoders. CLIPGuard identifies malicious regions by measuring _segment-wise embedding perturbations_ and selectively purifies only suspicious segments via semantic inpainting, preserving benign visual content and alignment quality. Extensive experiments on STL-10, ImageNet, and diverse trigger familiesincluding BadCLIP, BadNets, blended, patch-based, and typographic attacksdemonstrate that CLIPGuard reduces attack success rates to as low as _1.05%_ while maintaining clean accuracy up to _86.34%_ , consistently outperforming existing black-box defenses, including CleanCLIP and CleanerCLIP. Our code is available `https://github.com/wsu-cyber-security-lab-ai/CLIPGuard.git`