Adversarial Diffusion: Cross-Modality AI Attacks

Arxiv pdf 2026-06-01T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

Adversarial evaluation of AI systems has matured along four largely disconnected research tracks: difusion-based attacks on text and large language models (LLMs), difusion-based attacks on image classifers, jailbreak pipelines against vision-language models, and difusion-based input purfication defenses. Each track has developed its own vocabulary, threat models, and benchmarks, with denoising difusion models emerging as a shared generative mechanism whose recipes are now being actively ported between communities. This survey performs an information-fusion exercise at the meta-research level: we integrate these four tracks into a single conceptual framework with a unifed taxonomy, evaluation criteria, and research agenda, with primary focus on the LLM-side slice. We catalog ffty published papers across four scope areas (text/LLM, image classifer, vision-language model, defense), plus four difusion-LLM-as-victim entries and ten non-difusion baselines that any new difusion-based attack must be compared against. We propose a six-class taxonomy of difusion roles in adversarial pipelines, augmented by a threat-model axis that records attacker knowledge, query budget, and target accessibility, and we apply a fve-dimension evaluation framework (attack success rate, transferability, query budget, perplexity, defense-evasion) uniformly across modalities. The review adopts a dual attacker-defender perspective: alongside the attack catalog we cover four difusion-based defenses that constitute the natural evaluation backdrop for any new attack. We provide a critical analysis that identifes fve recurring weaknesses of the current LLM-side literature, and we close with a research agenda of open questions and concrete experimental designs. The companion catalog and the underlying spreadsheet are released alongside the paper. We are explicit that this is a narrative review with quality assessment, not a PRISMA-compliant systematic review, and we discuss the implications of this choice for replication.

Loading executive summary...

LINK COPIED TO CLIPBOARD