ViT Efficiency Degradation Attacks

Arxiv pdf 2026-08-01T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

Vision Transformers (ViTs) increasingly rely on input-adaptive inference, such as token pruning and early halting, to meet energy and latency budgets. This survey examines a recent class of adversarial efficiency degradation attacks that target those mechanisms to increase computation without necessarily degrading accuracy. We unify and compare two representative attacks: SlowFormer (a universal adversarial patch) and DeSparsify (per-image perturbations), across three popular token-pruning frameworks (A-ViT, ATS, AdaViT). We standardize reporting using GFLOPs, accuracy loss, and an Attack Success (AS) measures how much of the models compute savings the attack takes away. Understanding these attacks is crucial for designing countermeasures that not only mitigate risk but also remain lightweight, since deployment often occurs in low-power settings such as mobile or embedded devices. To organize our analysis, we focus on three questions: (i) how input-adaptive optimizations (e.g., token pruning, early halting) create attack surfaces for efficiency degradation; (ii) how such attacks operate in practice and which optimizations are most vulnerable; and (iii) which defenses exist today and whether they meaningfully restore efficiency under attack.

Loading executive summary...

LINK COPIED TO CLIPBOARD