HPAA: Typographic LLM Moderation Evasion

Arxiv pdf 2025-12-08T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

Large language model (LLM)-powered content moderation systems primarily operate on tokenized text, largely ignoring the visual cues that humans use to interpret content. This research introduces Human-Perceptible Adversarial Attacks (HPAA), which utilize typographic manipulations—such as spacing, visual emphasis, and spatial arrangement—to embed harmful expressions into benign text. These manipulations allow content to remain easily recognizable to humans while becoming effectively invisible to automated moderation systems. Evaluations across ten widely deployed systems reveal that these attacks can achieve over 86% human recognition while maintaining detection rates below 1%.

Loading executive summary...

LINK COPIED TO CLIPBOARD