Automated LLM Adversarial Red-Teaming

Arxiv pdf 2025-12-01T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

Current LLM safety evaluation relies on manual expert prompting or static benchmarks, which lack scalability and adaptivity. This research introduces a learning-driven framework that treats automated red-teaming as a constrained adversarial search problem, combining category-aware attack generation with hierarchical vulnerability detection to discover diverse, high-severity security flaws.

Loading executive summary...

LINK COPIED TO CLIPBOARD