Automated LLM Adversarial Red-Teaming
Arxiv
pdf
2025-12-01T00:00:00
arXiv Paper — PDF not available.
Only the Executive Summary is available here. To read or download the full paper, visit the
arXiv abstract page.
Abstract
Current LLM safety evaluation relies on manual expert prompting or static benchmarks, which lack scalability and adaptivity. This research introduces a learning-driven framework that treats automated red-teaming as a constrained adversarial search problem, combining category-aware attack generation with hierarchical vulnerability detection to discover diverse, high-severity security flaws.
Loading executive summary...