LLM Backdoor Neuron Pruning

Arxiv pdf 2026-07-01T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

six open-source LLMs and two benchmark datasets demonstrate that DeCNIP achieves more than 95% relative reduction in Attack Success Rate (ASR) , outperforming seven state-of-the-art defenses with only 0.1% of the neurons intervened . Moreover, it maintains an average of 97% of the models foundational performance on normal benchmarks, illustrating its efficacy, robustness, and scalability in securing large-scale generative models.

Loading executive summary...

LINK COPIED TO CLIPBOARD