Fools Gold: Decoy Hardening for LLMs

Arxiv pdf 2026-08-01T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weights in minutes, and no release-time defense we are aware of prevents it durably. What cannot be prevented can be deceived. Our defense, decoy hardening (Fools Gold), concedes the refusal strip and poisons its payoff: once refusal is stripped, most answers to hazardous operational requests are confident, fluent decoys whose critical elements are falsified. The decoy behavior is trained inside a differentiable simulation of the attack, so it expresses in the attacked state, while a refusal pin and a benign leash hold clean-state behavior to the original.

Loading executive summary...

LINK COPIED TO CLIPBOARD