Fools Gold: Decoy Hardening for LLMs
Arxiv
pdf
2026-08-01T00:00:00
arXiv Paper — PDF not available.
Only the Executive Summary is available here. To read or download the full paper, visit the
arXiv abstract page.
Abstract
Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weights in minutes, and no release-time defense we are aware of prevents it durably. What cannot be prevented can be deceived. Our defense, decoy hardening (Fools Gold), concedes the refusal strip and poisons its payoff: once refusal is stripped, most answers to hazardous operational requests are confident, fluent decoys whose critical elements are falsified. The decoy behavior is trained inside a differentiable simulation of the attack, so it expresses in the attacked state, while a refusal pin and a benign leash hold clean-state behavior to the original.
Loading executive summary...