Turn-Based Structural Triggers: Stealthy Backdoors via Fine-Tuning Supply Chain Compromise
Research highlights a novel backdoor injection vector in multi-turn Large Language Models (LLMs) termed Turn-Based Structural Triggers (TST). By compromising the loss-computation component during the fine-tuning phase, adversaries can condition malicious model behavior on the dialogue turn position rather than specific text patterns. This attack leverages chat template structural cues to activate payloads at a predetermined target turn index. The vulnerability is highly effective, achieving a 98.10% success rate on target turns while maintaining 97.78% utility on clean tasks. Because the trigger is structural rather than lexical, current defense mechanisms like prompt filtering, sanitization, and paraphrasing are rendered obsolete, posing a severe threat to the AI training supply chain.