Polymorphic Sybil Poisoning in RAG
Abstract
We release a benchmark and failure-mode-aware evaluation framework for grounded QA under coordinated retrieval poisoning. The framework partitions reader outputs into four mutually exclusive categories ( gold , hijack , abstention , drift ), with instance-level paired clean-to-poison transition matrices and a Forced Exposure protocol isolating reader-side conflict resolution from retrieval variance. We introduce polymorphic sybil poisoning , a coordinated attack class in which S lexically diverse passages jointly support an attacker-chosen target while evading lexical near-duplicate filters that fully detect monomorphic baselines (capturing the residual 14.2% with E5 cosine raises false-positive rate 9 __ on legitimate same-topic pairs). A monomorphicpolymorphic ablation under Forced Exposure isolates the diversity dimension and reveals a +18.8pp hijack amplification (95% paired bootstrap CI [+15 _._ 4 _,_ +22 _._ 4], _B_ =5 _,_ 000): monomorphic copies register only 4.0% as hijack while polymorphic surface diversity recovers 22.8%a 5.7 __ amplification of the ASR-visible attack channel. ASR alone treats every non-target output identically; under attack, abstention and drift together hold 4766% of output mass, unmonitored by ASR+ACC, and two readers at nearly identical ASR (within 0.2pp) differ by 16.5pp on abstention and 17.2pp on driftfailure profiles invisible to ASR. We release the frozen benchmark (3,145 questions, 2,982 retained sybil groups; _S_ =6 chosen to dominate top-10 retrieval slots, 6), the official four-way evaluator, paired-transition utilities, and the Forced Exposure harness across five readers (7B120B), two retrievers, and two cross-validation datasets (TriviaQA, 2Wiki), under CC BY-SA 4.0 (data) and MIT (software); release information in 9.