SIGIL: Proactive LLM Training Data Watermarking

Arxiv pdf 2025-12-07T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

As large language models (LLMs) are increasingly trained on scraped web corpora without authorisation, content owners require forensic methods to prove that their documents were included in a models training set. We propose SIGIL (Subtle Injection for Ground-truth Inference of LLM training data), a framework that embeds imperceptible canary sequences into protected text and code such that any LLM trained on those documents exhibits statistically detectable behavioural signatures when probed with targeted queries.

Loading executive summary...

LINK COPIED TO CLIPBOARD