LLM Privacy Auditing via Synthetic Canaries
Abstract
Parameter-efficient fine-tuning of large language models (LLMs) can exhibit problematic memorization of individual training examples. To quantify this risk, this research proposes an improved Empirical Privacy Auditing (EPA) method using synthetic canaries generated via high-temperature sampling. These canaries act as high-influence outliers that enhance the identifiability of membership inference and reconstruction attacks without requiring access to the original private data. Furthermore, the paper introduces a novel model-based audit for synthetic data, which uses an auxiliary model to detect subtle, diffuse leakage signals in generated datasets, demonstrating that releasing synthetic data does not inherently guarantee privacy.