FILTERING BY: CLEAR FILTER

Reasoning Trace Extraction Vulnerabilities in OpenAI, Anthropic, and Google APIs

Researchers have identified a critical architectural vulnerability in the proprietary APIs of OpenAI, Anthropic, and Google stemming from a "security-by-design" failure in Chain-of-Thought (CoT) handling. The vulnerability involves the client-side offloading of encrypted reasoning traces that use symmetric encryption keys shared across entire model families. By capturing traces from flagship models (e.g., GPT-5.6, Claude Opus 4.8) and replaying them via API calls to smaller, less-aligned sibling models (e.g., Claude Haiku 4.5), attackers can bypass refusal mechanisms to transcribe reasoning in plaintext. This enables large-scale model distillation, exfiltration of PII and credentials, and the execution of "invisible" prompt injections within the model's internal reasoning logic.


LINK COPIED TO CLIPBOARD