CTI Environment Semantics Gap
Abstract
Structured Cyber Threat Intelligence (CTI) is increasingly used to support adversary emulation, detection evaluation, and cyber range design. However, each of those workflows still needs a target System Under Test (SUT) whose environment is not fully described by the public CTI. We measure how much of that environment is available in MITRE ATT&CK Structured Threat Information Expression (STIX) bundles. Across the ATT&CK Enterprise, Mobile, and Industrial Control Systems datasets, with the Common Attack Pattern Enumeration and Classification (CAPEC) and 5G Hierarchy of Threats (FiGHT) as contrast datasets, we report platform coverage, software specificity, vulnerability evidence, and deployment compatibility. Platform tags are common, but software references rarely pin versions or Common Platform Enumeration (CPE) identifiers. In Enterprise, 97.6% of software objects lack both, and campaign-level Common Vulnerabilities and Exposures (CVE) evidence remains sparse and fragmented. The measurements show that public ATT&CK-style structured CTI narrows candidate environments and supports lower-bound backend-family assignment, but structured fields alone do not justify deriving a replay-ready SUT. Profile confusion falls from 1.3% at one linked software item to 0% at two. These numbers locate the boundary between corpus-supported environment narrowing and the version, vulnerability, and deployment detail that must still come from outside the corpus. Beyond locating the boundary, we are the first to show it has an operational consequence: holding every corpus-supported element fixed and varying only the analyst-authored region yields multiple distinct, campaign-compatible SUTs, including an executable witness that runs the same real vulnerability. Structured CTI, therefore, constrains but does not uniquely determine the environment. Replay-ready emulation should accordingly declare which environment commitments the corpus supports and which remain analyst-authored.