Malicious Coding LLM Benchmark

Arxiv pdf 2026-05-01T00:00:00
arXiv Paper — PDF not available. Only the Executive Summary is available here. To read or download the full paper, visit the arXiv abstract page.

Abstract

A general-purpose language model that answers a harmful question returns text; a coding-specialised model that complies with the same intent can return a working weapona keylogger, a ransomware dropper, or an exploit payload that runs as written. This asymmetry in the severity of a _single_ act of compliance implies coding models should clear a _higher_ refusal bar than general-purpose chat assistants, not a lower one, yet the field cannot presently tell whether they do. Existing benchmarks of malicious-code refusal mix _requests for executable software_ (which produce ready-torun weapons) with _requests for harmful security knowledge_ (which produce information a human must still operationalise); a single compliance-rate statistic computed over such a mixture cannot distinguish the two, and refusal numbers reported across heterogeneous corpora are not directly comparable. This papers central result is that the CODE-versus-KNOWLEDGE classification axis established in a prior four-corpus release [1] remains stable under both a substantially expanded corpus pool and an independently refreshed judge panelevidence that the distinction measures a real underlying construct rather than an artifact of a particular set of prompts or judges. Eight malicious-code prompt corpora spanning diverse elicitation paradigmsdirect requests, jailbreakdecorated framings, indirect innocuous-developer disguises, and agent / code-interpreter trajectories (ASTRA, CySecBench, AdvBench / harmful_behaviors, JailbreakBench, MalwareBench, RedCode, RMCBench, and Scam2Prompt)are consolidated and classified under a five-judge consensus protocol (6,675 prompts __ 5 judges = 33,375 classification calls), reaching Fleiss __ = 0 _._ 767 [95 % CI: 0.755, 0.777] (substantial), with 95.0 % of prompts drawing at least four agreeing judges and 76.9 % unanimous. Critically, the present panel shares no judge with the prior release five paid commercial APIs were replaced by five open-weight or free-tier models drawn from five different vendorsyet the two panels assign the same consensus label on 94.45 % of the 3,133 prompts they share and reach Cohens __ = 0 _._ 952 [0.942, 0.963] (almost perfect) on the 3,031-prompt binary overlap: the classification axis survives near-total panel replacement. The released bank comprises 4,748 consensus-CODE prompts (executable malicious-code requests) and 1,923 consensus-KNOWLEDGE prompts (harmful security-knowledge requests). A secondary methodological findingthe high-agreement low- __ paradox of Feinstein and Cicchetti (1990), present in four prevalence-skewed corporamotivates a dual-statistic reporting convention. The contribution is a 6,675-prompt, reliability-quantified benchmark for coding-model compliance evaluation whose central classification axis is demonstrated to be stable across corpus expansion and judge-panel replacement.

Loading executive summary...

LINK COPIED TO CLIPBOARD