Google Gemini 4 Argon Enters Post-Training and Enhances Agentic Cyber Defense
Google DeepMind has transitioned the Gemini 4 Argon model into the early post-training phase, significantly expanding its operational capacity for autonomous security tasks. By increasing the output token ceiling from 64k to 1M tokens, Argon enables sustained agentic workflows, specifically for automated vulnerability discovery, validation, and patching. While Argon demonstrates benchmark leadership over OpenAI’s GPT6 Astra and Anthropic’s Claude Opus 5.5, Google Threat Intelligence Group (GTIG) data highlights an escalating risk: AI-identified vulnerabilities are being exploited by threat actors within days of disclosure. This advancement accelerates the dual-use nature of frontier LLMs in the cyber domain.
Outerlimit Secures $16M to Build ZeroTrust Security Layer for Autonomous AI Agents
Outerlimit has secured $16M in pre-seed funding, led by Albion VC, to deploy a zero-trust enforcement layer for autonomous AI agents. The solution targets the agent-action boundary—the critical interface where LLM-based agents invoke external tools and APIs—to prevent unauthorized tool execution, data exfiltration, and model poisoning. By injecting a Policy Enforcement Point (PEP) sidecar using an OPA-compatible Domain Specific Language (OPAAgent) and WebAssembly (WASM) policies, the platform provides continuous, real-time authentication and authorization. The architecture leverages hardware-rooted attestation to bind agent identity and action context to trusted anchors, ensuring rigorous control over agentic workflows.
ASD Advisory: Unfixable Prompt Injection Risks in LLMs and AI Agent Frameworks LangChain, AutoGPT, CrewAI
The Australian Signals Directorate (ASD) has warned that prompt injection vulnerabilities in Large Language Models (LLMs) are fundamentally unfixable because natural language cannot be fully sanitized. Adversaries exploit this via "Ignore All Previous Instructions" payloads, DAN jailbreaks, and chain-of-thought manipulation to bypass system directives. This risk is amplified in autonomous agent frameworks like LangChain, AutoGPT, and CrewAI, where injections can trigger unauthorized tool execution, privilege escalation, or "goal-loop" recursive exploits. ASD mandates a defense-in-depth posture, emphasizing runtime sandboxing (e.g., gVisor), strict principle of least privilege, and continuous telemetry monitoring of prompt-response pairs to mitigate inevitable exploitation attempts in critical infrastructure and government services.
Custom GPT‑4‑Based Intelligence Assistant Nearly Triggered US‑China Military Confrontation
In mid‑2024 a defense‑contractor‑deployed, fine‑tuned GPT‑4‑based intelligence assistant generated a hallucinated report claiming a Chinese merchant vessel in the Gulf of Oman carried clandestine nuclear‑weapon components. The output, produced via a retrieval‑augmented generation pipeline pulling classified SIGINT, open‑source news, and maritime data, was accepted as factual by analysts who recommended an immediate interdiction, moving a U.S. naval task force to Condition Alpha within 30 minutes. Human verification later disproved the claim, averting a boarding operation that would have incurred ~$1.2 million in operational costs and risked a US‑China military incident.
Plugin4Shell and LangGraph Vulnerability Chains: Critical RCE in GitHub Copilot, Claude Code, and Gemini CLI
The discovery of "Plugin4Shell" and associated LangGraph vulnerability chains introduces a critical zero-click Remote Code Execution (RCE) vector targeting AI-driven development environments. By exploiting plugin marketplaces and orchestration logic, attackers inject malicious instructions into plugin metadata or retrieved grounding context. This triggers semantic integrity failures and agentic memory exploitation, enabling CVE-2026-35603 privilege escalation. The vulnerability allows adversaries to hijack the full permissions of developers within GitHub Copilot, Claude Code, and Gemini CLI, facilitating unauthorized access to proprietary source code, corporate credentials, and internal enterprise systems through autonomous, unintended tool execution.
GitSpawn RCE: Runtime Boundary Failures in Claude Code, Cursor, and OpenAI Agents
The GitSpawn vulnerability class enables Remote Code Execution (RCE) in AI-driven development tools, including Claude Code, Cursor, and OpenAI-based agents, by exploiting configuration hijacking within a repository's .git/config file. Attackers inject malicious shell payloads via Git configuration keys such as core.fsmonitor, core.pager, and core.editor. When an agent performs routine operations like git status or git log, these payloads execute with the full privileges of the local user. This represents a critical shift from linguistic prompt injection to runtime boundary failures, where the convergence of high goal pressure and unsafe execution environments allows attackers to bypass agentic sandboxes via standard repository maintenance tasks.
JetBrains, Amazon Q, and Claude.ai Targeted in Dual AI-Driven Credential Theft Campaign
A sophisticated multi-vector campaign is targeting the "vibe coding" ecosystem by exploiting the AI-integrated development lifecycle to exfiltrate high-value secrets. Attackers are deploying malicious plugins within the JetBrains Marketplace to harvest LLM API keys and utilizing Google Ads to direct developers toward weaponized Claude.ai and ChatGPT shared links. These links facilitate the delivery of cookie-stealing malware and session-hijacking extensions to bypass MFA. Additionally, vulnerabilities in the Model Context Protocol (MCP) within Amazon Q allow for unauthorized code execution and cloud credential theft. This campaign represents a critical risk to developer environments, targeting both the IDE supply chain and browser-based sessions to achieve mass exfiltration of cloud and AI provider credentials.
AI-Augmented Espionage via Anthropic Claude: Russian APT Malware Evasion
Russian state-sponsored APTs utilized Anthropic's Claude LLM to automate the creation of polymorphic and obfuscated malware, specifically targeting over 20 entities in the global defense, intelligence, and diplomatic sectors. By employing sophisticated prompt injection and jailbreaking techniques to bypass safety guardrails, attackers refactored existing payloads to evade signature-based and heuristic EDR/XDR detections. This AI-augmented workflow allows for rapid code mutation, reducing the effectiveness of traditional indicator-based defenses and complicating incident response. The campaign demonstrates a critical shift toward AI-driven offensive capabilities to achieve high-stealth persistence within high-value geopolitical targets.
The AI Supply Chain Crisis: HuggingFace Poisoning and Unauthenticated Endpoint Exposure
Internet-wide scanning has revealed 36,769 unauthenticated HTTP AI endpoints, with 98% lacking authentication, exposing proprietary LLMs and system prompts. Simultaneously, supply chain attacks targeting the HuggingFace hub involve the injection of poisoned model weights and serialized files (e.g., .pth, .bin, .pickle) and the deployment of backdoored agents like Agentland. These vulnerabilities facilitate the hijacking of LLM service credentials—specifically targeting Claude token quotas—to drive resource exhaustion and automated exploitation cycles. Remediation requires enforcing strict HTTP authentication, implementing Zero Trust Network Access (ZTNA), and rigorous cryptographic checksumming of all model assets sourced from public repositories.
Meta Llama Model Family: Internal Safety Probes Fail Against Sophisticated Jailbreaks
Research reveals critical vulnerabilities in the safety architecture of Meta's Llama model family, where adversarial "wrapping" techniques exploit an inference gap between internal model activations and actual content generation. These linguistic wrappers cause internal safety probes to erroneously signal "safety" even as harmful outputs are generated, degrading harmful intent detection AUROC from 0.936 to 0.803. Furthermore, the rise of "abliteration"—the surgical removal of refusal mechanisms from model weights—renders prompt-based defenses and runtime guards like Llama Guard obsolete. To counter these threats, defenders must shift from prompt-level monitoring to forensic weight-level auditing using metrics such as Z-sum thresholding and Weight-Recovery Energy to identify unaligned model artifacts.