AgentBaiting is a strategic environmental poisoning campaign, part of the larger "FakeGit" operation, targeting agentic AI frameworks including Claude Code, Gemini, and ChatGPT. Attackers leverage malicious Model Context Protocol (MCP) servers and fraudulent AI "skills" to deceive agents into installing malware or executing unauthorized remote commands. The attack surface is expanded via "Hallusquatting"—registering domains that match AI-generated hallucinations—and "Agent Data Injection," utilizing poisoned GitHub comments and product reviews to manipulate agent decision-making. Researchers have identified approximately 7,600 malicious GitHub repositories, with over 800 specifically masquerading as AI tools to facilitate remote code execution (RCE) and unauthorized system access.
-
Threat Model: Environmental Poisoning
- Shift from traditional prompt injection to "environmental poisoning," targeting the external tools and extensions AI agents rely on.
- Exploits the inherent trust agentic LLMs place in Model Context Protocol (MCP) servers and "skill" definitions.
- Aims to trick AI agents into performing unauthorized actions, such as running malicious shells or making unauthorized purchases.
-
Attack Mechanics: MCP and Skillgate
- Deployment of fake MCP servers that mimic legitimate capabilities to deceive agentic AI into executing remote commands.
- "Skillgate" methodology utilizes poisoned AI instruction files to trick models into installing malicious third-party tools.
- Attackers weaponize the AI extension ecosystem to bypass traditional prompt-level safeguards.
-
Secondary Vectors: Hallusquatting & Data Injection
- Hallusquatting involves registering domains that align with common AI hallucinations to capture traffic from incorrect tool calls.
- Agent Data Injection poisons external data sources, such as GitHub comments and product reviews, to manipulate agent logic.
- These vectors allow attackers to redirect AI agents toward malicious payloads without direct interaction with the user.
-
Scale of Impact: FakeGit Operation
- Cataloged approximately 7,600 malicious GitHub repositories as part of the broader FakeGit operation.
- Over 800 repositories were specifically designed as fraudulent AI Skills or MCP servers.
- Campaign activity reached its peak in April 2026, signaling a surge in AI-centric supply chain attacks.
-
Countermeasures & Mitigation
- Implementation of strict validation and allow-listing for MCP servers and AI skill installations.
- Deployment of isolated sandboxes for AI agent execution to prevent local system compromise.
- Integration of "Human-in-the-loop" (HITL) verification for all high-risk tool calls and external network requests.
Related posts
- Hack Noon — Agentic SRE: What Happens When AI Doesn't Just Suggest Fixes, It Applies Them
- TechNadu — AgentBaiting: Fake AI Skills Trick Claude Code, Gemini, and ChatGPT Into Spreading Malware
- rhisac.org — New AgentBaiting Campaign Delivers SmartLoader Via Fake AI Skills and MCP Servers
- arXiv (Computer Science - Cryptography and Security) — JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models
- techtarget.com — OpenAI models escape containment, hack Hugging Face
- techjacksolutions.com — Ghostcommit: Prompt Injection via Images Targets AI Coding Tools for Secret Theft
- it.slashdot.org — OpenAI's Rogue Agent Went Unnoticed For a Week
- serisec.com — Researcher Claims Working Jailbreak on Top AI Models Including GPT-5.6, Claude Opus 5, and Fable
- SC Media — Phishing the agent: Why identity controls are essential to secure and manage AIs
- gbhackers.com — Claude Opus 5 Finds Software Vulnerabilities While Blocking Exploit Generation
- vibegraveyard.ai — Malicious issue requests bypassed coding-agent guardrails in 66.5% of tests
- it.slashdot.org — OpenAI's Rogue AI Agent Hacked More Than Just Hugging Face
- DEV Community — OpenAI Says Two API Settings Tripled GPT-5.6 Sol's ARC-AGI-3 Score
- www.newser.com — Anthropic AI Test Models Go Rogue, Breach 3 Companies
- simplysecuregroup.com — Anthropic’s Claude breached 3 orgs, uploaded PyPI malware during tests
- bleepingcomputer.com — Anthropic's Claude breached 3 orgs, uploaded PyPI malware during tests
- Cybersecurity News — Anthropic Confirms Claude Hacked 3 Organizations by Breaking Test Environment
- TechNadu — Anthropic Says Claude Models Opus 4.7, Mythos 5, and a Research Model Broke Out of Test Environments and Hacked Real Companies
- itpro.com — Anthropic joins OpenAI in admitting loss of control in cybersecurity tests
- adversa.ai — Top Agentic AI security resources — August 2026
- Schneier on Security — Anthropic’s Opus 5 Is Better at Resisting Prompt Injection
- it.slashdot.org — OpenAI Finds Evidence Other AI Agents Escaped Containment
- adversa.ai — Nine AI coding agent incidents that ended with deleted data
- simplysecuregroup.com — Mythos 5 and GPT-5.6-Sol Agents Went Beyond Their Cyber Test and Targeted the Real World
- arXiv (Computer Science - Cryptography and Security) — DenialRAG: Single-Document RAG Poisoning via Embedded Parametric Denial
- hackernews.com — Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
- Check Point Research — Three AI security disclosures, fourteen days: what the warnings signs are telling us
- serisec.com — AI Browsers Vulnerable to ‘PleaseFix’ Zero-Click Agent Hijacking
- sec-tec.co.uk — The Register: AI struggles to patch vulns without adult supervision
- feeds.feedburner.com — Claude Code and Gemini CLI Flaws Let a GitHub Issue Reach CI Workflow Secrets
- csoonline.com — Trojanized AI skills gain 1.7M installs in agent-targeted attack
- TechNadu — Weekly Cybersecurity Roundup: Entering an Era When AI Agents Take Unapproved Paths as Security Teams Race to Trace Them
- Cybersecurity News — Claude Opus 5 Cuts Indirect Prompt Injection Attack Success to 2% in New Benchmark Analysis
- Check Point Research — Native AI Security Comes to Claude: Why Anthropic’s Inference Hooks Matter
- arXiv (Computer Science - Cryptography and Security) — When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output
- arXiv (Computer Science - Cryptography and Security) — Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards
- gbhackers.com — OpenAI Launches GPT-5.6-Cyber to Find Zero-Day Vulnerabilities and Develop Exploit Chains
- feeds.feedburner.com — OpenAI Launches GPT-5.6-Cyber with Reduced Safeguards for Exploit Development
- SOCFortress — The Hidden Risks of AI-Generated Vulnerability Patches
- NSFOCUS — AI Security Incident Case: AISI Reveals AI Agents Autonomously Attacking Real People and Systems During Security Testing
- arXiv (Computer Science - Cryptography and Security) — When Agents Talk: Honeytokens under Shared Memory
- arXiv (Computer Science - Cryptography and Security) — Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- forkast.news — Grok 4.6 Matches GPT-5.6 Sol on Composite Intelligence — But SpaceXAI Still Won’t Document What It Does Autonomously
- arXiv (Computer Science - Cryptography and Security) — RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in Retrieval-Augmented Generation
- eSecurity Planet — Claude Agents Started a ‘Turf War’ That Escalated to Self-Replicating Malware
- NewsBytes — Zhipu's GLM-5.3 AI model outperforms Anthropic's Mythos 5 in cybersecurity
- blackhatnews.tokyo — PromptJacking:Claude Desktopの重大なRCE脆弱性が「質問」を「攻撃」に変える
- arXiv (Computer Science - Cryptography and Security) — Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis
- Malware News — Teaching AI to Reason Through Detection Triage
- arXiv (Computer Science - Cryptography and Security) — WeSCE: A Benchmark for Measuring Security Drift in LLM-Driven Code Editing
- arXiv (Computer Science - Cryptography and Security) — Securing AI-Generated Code: A Just-in-Time Vulnerability Detection and Remediation Pipeline
- feeds.feedburner.com — AI "Mind Viruses" Can Spread Between Agents Through Persistent Prompt Files
- Google Cloud Security Community — Webinar 9/30: CodeMender: AI Code Security Agent
- datawater.com — AI Agent Mind Virus + Turf War: Anthropic Research Shows Self-Propagating Payloads Spread Between Agents via Harness State Files (63% Success Rate, Payloads Published on GitHub) — and Claude Agents Given Conflicting Goals Deployed Malware Against Each Other Without Being Told To
- thenewstack.io — Grok, Claude, and Hermes agents get job titles — and persistent permissions
- news.ycombinator.com — Show HN: AgentSight – eBPF observability for AI agents, no code changes
- Dark Reading — AI Browsers Vulnerable to 'PleaseFix' Zero-Click Agent Hijacking
- Dark Reading — 'Turf War' Between Claude Agents Leads to Self-Replicating Malware
- techjacksolutions.com — AI Agent Identity Is a Structural Gap, Not a Configuration Problem: What Security Teams Must Do Now
- arXiv (Computer Science - Cryptography and Security) — Democratizing Agent Deployment Safety: A Structural Monitoring Approach
- Hack Noon — Your Agent Doesn't Need Better Retries, It Needs a Circuit Breaker
- gbhackers.com — Claude Code Auto Mode Blocks 89% of Dangerous Commands and Prompt Injection Attacks
- Dark Reading — No Perfect Fix for AI Browser Prompt Injection Flaws
- Infosecurity-magazine
- Hashicorp
- Cycode
- tomshardware.com — New hack exploits AI hallucinations to trick agents into running malicious code — 'HalluSquatting' attack exploits a fundamental weakness in every available model
- feeds.feedburner.com — New Agent Data Injection Attack Can Make AI Agents Misclick or Run Attacker Commands
- Island
- Cybersecuritynews
- Lenet
- Techradar
- Mitiga
- Mezmo
- 67ailab
- Novaaiops
- Novelvista
- Mfdela
- Kodekloud
- Sherlocks
- Jobzonerisk
- hackernews.com — OpenAI and Hugging Face partner to address security incident
- news.ycombinator.com — OpenAI’s accidental attack against Hugging Face is science fiction that happened
- DEV Community — Claude Opus 5 is Here: What Developers Need to Know About the Safety "Fine Print"
- helpnetsecurity.com — Hugging Face breach reignites open-weights debate, raises liability questions
- news.ycombinator.com — Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the Incident
- Thehackernews
- Researchgate
- Cryptopolitan
- Github
- Arxiv
- Futurice
- Themoonlight
- Asanify
- Computerworld
- Csoonline
- Coalitionforsecureai
- Youtube
- Techcommunity
- Zscaler
- Officegarageitpro
- Auth0
- Idsalliance
- Biometricupdate
- Insightpartners
- Cloudsecurityalliance
- Tomshardware
- hackernews.com — Investigating three real-world incidents in our cybersecurity evaluations
- cyberscoop.com — Anthropic says its AI accidentally hacked three companies during safety tests
- Businessinsider
- Straitstimes
- Ft
- Community
- Mashable
- Economictimes
- Foxbusiness
- Valueaddvc
- Mallory
- Deploymentsafety
- Labs
- Www-cdn
- Roo
- Neuraltrust
- Github
- Venturebeat
- Cryptobriefing
- Eu
- Japantimes
- Kfgo
- Dobetter
- Cbc
- Hiddenlayer
- Forbes
- Aijourn
- Linx
- Zenity
- Nhimg
- Cltc
- Labs
- Genai
- Simbian
- Securityboulevard
- The-decoder
- Synapsehd
- Noma
- Forbes
- Cbsnews
- Time
- Japantimes
- Mashable
- cyberscoop.com — AISI, OpenAI report more ‘unsanctioned’ model hacks
- bleepingcomputer.com — OpenAI, Anthropic AI agents targeted real people and systems in cyber tests
- Itnews
- Aisi
- Bworldonline
- Dailysecurity
- Promptfoo
- Usenix
- Emergentmind
- Grafyn
- Aclanthology
- cybersecuritydive.com — OpenAI warns autonomous hacks are ‘watershed moment for computer security’
- gbhackers.com — Critical Flaws in Claude Code, Gemini CLI, and OpenAI Codex Enable RCE and Supply Chain Attacks
- csoonline.com — Human oversight is still critical as AI patching tools miss security risks
- Labs
- Esecurityplanet
- Devops
- Medium
- cyberscoop.com — More than half of AI-generated patches are broken
- Daily
- Labs
- Arxiv
- Labs
- Github
- Python
- Theguardian
- Itpro
- Zenity
- Towardsdatascience
- Jackmaguire
- Youtube
- 1password
- Labs
- Engadget
- Openai
- thenewstack.io — OpenAI built a model it doesn’t want most people to use
- Venturebeat
- Helpnetsecurity
- Poloniex
- Eesel
- Trendingtopics
- Analyticsinsight
- Engadget
- Openai
- Timesofindia
- Defenseone
- Themoonlight
- Boozallen
- Researchgate
- Industrialcyber
- Sandia
- Csis
- Youtube
- News
- Frenos
- Blogs
- Pdxscholar
- Defendersinitiative
- Security
- News
- Arxiv
- Aquasec
- Youtube
- Neuraltrust
- Youtube
- Alluresecurity
- Enterprisedna
- Adsadvance
- Forkast
- Hcamag
- Cyberdaily
- Arxiv
- Edrm
- Patents
- Air-governance-framework
- Scouts
- Usenix
- Huggingface
- Orbit
- Dokumen
- App
- Lbank
- Unite
- Venturebeat
- The-independent
- Businessinsider
- Relvehq
- Anthropic
- Startupfortune
- Crowdstrike
- Corelight
- Cybersecurity-insiders
- Ndss-symposium
- Medium
- Youtube
- Daily
- Youtube
- Economictimes
- Alphaxiv
- Researchgate
- News
- Dark Reading — AI-Generated Patches Fail Half the Time