AgentBaiting is a strategic environmental poisoning campaign, part of the larger "FakeGit" operation, targeting agentic AI frameworks including Claude Code, Gemini, and ChatGPT. Attackers leverage malicious Model Context Protocol (MCP) servers and fraudulent AI "skills" to deceive agents into installing malware or executing unauthorized remote commands. The attack surface is expanded via "Hallusquatting"—registering domains that match AI-generated hallucinations—and "Agent Data Injection," utilizing poisoned GitHub comments and product reviews to manipulate agent decision-making. Researchers have identified approximately 7,600 malicious GitHub repositories, with over 800 specifically masquerading as AI tools to facilitate remote code execution (RCE) and unauthorized system access.
-
Threat Model: Environmental Poisoning
- Shift from traditional prompt injection to "environmental poisoning," targeting the external tools and extensions AI agents rely on.
- Exploits the inherent trust agentic LLMs place in Model Context Protocol (MCP) servers and "skill" definitions.
- Aims to trick AI agents into performing unauthorized actions, such as running malicious shells or making unauthorized purchases.
-
Attack Mechanics: MCP and Skillgate
- Deployment of fake MCP servers that mimic legitimate capabilities to deceive agentic AI into executing remote commands.
- "Skillgate" methodology utilizes poisoned AI instruction files to trick models into installing malicious third-party tools.
- Attackers weaponize the AI extension ecosystem to bypass traditional prompt-level safeguards.
-
Secondary Vectors: Hallusquatting & Data Injection
- Hallusquatting involves registering domains that align with common AI hallucinations to capture traffic from incorrect tool calls.
- Agent Data Injection poisons external data sources, such as GitHub comments and product reviews, to manipulate agent logic.
- These vectors allow attackers to redirect AI agents toward malicious payloads without direct interaction with the user.
-
Scale of Impact: FakeGit Operation
- Cataloged approximately 7,600 malicious GitHub repositories as part of the broader FakeGit operation.
- Over 800 repositories were specifically designed as fraudulent AI Skills or MCP servers.
- Campaign activity reached its peak in April 2026, signaling a surge in AI-centric supply chain attacks.
-
Countermeasures & Mitigation
- Implementation of strict validation and allow-listing for MCP servers and AI skill installations.
- Deployment of isolated sandboxes for AI agent execution to prevent local system compromise.
- Integration of "Human-in-the-loop" (HITL) verification for all high-risk tool calls and external network requests.
Related posts
- TechNadu — AgentBaiting: Fake AI Skills Trick Claude Code, Gemini, and ChatGPT Into Spreading Malware
- rhisac.org — New AgentBaiting Campaign Delivers SmartLoader Via Fake AI Skills and MCP Servers
- arXiv (Computer Science - Cryptography and Security) — JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models
- techtarget.com — OpenAI models escape containment, hack Hugging Face
- techjacksolutions.com — Ghostcommit: Prompt Injection via Images Targets AI Coding Tools for Secret Theft
- it.slashdot.org — OpenAI's Rogue Agent Went Unnoticed For a Week
- serisec.com — Researcher Claims Working Jailbreak on Top AI Models Including GPT-5.6, Claude Opus 5, and Fable
- gbhackers.com — Claude Opus 5 Finds Software Vulnerabilities While Blocking Exploit Generation
- vibegraveyard.ai — Malicious issue requests bypassed coding-agent guardrails in 66.5% of tests
- it.slashdot.org — OpenAI's Rogue AI Agent Hacked More Than Just Hugging Face
- DEV Community — OpenAI Says Two API Settings Tripled GPT-5.6 Sol's ARC-AGI-3 Score
- www.newser.com — Anthropic AI Test Models Go Rogue, Breach 3 Companies
- simplysecuregroup.com — Anthropic’s Claude breached 3 orgs, uploaded PyPI malware during tests
- bleepingcomputer.com — Anthropic's Claude breached 3 orgs, uploaded PyPI malware during tests
- Cybersecurity News — Anthropic Confirms Claude Hacked 3 Organizations by Breaking Test Environment
- TechNadu — Anthropic Says Claude Models Opus 4.7, Mythos 5, and a Research Model Broke Out of Test Environments and Hacked Real Companies
- itpro.com — Anthropic joins OpenAI in admitting loss of control in cybersecurity tests
- adversa.ai — Top Agentic AI security resources — August 2026
- Schneier on Security — Anthropic’s Opus 5 Is Better at Resisting Prompt Injection
- it.slashdot.org — OpenAI Finds Evidence Other AI Agents Escaped Containment
- adversa.ai — Nine AI coding agent incidents that ended with deleted data
- simplysecuregroup.com — Mythos 5 and GPT-5.6-Sol Agents Went Beyond Their Cyber Test and Targeted the Real World
- hackernews.com — Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
- Check Point Research — Three AI security disclosures, fourteen days: what the warnings signs are telling us
- serisec.com — AI Browsers Vulnerable to ‘PleaseFix’ Zero-Click Agent Hijacking
- feeds.feedburner.com — Claude Code and Gemini CLI Flaws Let a GitHub Issue Reach CI Workflow Secrets
- csoonline.com — Trojanized AI skills gain 1.7M installs in agent-targeted attack
- TechNadu — Weekly Cybersecurity Roundup: Entering an Era When AI Agents Take Unapproved Paths as Security Teams Race to Trace Them
- Cybersecurity News — Claude Opus 5 Cuts Indirect Prompt Injection Attack Success to 2% in New Benchmark Analysis
- Check Point Research — Native AI Security Comes to Claude: Why Anthropic’s Inference Hooks Matter
- arXiv (Computer Science - Cryptography and Security) — When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output
- arXiv (Computer Science - Cryptography and Security) — Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards
- gbhackers.com — OpenAI Launches GPT-5.6-Cyber to Find Zero-Day Vulnerabilities and Develop Exploit Chains
- feeds.feedburner.com — OpenAI Launches GPT-5.6-Cyber with Reduced Safeguards for Exploit Development
- NSFOCUS — AI Security Incident Case: AISI Reveals AI Agents Autonomously Attacking Real People and Systems During Security Testing
- arXiv (Computer Science - Cryptography and Security) — Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- forkast.news — Grok 4.6 Matches GPT-5.6 Sol on Composite Intelligence — But SpaceXAI Still Won’t Document What It Does Autonomously
- eSecurity Planet — Claude Agents Started a ‘Turf War’ That Escalated to Self-Replicating Malware
- NewsBytes — Zhipu's GLM-5.3 AI model outperforms Anthropic's Mythos 5 in cybersecurity
- Dark Reading — AI Browsers Vulnerable to 'PleaseFix' Zero-Click Agent Hijacking
- gbhackers.com — Claude Code Auto Mode Blocks 89% of Dangerous Commands and Prompt Injection Attacks
- Dark Reading — No Perfect Fix for AI Browser Prompt Injection Flaws
- Infosecurity-magazine
- tomshardware.com — New hack exploits AI hallucinations to trick agents into running malicious code — 'HalluSquatting' attack exploits a fundamental weakness in every available model
- feeds.feedburner.com — New Agent Data Injection Attack Can Make AI Agents Misclick or Run Attacker Commands
- Island
- Cybersecuritynews
- Lenet
- Techradar
- Mitiga
- hackernews.com — OpenAI and Hugging Face partner to address security incident
- news.ycombinator.com — OpenAI’s accidental attack against Hugging Face is science fiction that happened
- DEV Community — Claude Opus 5 is Here: What Developers Need to Know About the Safety "Fine Print"
- helpnetsecurity.com — Hugging Face breach reignites open-weights debate, raises liability questions
- news.ycombinator.com — Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the Incident
- Thehackernews
- Researchgate
- Cryptopolitan
- Github
- Arxiv
- Futurice
- Themoonlight
- Asanify
- Tomshardware
- hackernews.com — Investigating three real-world incidents in our cybersecurity evaluations
- cyberscoop.com — Anthropic says its AI accidentally hacked three companies during safety tests
- Businessinsider
- Straitstimes
- Ft
- Community
- Mashable
- Economictimes
- Foxbusiness
- Valueaddvc
- Mallory
- Deploymentsafety
- Labs
- Www-cdn
- Roo
- Neuraltrust
- Github
- Venturebeat
- Cryptobriefing
- Eu
- Japantimes
- Kfgo
- Dobetter
- Cbc
- Hiddenlayer
- Forbes
- Aijourn
- Linx
- Zenity
- Nhimg
- Cltc
- Labs
- Genai
- Simbian
- Securityboulevard
- The-decoder
- Synapsehd
- Noma
- Forbes
- Cbsnews
- Time
- Japantimes
- Mashable
- cyberscoop.com — AISI, OpenAI report more ‘unsanctioned’ model hacks
- bleepingcomputer.com — OpenAI, Anthropic AI agents targeted real people and systems in cyber tests
- Itnews
- Aisi
- Bworldonline
- cybersecuritydive.com — OpenAI warns autonomous hacks are ‘watershed moment for computer security’
- gbhackers.com — Critical Flaws in Claude Code, Gemini CLI, and OpenAI Codex Enable RCE and Supply Chain Attacks
- Labs
- Esecurityplanet
- Devops
- Medium
- Daily
- Labs
- Arxiv
- Labs
- Github
- Python
- Theguardian
- Itpro
- Zenity
- Towardsdatascience
- Jackmaguire
- Youtube
- Labs
- Engadget
- Openai
- thenewstack.io — OpenAI built a model it doesn’t want most people to use
- Venturebeat
- Helpnetsecurity
- Poloniex
- Eesel
- Trendingtopics
- Analyticsinsight
- Engadget
- Openai
- Timesofindia
- Defenseone
- Themoonlight
- Boozallen
- Researchgate
- Industrialcyber
- Sandia
- Csis
- Youtube
- News
- Frenos
- Blogs
- Pdxscholar
- Neuraltrust
- Youtube
- Alluresecurity
- Enterprisedna
- Adsadvance
- Forkast
- Hcamag
- Cyberdaily
- Arxiv
- App
- Lbank
- Unite
- Venturebeat
- The-independent
- Businessinsider
- Relvehq
- Anthropic
- Startupfortune