Skip to content
AI Goat - AI Security Playground AI GoatTM
Β· 12 min read Β· AI Security Consortium

Announcing AI Goat v2.0: AI Red Teaming for MCP Servers, AI Agents and RAG

AI Goat v2.0 grows from 17 to 46 labs across the OWASP LLM, Agentic and MCP Top 10, adding a tool-calling agent, real MCP servers, agent memory and a kill chain.

AI Goat v2.0 release MCP security agentic AI security AI red teaming OWASP Agentic Top 10 OWASP MCP Top 10 OWASP LLM Top 10 2026

AI Goat v2.0 is out. It’s the largest release of the open source AI security playground so far, and we’re announcing it at c0c0n 2026 in Kochi.

In v0.3, AI Goat was an LLM security lab with one target, the Cracky chatbot, and 17 labs on the OWASP LLM Top 10. That was the right scope when most AI features were chatbots. It isn’t any more. Production AI applications now retrieve documents, connect to MCP servers, call tools, remember what they saw and act on their own, and every one of those steps is a place where an attacker can plant instructions.

v2.0 adds the attack surfaces that appear once a model can act: a tool-calling agent, an MCP host and client connected to real MCP servers, agent memory, and an end-to-end kill chain. The number of labs grows from 17 to 46, mapped to three OWASP frameworks. The whole platform still runs on your own machine, and it now also runs in Google Colab.

TL;DR

  • 46 hands-on labs (up from 17): 23 for the OWASP LLM Top 10 (2026), 12 for the OWASP Agentic Top 10 (2026) and 11 for the OWASP MCP Top 10 (2025, beta).
  • New attack surfaces: a tool-calling shop agent with memory, plus an MCP client and host built on the official MCP Python SDK, with four local MCP servers.
  • Agentic Kill Chain capstone: a poisoned product review becomes a compromised agent, and you replay it in Vulnerable, Defended and Guardrailed modes.
  • Defenses rebuilt as 19 composable controls, with three defense levels on every surface, not just the chatbot.
  • New learning tools: an OWASP Top 10 hub with risk pages for all three lists, and a Threat Modeling workbench.
  • Run it anywhere: locally, in Docker (now reachable from your lab network) or in Google Colab.
AI Goat v2.0 Attack Labs hub with tabs for OWASP LLM 2026 (23 labs), MCP 2025 (11 labs) and Agentic 2026 (12 labs)
The v2.0 Attack Labs hub: pick a framework, pick a risk, open a lab.

v0.3 vs v2.0 at a glance

Areav0.3v2.0
FrameworksOWASP LLM Top 10 (2025)OWASP LLM Top 10 (2026), Agentic Top 10 (2026), MCP Top 10 (2025, beta)
Labs1746 (23 LLM, 12 Agentic, 11 MCP)
Attack surfacesChatChat, RAG, agent, MCP client, MCP host
Defenses3 levels on the chatbot3 levels on every surface, built from 19 composable controls
CapstoneNoneAgentic Kill Chain: The Compromised eCommerce Agent
LearningAttack Labs and ChallengesAdds an OWASP Top 10 hub, risk pages and a Threat Modeling workbench
Run optionsLocal, DockerLocal, Docker (network-reachable), Google Colab
Databasecreate_all on startupAlembic migrations
Backend test files1267, plus frontend component tests in CI

Why v2.0: the model is only one way in

A prompt injection against a chatbot gets you a bad answer. The same injection against an agent that can issue refunds, export customer records or call an MCP server gets you a bad action, and that action may happen days later, on a request nobody would think twice about.

That’s the shift v2.0 is built around. The new OWASP lists reflect it too: the OWASP Top 10 for Agentic Applications describes what goes wrong when AI systems act, and the OWASP MCP Top 10 covers the protocol that connects them to tools. AI Goat v2.0 gives you a working, intentionally vulnerable application where you can exploit each of those risks and then see which defenses actually stop them.

New attack surfaces: agents and MCP

A tool-calling shop agent

The new shop agent (agent.runner) plans, calls tools and answers. AI Goat treats the planner’s output as untrusted: every tool call passes an Intent Gate, which validates the arguments and applies the active defense profile before the tool runs. Each run shows the plan, every tool call and any pending approvals. Sensitive tools such as refunds and customer exports can be held for human approval.

The key lesson in these labs: the tool call is the exploit. Labs score the tool_call the agent actually made, not what it said in its reply.

Agent memory

The agent can remember and recall notes per user and per lab. At Level 0, recalled notes are trusted as policy, which is exactly the problem. Level 2 adds a memory scan. If you’ve only ever thought of prompt injection as something that ends with the conversation, memory is where that assumption breaks.

AI Goat v2.0 ASI06-1 Memory Poisoning lab with objective, steps, goal input and standing shop notes
ASI06-1 Memory Poisoning: plant a standing note, then show that a later run treats it as trusted policy.

An MCP client and host connected to real MCP servers

The MCP labs use the official mcp Python SDK (protocol 2026-07-28, stateless) over stdio. This is real protocol traffic, not a mock. Four allow-listed servers ship with AI Goat:

  • an official catalog server,
  • a community support server,
  • an untrusted lookalike, and
  • an internal management server.

Servers are spawned with a scrubbed environment and a fixed command, so a request can never choose what runs. A staff admin assistant, an MCP-speaking assistant used by the host labs, is one click away through the new Switch to Admin shortcut in the header.

AI Goat v2.0 MCP03-1 Tool Description Poisoning lab with MCP interaction history, server identity and submit evidence panels
MCP03-1 Tool Description Poisoning: identify the server, list its tools, then submit the evidence that supports your finding.

A support desk

Customers can now open support tickets with attachments, and staff manage them in a feedback inbox. Tickets are a realistic injection channel: several agent and MCP labs start with a planted ticket.

Labs: from 17 to 46, on three frameworks

OWASP LLM Top 10 (2026): 23 labs

All LLM labs are renumbered to the 2026 edition of the OWASP LLM Top 10. New labs cover indirect injection through retrieved documents, unauthorized retrieval, tool allowlist bypass, retrieval keyword stuffing, trust-tier spoofing and stale knowledge, and LLM05 is rewritten as Data and Model Poisoning.

OWASP Agentic Top 10 (2026): 12 labs

There is one lab per risk, from ASI01 Agent Goal Hijack to ASI10 Rogue Agents, plus a second memory-poisoning lab for practising Level 2, and the kill chain capstone. Highlights:

  • ASI01: a planted ticket rewrites the admin assistant’s goal.
  • ASI02: talk the shop agent into applying a staff-only coupon.
  • ASI03: export another customer’s data.
  • ASI05: reach a shell executor that runs in a disposable Docker container with no network and a read-only root filesystem. Nothing runs on your host.
  • ASI06: memory poisoning, then watch the Level 2 memory scan catch it.
  • ASI09: an approval dialog with a plausible summary that doesn’t match the real tool arguments.

Each agentic lab has a learner guide in the repository under docs/labs/agentic/.

OWASP MCP Top 10 (2025, beta): 11 labs

The MCP labs cover tool description poisoning, rug-pull redefinition, schema drift, decoy tokens, lookalike integrations, poisoned tickets, token passthrough, session reconstruction, unverified server identity and context over-sharing.

MCP05 (command injection) is documented, but deliberately has no executable sink. AI Goat keeps its vulnerabilities in the AI layer, never in the host OS. If you want to practise code execution, ASI05 does it safely in a sandbox.

Labs are now data, not code

Lab content has moved out of a hardcoded frontend object into YAML manifests (config/labs/*.yml: llm, rag, agent, mcp), served by GET /api/labs/. Each lab declares its risks, surface, difficulty, objective, example payloads and the expected result at each defense level. That makes labs much easier to review, and to contribute.

Evidence-based scoring for MCP labs

MCP labs record what actually happened (tool listings, calls and results) and score your submitted finding against that evidence. They also support hints, replay and per-learner rug-pull state. You can’t finish an MCP lab by guessing; your conclusion has to be supported by the transcript.

The capstone: Agentic Kill Chain

killchain-1, The Compromised eCommerce Agent, shows why memory poisoning is more than a one-shot prompt injection.

  1. A hidden instruction goes into a product review, or into a PDF invoice as white 7 pt text.
  2. A naive ingestion step writes it into connector memory.
  3. The agent turns it into a standing note in its own memory.
  4. Later, a routine admin request retrieves that note.
  5. The agent acts on it: a customer data export to an attacker, an internal coupon disclosure, or a $1.00 price quote.

Then you replay the chain in three modes:

  • Vulnerable: the poison persists and sensitive tools run with no approval.
  • Defended: the poison is still there, but each sensitive operation is held for an administrator approval bound to the exact arguments.
  • Guardrailed: adds rails that an approval cannot override (an egress allowlist, card-number masking, a staff-coupon rule, an ingestion scan and answer masking), so approving the wrong call still doesn’t send data outside the shop.
AI Goat v2.0 Agentic Kill Chain workbench with the review-poisoning attack source and the Agentic Cracky operations agent
The kill chain workbench: attack sources, the operations agent, and below them a memory inspector, live trace and attacker inbox.

The workbench at /challenges?killchain=1 includes a live trace, a memory inspector, an attacker inbox, layered cleanup and a hard reset. Try clearing only the agent’s memory and see what happens on the next request: the connector still holds the original text, and it comes straight back. Lab data is copied from the shop into its own tables, and nothing leaves your machine.

Defenses, rebuilt as composable controls

In v0.3, the three defense levels applied only to the chatbot. In v2.0 they apply to every surface, through per-surface defense profiles (config/defense_profiles.yml) that map each surface and level to an ordered chain of controls. Level 0 is always empty.

New controls cover:

  • Retrieval: provenance, access control and an injection scan.
  • Tools: an allowlist, a coupon policy, human approval and a result scan.
  • Memory: a memory scan.
  • MCP: tool pinning, schema pinning, origin pinning, a description scan and a result scan.

These sit alongside the existing input validation, intent classification, output moderation and NeMo Guardrails rails, for 19 controls in total. At Level 2, NeMo Guardrails now also checks agent goals, tool results and answers. If nemoguardrails isn’t installed, the same checks run locally and the result tells you so.

RAG upgrades

AI Goat v2.0 RAG page with knowledge documents, trust tiers, retrieval trace and index pipeline status
The redesigned RAG page: documents with provenance and trust tiers, a retrieval trace and vector sync status.
  • Knowledge-base documents now carry retrieval provenance and trust tiers, which are the target of the new trust-tier spoofing lab.
  • Optional hybrid retrieval (BM25 plus vectors, merged with reciprocal rank fusion) is available. It’s off by default.
  • A retrieval trace endpoint and vector sync status show what was ranked, dropped and placed in the prompt.

A better learning experience

  • Navigation regrouped into Learn (OWASP Top 10, Threat Modeling), Try (Attack Labs, Challenges) and Console (RAG, MCP, Agent).
  • OWASP Top 10 hub and risk pages show the three frameworks side by side, with the labs and challenges mapped to each risk. You can read our guide to all 30 risks here too.
  • Threat Modeling workbench, with an architecture diagram of AI Goat itself, a framework picker (STRIDE, MITRE ATLAS and others) and six worked scenarios.
  • One hub layout across Attack Labs, Challenges, RAG, MCP, Agent, OWASP and Threat Modeling.

Run it anywhere

  • Google Colab: a community-contributed notebook (colab_notebooks/AIGoat_Colab.ipynb) runs the full stack in a Colab runtime, with an optional tool-calling model. It prints the app link on Colab’s own domain and a separate link to the API docs.
  • Docker, reachable from your network: published ports now bind to 0.0.0.0, and the frontend calls whichever host opened the page, so one lab machine can serve a whole classroom.
  • start.sh pins Python 3.11 to match CI, runs Alembic migrations (stamping an existing v0.3 database first), and syncs the support inbox and challenge metadata on every start.

AI Goat is deliberately vulnerable. Run it on your own machine or a trusted lab network, never on the internet. If you use Docker on a shared network, publish the ports on 127.0.0.1 yourself.

Reliability fixes that change model behavior

  • System prompts now reach the model on /api/chat. Ollama ignores a top-level system field on that endpoint, so agent, MCP and RAG calls used to run without their system prompt. They now send it as a leading system message, so expect live-model behavior on those labs to differ from earlier builds.
  • Agent step budget: agent runs reserve their last step for a written answer, so runs no longer end with a blank reply, and native tool-call history is replayed to Ollama correctly.
  • Stopping a chat now cancels the Ollama stream on the server.

Get started with AI Goat v2.0

New install:

git clone https://github.com/AISecurityConsortium/AIGoat.git
cd AIGoat
./scripts/start.sh

Or with Docker (give Docker Desktop at least 12 GB of RAM):

docker volume create ollama_models
cd AIGoat/docker
docker compose up --build

Then open http://localhost:3000, log in as alice / password123, and use Switch to Admin for the staff-only labs.

For the agent, MCP host and kill chain labs, pull a model with native tool calls and select it in the console:

ollama pull qwen3.5:9b

Mistral remains the default and is enough for the chat and RAG labs. The getting started guide covers hardware requirements and your first attack.

Upgrading from v0.3

  1. Pull and restart: ./scripts/stop.sh && git pull && ./scripts/start.sh. Your existing database is stamped and migrated in place. For a clean slate, use ./scripts/start.sh --fresh. Docker users run docker compose up --build.
  2. New Python dependencies (installed by start.sh): alembic, mcp>=2.2.0, pypdf, bm25s and eval-type-backport.
  3. Lab IDs follow the 2026 numbering. For example, the old llm10-1 (Token Flood) is now llm06-1. A migration renames stored lab sessions, agent runs and memory, and the frontend migrates saved lab completion once. Update any bookmarks, scripts or workshop handouts that use old IDs.
  4. Unchanged: seed users and passwords, the 9 CTF challenges and their points, the AIGOAT{32 hex} flag format and its per-user derivation, and existing API paths. All API changes are additive.

Known limitations

We’d rather tell you than have you find out mid-workshop:

  • Local models follow planted instructions probabilistically. A lab that doesn’t trigger records that honestly instead of faking the impact. Retry, or use a tool-capable model.
  • Mistral doesn’t emit native tool calls, so the agent and MCP host labs show no tool calls with it.
  • The Colab notebook clones main and keeps no state between sessions, and Colab links expire when the runtime disconnects.
  • MCP05 has no lab, by design.

Thank you, community

v2.0 includes community contributions:

  • The Google Colab notebook and guide (#5).
  • Python 3.9 type hint compatibility for the SQLAlchemy models (#7).
  • A fix for a KeyError when creating a coupon without optional fields (#3).

Thank you to everyone who opened issues and pull requests. If you’d like to contribute a lab, a defense or a fix, start with the AI Goat repository on GitHub. Labs are YAML now, so it’s easier than ever.

FAQ

What is AI Goat?

AI Goat (AIGoat) is a free, open source, intentionally vulnerable AI security playground built by the AI Security Consortium. It’s an AI-powered e-commerce shop with a chatbot, a RAG knowledge base, MCP integrations and a tool-calling agent that you attack and defend through guided labs, entirely on your own machine.

What’s new in AI Goat v2.0?

v2.0 adds a tool-calling agent with memory, an MCP client and host with four real MCP servers, an Agentic Kill Chain capstone, composable defenses on every surface, RAG provenance and retrieval tracing, an OWASP Top 10 hub, a Threat Modeling workbench and Google Colab support. Labs grow from 17 to 46.

Which OWASP frameworks does AI Goat v2.0 cover?

The OWASP Top 10 for LLM Applications (2026) with 23 labs, the OWASP Top 10 for Agentic Applications (2026) with 12 labs, and the OWASP MCP Top 10 (2025, beta) with 11 labs.

Do I need a GPU or a cloud API key?

No. AI Goat runs locally with Ollama and needs no API keys. A GPU is optional but makes replies much faster. Plan for about 8 GB of free RAM, and use a tool-calling model such as qwen3.5:9b for the agent and MCP labs.

Happy hacking, and see you at c0c0n.