AI Goat v2.0 is out. Itβs the largest release of the open source AI security playground so far, and weβre announcing it at c0c0n 2026 in Kochi.
In v0.3, AI Goat was an LLM security lab with one target, the Cracky chatbot, and 17 labs on the OWASP LLM Top 10. That was the right scope when most AI features were chatbots. It isnβt any more. Production AI applications now retrieve documents, connect to MCP servers, call tools, remember what they saw and act on their own, and every one of those steps is a place where an attacker can plant instructions.
v2.0 adds the attack surfaces that appear once a model can act: a tool-calling agent, an MCP host and client connected to real MCP servers, agent memory, and an end-to-end kill chain. The number of labs grows from 17 to 46, mapped to three OWASP frameworks. The whole platform still runs on your own machine, and it now also runs in Google Colab.
TL;DR
- 46 hands-on labs (up from 17): 23 for the OWASP LLM Top 10 (2026), 12 for the OWASP Agentic Top 10 (2026) and 11 for the OWASP MCP Top 10 (2025, beta).
- New attack surfaces: a tool-calling shop agent with memory, plus an MCP client and host built on the official MCP Python SDK, with four local MCP servers.
- Agentic Kill Chain capstone: a poisoned product review becomes a compromised agent, and you replay it in Vulnerable, Defended and Guardrailed modes.
- Defenses rebuilt as 19 composable controls, with three defense levels on every surface, not just the chatbot.
- New learning tools: an OWASP Top 10 hub with risk pages for all three lists, and a Threat Modeling workbench.
- Run it anywhere: locally, in Docker (now reachable from your lab network) or in Google Colab.

v0.3 vs v2.0 at a glance
| Area | v0.3 | v2.0 |
|---|---|---|
| Frameworks | OWASP LLM Top 10 (2025) | OWASP LLM Top 10 (2026), Agentic Top 10 (2026), MCP Top 10 (2025, beta) |
| Labs | 17 | 46 (23 LLM, 12 Agentic, 11 MCP) |
| Attack surfaces | Chat | Chat, RAG, agent, MCP client, MCP host |
| Defenses | 3 levels on the chatbot | 3 levels on every surface, built from 19 composable controls |
| Capstone | None | Agentic Kill Chain: The Compromised eCommerce Agent |
| Learning | Attack Labs and Challenges | Adds an OWASP Top 10 hub, risk pages and a Threat Modeling workbench |
| Run options | Local, Docker | Local, Docker (network-reachable), Google Colab |
| Database | create_all on startup | Alembic migrations |
| Backend test files | 12 | 67, plus frontend component tests in CI |
Why v2.0: the model is only one way in
A prompt injection against a chatbot gets you a bad answer. The same injection against an agent that can issue refunds, export customer records or call an MCP server gets you a bad action, and that action may happen days later, on a request nobody would think twice about.
Thatβs the shift v2.0 is built around. The new OWASP lists reflect it too: the OWASP Top 10 for Agentic Applications describes what goes wrong when AI systems act, and the OWASP MCP Top 10 covers the protocol that connects them to tools. AI Goat v2.0 gives you a working, intentionally vulnerable application where you can exploit each of those risks and then see which defenses actually stop them.
New attack surfaces: agents and MCP
A tool-calling shop agent
The new shop agent (agent.runner) plans, calls tools and answers. AI Goat treats the plannerβs output as untrusted: every tool call passes an Intent Gate, which validates the arguments and applies the active defense profile before the tool runs. Each run shows the plan, every tool call and any pending approvals. Sensitive tools such as refunds and customer exports can be held for human approval.
The key lesson in these labs: the tool call is the exploit. Labs score the tool_call the agent actually made, not what it said in its reply.
Agent memory
The agent can remember and recall notes per user and per lab. At Level 0, recalled notes are trusted as policy, which is exactly the problem. Level 2 adds a memory scan. If youβve only ever thought of prompt injection as something that ends with the conversation, memory is where that assumption breaks.

An MCP client and host connected to real MCP servers
The MCP labs use the official mcp Python SDK (protocol 2026-07-28, stateless) over stdio. This is real protocol traffic, not a mock. Four allow-listed servers ship with AI Goat:
- an official catalog server,
- a community support server,
- an untrusted lookalike, and
- an internal management server.
Servers are spawned with a scrubbed environment and a fixed command, so a request can never choose what runs. A staff admin assistant, an MCP-speaking assistant used by the host labs, is one click away through the new Switch to Admin shortcut in the header.

A support desk
Customers can now open support tickets with attachments, and staff manage them in a feedback inbox. Tickets are a realistic injection channel: several agent and MCP labs start with a planted ticket.
Labs: from 17 to 46, on three frameworks
OWASP LLM Top 10 (2026): 23 labs
All LLM labs are renumbered to the 2026 edition of the OWASP LLM Top 10. New labs cover indirect injection through retrieved documents, unauthorized retrieval, tool allowlist bypass, retrieval keyword stuffing, trust-tier spoofing and stale knowledge, and LLM05 is rewritten as Data and Model Poisoning.
OWASP Agentic Top 10 (2026): 12 labs
There is one lab per risk, from ASI01 Agent Goal Hijack to ASI10 Rogue Agents, plus a second memory-poisoning lab for practising Level 2, and the kill chain capstone. Highlights:
- ASI01: a planted ticket rewrites the admin assistantβs goal.
- ASI02: talk the shop agent into applying a staff-only coupon.
- ASI03: export another customerβs data.
- ASI05: reach a shell executor that runs in a disposable Docker container with no network and a read-only root filesystem. Nothing runs on your host.
- ASI06: memory poisoning, then watch the Level 2 memory scan catch it.
- ASI09: an approval dialog with a plausible summary that doesnβt match the real tool arguments.
Each agentic lab has a learner guide in the repository under docs/labs/agentic/.
OWASP MCP Top 10 (2025, beta): 11 labs
The MCP labs cover tool description poisoning, rug-pull redefinition, schema drift, decoy tokens, lookalike integrations, poisoned tickets, token passthrough, session reconstruction, unverified server identity and context over-sharing.
MCP05 (command injection) is documented, but deliberately has no executable sink. AI Goat keeps its vulnerabilities in the AI layer, never in the host OS. If you want to practise code execution, ASI05 does it safely in a sandbox.
Labs are now data, not code
Lab content has moved out of a hardcoded frontend object into YAML manifests (config/labs/*.yml: llm, rag, agent, mcp), served by GET /api/labs/. Each lab declares its risks, surface, difficulty, objective, example payloads and the expected result at each defense level. That makes labs much easier to review, and to contribute.
Evidence-based scoring for MCP labs
MCP labs record what actually happened (tool listings, calls and results) and score your submitted finding against that evidence. They also support hints, replay and per-learner rug-pull state. You canβt finish an MCP lab by guessing; your conclusion has to be supported by the transcript.
The capstone: Agentic Kill Chain
killchain-1, The Compromised eCommerce Agent, shows why memory poisoning is more than a one-shot prompt injection.
- A hidden instruction goes into a product review, or into a PDF invoice as white 7 pt text.
- A naive ingestion step writes it into connector memory.
- The agent turns it into a standing note in its own memory.
- Later, a routine admin request retrieves that note.
- The agent acts on it: a customer data export to an attacker, an internal coupon disclosure, or a $1.00 price quote.
Then you replay the chain in three modes:
- Vulnerable: the poison persists and sensitive tools run with no approval.
- Defended: the poison is still there, but each sensitive operation is held for an administrator approval bound to the exact arguments.
- Guardrailed: adds rails that an approval cannot override (an egress allowlist, card-number masking, a staff-coupon rule, an ingestion scan and answer masking), so approving the wrong call still doesnβt send data outside the shop.

The workbench at /challenges?killchain=1 includes a live trace, a memory inspector, an attacker inbox, layered cleanup and a hard reset. Try clearing only the agentβs memory and see what happens on the next request: the connector still holds the original text, and it comes straight back. Lab data is copied from the shop into its own tables, and nothing leaves your machine.
Defenses, rebuilt as composable controls
In v0.3, the three defense levels applied only to the chatbot. In v2.0 they apply to every surface, through per-surface defense profiles (config/defense_profiles.yml) that map each surface and level to an ordered chain of controls. Level 0 is always empty.
New controls cover:
- Retrieval: provenance, access control and an injection scan.
- Tools: an allowlist, a coupon policy, human approval and a result scan.
- Memory: a memory scan.
- MCP: tool pinning, schema pinning, origin pinning, a description scan and a result scan.
These sit alongside the existing input validation, intent classification, output moderation and NeMo Guardrails rails, for 19 controls in total. At Level 2, NeMo Guardrails now also checks agent goals, tool results and answers. If nemoguardrails isnβt installed, the same checks run locally and the result tells you so.
RAG upgrades

- Knowledge-base documents now carry retrieval provenance and trust tiers, which are the target of the new trust-tier spoofing lab.
- Optional hybrid retrieval (BM25 plus vectors, merged with reciprocal rank fusion) is available. Itβs off by default.
- A retrieval trace endpoint and vector sync status show what was ranked, dropped and placed in the prompt.
A better learning experience
- Navigation regrouped into Learn (OWASP Top 10, Threat Modeling), Try (Attack Labs, Challenges) and Console (RAG, MCP, Agent).
- OWASP Top 10 hub and risk pages show the three frameworks side by side, with the labs and challenges mapped to each risk. You can read our guide to all 30 risks here too.
- Threat Modeling workbench, with an architecture diagram of AI Goat itself, a framework picker (STRIDE, MITRE ATLAS and others) and six worked scenarios.
- One hub layout across Attack Labs, Challenges, RAG, MCP, Agent, OWASP and Threat Modeling.
Run it anywhere
- Google Colab: a community-contributed notebook (
colab_notebooks/AIGoat_Colab.ipynb) runs the full stack in a Colab runtime, with an optional tool-calling model. It prints the app link on Colabβs own domain and a separate link to the API docs. - Docker, reachable from your network: published ports now bind to
0.0.0.0, and the frontend calls whichever host opened the page, so one lab machine can serve a whole classroom. start.shpins Python 3.11 to match CI, runs Alembic migrations (stamping an existing v0.3 database first), and syncs the support inbox and challenge metadata on every start.
AI Goat is deliberately vulnerable. Run it on your own machine or a trusted lab network, never on the internet. If you use Docker on a shared network, publish the ports on
127.0.0.1yourself.
Reliability fixes that change model behavior
- System prompts now reach the model on
/api/chat. Ollama ignores a top-levelsystemfield on that endpoint, so agent, MCP and RAG calls used to run without their system prompt. They now send it as a leading system message, so expect live-model behavior on those labs to differ from earlier builds. - Agent step budget: agent runs reserve their last step for a written answer, so runs no longer end with a blank reply, and native tool-call history is replayed to Ollama correctly.
- Stopping a chat now cancels the Ollama stream on the server.
Get started with AI Goat v2.0
New install:
git clone https://github.com/AISecurityConsortium/AIGoat.git
cd AIGoat
./scripts/start.sh
Or with Docker (give Docker Desktop at least 12 GB of RAM):
docker volume create ollama_models
cd AIGoat/docker
docker compose up --build
Then open http://localhost:3000, log in as alice / password123, and use Switch to Admin for the staff-only labs.
For the agent, MCP host and kill chain labs, pull a model with native tool calls and select it in the console:
ollama pull qwen3.5:9b
Mistral remains the default and is enough for the chat and RAG labs. The getting started guide covers hardware requirements and your first attack.
Upgrading from v0.3
- Pull and restart:
./scripts/stop.sh && git pull && ./scripts/start.sh. Your existing database is stamped and migrated in place. For a clean slate, use./scripts/start.sh --fresh. Docker users rundocker compose up --build. - New Python dependencies (installed by
start.sh):alembic,mcp>=2.2.0,pypdf,bm25sandeval-type-backport. - Lab IDs follow the 2026 numbering. For example, the old
llm10-1(Token Flood) is nowllm06-1. A migration renames stored lab sessions, agent runs and memory, and the frontend migrates saved lab completion once. Update any bookmarks, scripts or workshop handouts that use old IDs. - Unchanged: seed users and passwords, the 9 CTF challenges and their points, the
AIGOAT{32 hex}flag format and its per-user derivation, and existing API paths. All API changes are additive.
Known limitations
Weβd rather tell you than have you find out mid-workshop:
- Local models follow planted instructions probabilistically. A lab that doesnβt trigger records that honestly instead of faking the impact. Retry, or use a tool-capable model.
- Mistral doesnβt emit native tool calls, so the agent and MCP host labs show no tool calls with it.
- The Colab notebook clones
mainand keeps no state between sessions, and Colab links expire when the runtime disconnects. - MCP05 has no lab, by design.
Thank you, community
v2.0 includes community contributions:
- The Google Colab notebook and guide (#5).
- Python 3.9 type hint compatibility for the SQLAlchemy models (#7).
- A fix for a
KeyErrorwhen creating a coupon without optional fields (#3).
Thank you to everyone who opened issues and pull requests. If youβd like to contribute a lab, a defense or a fix, start with the AI Goat repository on GitHub. Labs are YAML now, so itβs easier than ever.
FAQ
What is AI Goat?
AI Goat (AIGoat) is a free, open source, intentionally vulnerable AI security playground built by the AI Security Consortium. Itβs an AI-powered e-commerce shop with a chatbot, a RAG knowledge base, MCP integrations and a tool-calling agent that you attack and defend through guided labs, entirely on your own machine.
Whatβs new in AI Goat v2.0?
v2.0 adds a tool-calling agent with memory, an MCP client and host with four real MCP servers, an Agentic Kill Chain capstone, composable defenses on every surface, RAG provenance and retrieval tracing, an OWASP Top 10 hub, a Threat Modeling workbench and Google Colab support. Labs grow from 17 to 46.
Which OWASP frameworks does AI Goat v2.0 cover?
The OWASP Top 10 for LLM Applications (2026) with 23 labs, the OWASP Top 10 for Agentic Applications (2026) with 12 labs, and the OWASP MCP Top 10 (2025, beta) with 11 labs.
Do I need a GPU or a cloud API key?
No. AI Goat runs locally with Ollama and needs no API keys. A GPU is optional but makes replies much faster. Plan for about 8 GB of free RAM, and use a tool-calling model such as qwen3.5:9b for the agent and MCP labs.
Happy hacking, and see you at c0c0n.