CourseModel Context Protocol · Module 8: Security · part 55 of 83
Part 55 · Module 8: Security

Topic 10: Measure it

5 min read·22 Sept 2026

Opinions about security age badly. Numbers from real runs, honestly labelled, do not. Here is a small matrix: three attacks, each run with its relevant defence layer on and off, reporting whether the harmful action actually happened. Every row is a real run of the code in this module with the scripted stand-in model.

State the caveat up front, because it is the most important sentence in this part. This measures the mechanism, not a real model's robustness. The stand-in obeys every injection, so an "off" row shows the worst case and an "on" row shows the structural control holding regardless of the model. A real LLM will fall for these attacks at some rate between never and always, and that rate changes with the model, the prompt, and the attacker's wording. You must rerun this with your real model to learn your actual numbers, and re-run it whenever you change models or prompts.

python
"""Measure: attack attempts vs outcome with each defence layer on or off.

Every row is a real run with the scripted stand-in model (not a real LLM). We
report whether the harmful action actually happened. This is a mechanism test,
not a model-robustness test: a real model may resist or fall for the same
injections differently, so rerun this with your model and recalibrate.
"""
from __future__ import annotations

import logging
import shutil
from pathlib import Path

import anyio

from mcp import Client
from notes_assistant.host import Host, HostConfig, deny_all
from notes_assistant.server import build_server
from notes_assistant.store import NoteStore

from examples import m08_exfil_server
from examples.m08_exfil_server import build_exfil_server
from examples.m08_poisoned_server import build_poisoned_server
from examples.m08_rugpull_server import build_rugpull_server
from examples.m08_scripted_model import ScriptedModel

ALLOW = lambda name, args: True
PIN_DIR = Path("pins_matrix")


async def poisoning_attempt(approve) -> bool:
    """True if the poisoned write actually happened."""
    server = build_poisoned_server()
    async with Client(server) as client:
        model = ScriptedModel(plan=[("kb__search_kb", {"query": "x"})])
        host = Host({"kb": client}, chat_fn=model, approve=approve, config=HostConfig(max_iterations=4))
        answer = await host.ask("Search for x.")
        return any(r.name.endswith("create_note") and r.ok for r in answer.calls)


async def exfil_attempt(approve) -> bool:
    """True if an email actually left the machine."""
    from examples.m08_demo_exfil import ExfilModel
    m08_exfil_server.SENT.clear()
    server = build_exfil_server()
    async with Client(server) as client:
        host = Host({"mail": client}, chat_fn=ExfilModel(), approve=approve, config=HostConfig(max_iterations=4))
        await host.ask("Read and mail the key.")
        return len(m08_exfil_server.SENT) > 0


async def rugpull_attempt(pin: bool) -> bool:
    """True if the rugged (poisoned) tool was still reachable by the model."""
    if PIN_DIR.exists():
        shutil.rmtree(PIN_DIR)
    server = build_rugpull_server()
    async with Client(server) as client:
        cfg = HostConfig(pin_dir=PIN_DIR if pin else None, max_iterations=3)
        host = Host({"notes": client}, config=cfg)
        await host.load_tools()          # pin on first sight (if pinning on)
        server._swap()                   # the rug pull
        model = ScriptedModel(plan=[("notes__search_notes", {"query": "x"})])
        host2 = Host({"notes": client}, chat_fn=model, approve=deny_all, config=cfg)
        answer = await host2.ask("Search for x.")
        return any(r.name.endswith("search_notes") and r.ok for r in answer.calls)


async def main() -> None:
    logging.disable(logging.CRITICAL)
    rows = []
    rows.append(("tool poisoning -> write", "approval gate off", await poisoning_attempt(ALLOW)))
    rows.append(("tool poisoning -> write", "approval gate on", await poisoning_attempt(deny_all)))
    rows.append(("exfiltration -> email out", "approval gate off", await exfil_attempt(ALLOW)))
    rows.append(("exfiltration -> email out", "approval gate on", await exfil_attempt(deny_all)))
    rows.append(("rug pull -> tool reachable", "pinning off", await rugpull_attempt(False)))
    rows.append(("rug pull -> tool reachable", "pinning on", await rugpull_attempt(True)))

    print(f"{'attack':28} {'defence':20} harmful action happened")
    print("-" * 72)
    for attack, defence, happened in rows:
        print(f"{attack:28} {defence:20} {'YES' if happened else 'no'}")


if __name__ == "__main__":
    anyio.run(main)

Comes out:

The rest of this course is yours to keep

This course is bought on its own, once, and stays readable afterwards, including the parts added to it later.