CourseLarge Language Models · Module 8: Agents · part 41 of 80
Part 41 · Module 8: Agents

Part 6: Capabilities

14 min read·22 Sept 2026

MCP as an interop layer

So far our tools are Python functions in the same process. The Model Context Protocol (MCP) is an open protocol for exposing tools and context to LLM applications, so a tool written once can be used by any MCP-capable host (a chat app, an IDE, your agent). The current specification is version 2026-07-28 at modelcontextprotocol.io/specification/2026-07-28. In its terms:

  • A host is the LLM application (our agent). A client is the connector inside the host that talks to one server. A server provides capabilities.
  • Messages are JSON-RPC 2.0. Servers offer tools (functions the model can call), resources (data for context), and prompts (templates for users).
  • Transports are stdio (a local subprocess) and Streamable HTTP.

The 2026-07-28 revision made the protocol stateless. Per its changelog: the initialize handshake is gone and every request carries its protocol version and client capabilities in _meta; servers must implement a server/discover method; protocol-level sessions (and the Mcp-Session-Id header) are removed, so servers that need state across calls return explicit handles as ordinary tool arguments; every result carries a resultType ("complete", or "input_required" for the new multi round-trip pattern); tasks moved into an optional extension; and Roots, Sampling, and Logging are deprecated. If you read MCP tutorials from 2025, expect them to show the older handshake.

Here is our tool registry answering MCP-shaped requests. It is an in-process teaching shim, not a server: no transport, no authorization, three methods. It exists so you can read real MCP traffic and see how an MCP tool maps onto the tool specs our loop already uses.

python
"""Module 8: what our tools look like on the wire as MCP (protocol version 2026-07-28).

This is an in-process teaching shim, not a compliant MCP server: it has no
transport, no authorization, and only three methods. It shows the message shapes
so you can read real MCP traffic and see how any MCP server's tools become the
same tool specs our agent loop already uses. For real servers use an official SDK.
"""
from __future__ import annotations

import json
from pathlib import Path

from m08_agent import BillingStore, Tool, make_tools

PROTOCOL_VERSION = "2026-07-28"
READ_ONLY = {"search_kb", "read_article", "get_account", "get_invoice"}


def to_mcp_tool(tool: Tool) -> dict:
    """An MCP Tool definition. Annotations are hints for the host, and hosts must treat them as untrusted."""
    return {"name": tool.name, "description": tool.description, "inputSchema": tool.parameters,
            "annotations": {"readOnlyHint": tool.name in READ_ONLY,
                            "destructiveHint": tool.irreversible,
                            "idempotentHint": tool.name in READ_ONLY or tool.name == "issue_refund",
                            "openWorldHint": False}}


class McpShim:
    """Answers server/discover, tools/list, and tools/call from our tool registry."""

    def __init__(self, tools: dict[str, Tool]) -> None:
        self.tools = tools

    def handle(self, request: dict) -> dict:
        rid, method, params = request["id"], request["method"], request.get("params", {})
        if method == "server/discover":
            result = {"supportedVersions": [PROTOCOL_VERSION], "capabilities": {"tools": {"listChanged": False}},
                      "instructions": "Brightlane support tools. issue_refund needs human approval in the host."}
        elif method == "tools/list":
            result = {"tools": [to_mcp_tool(t) for t in self.tools.values()], "ttlMs": 300000, "cacheScope": "private"}
        elif method == "tools/call":
            tool = self.tools.get(params.get("name"))
            if tool is None:  # protocol error: the model cannot fix an unknown tool name by retrying arguments
                return {"jsonrpc": "2.0", "id": rid, "error": {"code": -32602, "message": f"Unknown tool: {params.get('name')}"}}
            try:
                value = tool.fn(**params.get("arguments", {}))
                result = {"content": [{"type": "text", "text": json.dumps(value)}], "structuredContent": value, "isError": False}
            except Exception as exc:  # tool execution error: returned to the model so it can self-correct
                result = {"content": [{"type": "text", "text": f"{type(exc).__name__}: {exc}"}], "isError": True}
        else:
            return {"jsonrpc": "2.0", "id": rid, "error": {"code": -32601, "message": f"Method not found: {method}"}}
        return {"jsonrpc": "2.0", "id": rid, "result": {"resultType": "complete", **result}}


def mcp_to_chat_spec(mcp_tool: dict, server_prefix: str) -> dict:
    """Turn an MCP tool into a Chat Completions tool spec, prefixed to avoid clashes between servers."""
    return {"type": "function", "function": {"name": f"{server_prefix}__{mcp_tool['name']}",
                                             "description": mcp_tool.get("description", ""),
                                             "parameters": mcp_tool["inputSchema"]}}


if __name__ == "__main__":
    Path("runs/m08").mkdir(parents=True, exist_ok=True)
    server = McpShim(make_tools(BillingStore("runs/m08/mcp_billing.json"), flaky_failures=0))
    meta = {"io.modelcontextprotocol/protocolVersion": PROTOCOL_VERSION,
            "io.modelcontextprotocol/clientInfo": {"name": "brightlane-agent", "version": "0.8"}}
    print(json.dumps(server.handle({"jsonrpc": "2.0", "id": 1, "method": "server/discover", "params": {"_meta": meta}})))
    listed = server.handle({"jsonrpc": "2.0", "id": 2, "method": "tools/list", "params": {"_meta": meta}})
    for t in listed["result"]["tools"]:
        hints = t["annotations"]
        print(f"  {t['name']:18} readOnly={hints['readOnlyHint']!s:5} destructive={hints['destructiveHint']}")
    ok = server.handle({"jsonrpc": "2.0", "id": 3, "method": "tools/call",
                        "params": {"_meta": meta, "name": "search_kb", "arguments": {"query": "duplicate charge refund"}}})
    print(json.dumps(ok)[:220])
    bad = server.handle({"jsonrpc": "2.0", "id": 4, "method": "tools/call",
                         "params": {"_meta": meta, "name": "get_invoice", "arguments": {"invoice_id": "INV-2026-999999"}}})
    print(json.dumps(bad))
    unknown = server.handle({"jsonrpc": "2.0", "id": 5, "method": "tools/call", "params": {"_meta": meta, "name": "delete_account"}})
    print(json.dumps(unknown))
    spec = mcp_to_chat_spec(listed["result"]["tools"][0], "brightlane")
    print("as a chat tool spec:", spec["function"]["name"])

Code explained

  • In simple words: dress our seven tools in MCP's clothes, send four requests, and turn one MCP tool back into a chat tool spec.
  • What happens:
    • to_mcp_tool builds an MCP tool definition: name, description, inputSchema (our JSON Schema, unchanged), and annotations. The annotation hints (readOnlyHint, destructiveHint, idempotentHint, openWorldHint) are in the spec's schema; the spec also says clients must treat annotations as untrusted unless the server is trusted. That is why our gate keys on our own irreversible flag, not on a server's hint.
    • McpShim.handle answers server/discover (supported versions and capabilities), tools/list (with the ttlMs and cacheScope caching fields this revision requires on list results), and tools/call.
    • The spec's two error kinds are both here: a tool execution error (isError: true inside a normal result) for problems the model can fix, like a bad invoice id, and a protocol error (a JSON-RPC error with code -32602) for an unknown tool.
    • mcp_to_chat_spec is the interop point: any server's tools become Chat Completions tool specs, prefixed with a server name because the spec only guarantees names are unique within one server.
  • Comes out:
text
  {"jsonrpc": "2.0", "id": 1, "result": {"resultType": "complete", "supportedVersions": ["2026-07-28"], "capabilities": {"tools": {"listChanged": false}}, "instructions": "Brightlane support tools. issue_refund needs human approval in the host."}}
    search_kb          readOnly=True  destructive=False
    read_article       readOnly=True  destructive=False
    get_account        readOnly=True  destructive=False
    get_invoice        readOnly=True  destructive=False
    issue_refund       readOnly=False destructive=True
    add_internal_note  readOnly=False destructive=False
    escalate_to_human  readOnly=False destructive=False
  {"jsonrpc": "2.0", "id": 3, "result": {"resultType": "complete", "content": [{"type": "text", "text": "[{\"article_id\": \"billing-refunds\", \"title\": \"Refunds and cancellations\", \"score\": 4.044}]"}], "structuredCo
  {"jsonrpc": "2.0", "id": 4, "result": {"resultType": "complete", "content": [{"type": "text", "text": "ValueError: No invoice 'INV-2026-999999'"}], "isError": true}}
  {"jsonrpc": "2.0", "id": 5, "error": {"code": -32602, "message": "Unknown tool: delete_account"}}
  as a chat tool spec: brightlane__search_kb

The destructive flag lands only on issue_refund. The failed invoice lookup comes back as a normal result with isError: true, which the host should pass to the model so it can correct itself; the unknown tool is a protocol error.

For real servers, use an official MCP SDK (the Python one is the mcp package on PyPI); check that its version supports the protocol revision your hosts speak, because this revision changed the wire format.

SituationUse thisWhy
Tools used only by this one agent, in one codebasePlain function tools (as in m08_agent.py)No protocol overhead, easiest to test
The same tools should serve several hosts (IDE, chat app, agents)An MCP serverWrite once, connect anywhere
Using third-party toolsAn MCP client in your host, with your own approval gate and allowlistYou get their tools; you keep control of what runs
Stateful tools (a cart, an open browser)Explicit handles returned by a create tool, passed back as argumentsThe 2026-07-28 spec has no protocol sessions

Code execution as a general-purpose tool

One tool that runs code replaces dozens of narrow tools: proration math, CSV reshaping, date arithmetic, checking a regex. Models are unreliable at arithmetic (Module 1) and good at writing short programs, so "write code, run it, read the output" is often more accurate than "think harder". The cost is risk: you are executing text the model wrote, and that text can be steered by anything in its context, including a malicious ticket (Module 11).

Here is a subprocess runner with time, CPU, memory, and file-size limits, plus a workspace-confined file tool. The last two cases show what these limits do not stop.

python
"""Module 8: code execution and workspace tools, with limits, and a demonstration of what the limits do NOT stop.

Linux or macOS only (uses the `resource` module). This is a teaching sandbox: it
bounds time, memory, output, and file size. It is NOT an isolation boundary.
"""
from __future__ import annotations

import os
import resource
import subprocess
import sys
import tempfile
import time
from pathlib import Path


def run_python(code: str, timeout_s: float = 2.0, cpu_s: int = 2, mem_mb: int = 256,
               file_mb: int = 1, max_output: int = 2000) -> dict:
    """Run untrusted Python in a child process with resource limits. Returns a result dict; never raises."""
    def limit_child() -> None:  # runs in the child between fork and exec
        resource.setrlimit(resource.RLIMIT_CPU, (cpu_s, cpu_s))
        resource.setrlimit(resource.RLIMIT_AS, (mem_mb * 2**20, mem_mb * 2**20))
        resource.setrlimit(resource.RLIMIT_FSIZE, (file_mb * 2**20, file_mb * 2**20))
        os.setsid()  # own process group, so a timeout kill takes children with it

    with tempfile.TemporaryDirectory(prefix="agent-sbx-") as tmp:
        started = time.perf_counter()
        try:
            proc = subprocess.run([sys.executable, "-I", "-S", "-c", code], cwd=tmp, env={"PATH": "/usr/bin:/bin"},
                                  capture_output=True, text=True, timeout=timeout_s, preexec_fn=limit_child)
            outcome = "ok" if proc.returncode == 0 else f"exit {proc.returncode}"
            if proc.returncode < 0:
                outcome = f"killed by signal {-proc.returncode}"
            stdout, stderr = proc.stdout, proc.stderr
        except subprocess.TimeoutExpired as exc:
            outcome = f"timeout after {timeout_s} s"
            stdout = (exc.stdout or b"").decode() if isinstance(exc.stdout, bytes) else (exc.stdout or "")
            stderr = ""
        ms = round((time.perf_counter() - started) * 1000)
    return {"outcome": outcome, "ms": ms, "stdout": stdout[:max_output],
            "stderr_tail": stderr.strip().splitlines()[-1][:160] if stderr.strip() else ""}


class Workspace:
    """File tools confined to one directory: the agent's scratch area for drafts and exports."""

    def __init__(self, root: str | Path, max_bytes: int = 100_000) -> None:
        self.root = Path(root).resolve()
        self.root.mkdir(parents=True, exist_ok=True)
        self.max_bytes = max_bytes

    def _path(self, relative: str) -> Path:
        p = (self.root / relative).resolve()  # resolve() follows symlinks and removes ..
        if not p.is_relative_to(self.root):
            raise PermissionError(f"{relative!r} is outside the workspace")
        return p

    def write_file(self, path: str, content: str) -> dict:
        if len(content.encode()) > self.max_bytes:
            raise ValueError(f"content over {self.max_bytes} bytes")
        p = self._path(path)
        p.parent.mkdir(parents=True, exist_ok=True)
        p.write_text(content, encoding="utf-8")
        return {"written": str(p.relative_to(self.root)), "bytes": len(content.encode())}

    def read_file(self, path: str) -> dict:
        p = self._path(path)
        return {"path": path, "content": p.read_text(encoding="utf-8")[: self.max_bytes]}

    def list_files(self) -> list[str]:
        return sorted(str(p.relative_to(self.root)) for p in self.root.rglob("*") if p.is_file())


if __name__ == "__main__":
    repo_file = str(Path("data/tickets.jsonl").resolve())
    cases = {
        "proration math": "seats, monthly, annual = 24, 12, 10\n"
                          "print('monthly per year:', seats * monthly * 12)\n"
                          "print('annual per year :', seats * annual * 12)\n"
                          "print('saving        :', seats * (monthly - annual) * 12)",
        "infinite loop": "while True:\n    pass",
        "memory bomb": "x = bytearray(1024 * 1024 * 1024)\nprint(len(x))",
        "disk filler": "open('big.bin', 'wb').write(b'0' * 5 * 1024 * 1024)",
        "reads outside cwd": f"print(open({repo_file!r}).readline()[:60])",
        "env and network": "import os, socket\nprint(sorted(os.environ))\n"
                           "s = socket.create_connection(('pypi.org', 443), timeout=3)\nprint('connected:', s.getpeername()[1])",
    }
    for name, code in cases.items():
        r = run_python(code)
        print(f"{name:18} {r['outcome']:22} {r['ms']:>5} ms | out: {r['stdout'].strip().replace(chr(10), '; ')[:84]!r} | err: {r['stderr_tail'][:70]!r}")

    ws = Workspace("runs/m08/workspace")
    print("\n", ws.write_file("drafts/T-1001.md", "Hi, we refunded the duplicate charge."))
    print(ws.list_files())
    for bad in ["../../supportdesk/llm.py", "/etc/passwd", "drafts/../../escape.txt"]:
        try:
            ws.read_file(bad)
        except PermissionError as exc:
            print("blocked:", exc)

Code explained

  • In simple words: run model-written Python in a child process with a stopwatch and a memory cap, and give the agent a folder it cannot climb out of.
  • What happens:
    • run_python starts python -I -S -c code (isolated mode, no site-packages) in a fresh temporary directory with an almost empty environment. limit_child runs in the child before it starts: RLIMIT_CPU caps CPU seconds, RLIMIT_AS caps address space (memory), RLIMIT_FSIZE caps the size of any file it writes, and setsid gives it its own process group. subprocess.run(timeout=...) caps wall time. Output is truncated to max_output characters so a chatty program cannot flood the context.
    • Workspace resolves every path (following symlinks and ..) and refuses anything outside its root. It also caps write size.
    • The demo runs six programs: a legitimate calculation, an infinite loop, a 1 GiB allocation, a 5 MB file write, a read of a file outside the sandbox directory, and a network connection.
  • Comes out: (real runs; timings vary by machine and run)
text
  proration math     ok                        68 ms | out: 'monthly per year: 3456; annual per year : 2880; saving        : 576' | err: ''
  infinite loop      timeout after 2.0 s     2028 ms | out: '' | err: ''
  memory bomb        exit 1                    49 ms | out: '' | err: 'MemoryError'
  disk filler        exit 1                    67 ms | out: '' | err: 'OSError: [Errno 27] File too large'
  reads outside cwd  ok                        54 ms | out: '{"id": "T-1001", "subject": "Charged twice this month", "bod' | err: ''
  env and network    ok                       107 ms | out: "['LC_CTYPE', 'PATH']; connected: 443" | err: ''

   {'written': 'drafts/T-1001.md', 'bytes': 37}
  ['drafts/T-1001.md']
  blocked: '../../supportdesk/llm.py' is outside the workspace
  blocked: '/etc/passwd' is outside the workspace
  blocked: 'drafts/../../escape.txt' is outside the workspace

The first four rows are the limits working: the answer comes back in tens of milliseconds, the loop dies at the 2 s timeout, the allocation fails with MemoryError under the 256 MB cap, and the write fails at the 1 MB file limit. The last two rows are the honest part. The child read a ticket file by absolute path, because nothing stopped it; and it opened a TCP connection to pypi.org on port 443, because resource limits do nothing about the network. Stripping environment variables hid the proxy settings but did not remove network access on this machine.

So this is a resource limiter, not a sandbox. It stops accidents (runaway loops, memory blowups). It does not stop a hostile program from reading secrets on disk or sending them somewhere. For code a model writes from untrusted input, run it where there is nothing to steal and nowhere to send it: a container or microVM with no credentials mounted, a read-only filesystem except a scratch directory, and networking off or restricted to an allowlist. Hosted code-execution tools from model providers and dedicated sandbox services exist for exactly this (see Other Tools).

SituationUse thisWhy
Model-written code over data you trust, on a dev machineSubprocess with resource limits (as above)Stops accidents cheaply
Model-written code in production, or any untrusted input in contextContainer or microVM, no secrets, no network, disposableResource limits do not stop exfiltration
A fixed calculation (proration, tax)A normal function toolDeterministic, testable, no code generation at all
The agent needs files (drafts, exports)A workspace-confined file tool with size limitsPath traversal is blocked in one place

Computer and browser use

Computer use means the model operates a graphical interface: it receives a screenshot, returns an action (click at x,y; type text; scroll; press keys), your code performs it, takes a new screenshot, and the loop continues. It is the same perceive, decide, act, observe loop, with pixels as observations. Browser use is the same idea restricted to a web browser, often with page structure (the accessibility tree or DOM) in addition to screenshots.

What is on offer as of September 2026 (check the docs, this moves quickly):

  • Anthropic's Claude API documents a computer-use toolset, computer_toolset_20260801, which its docs describe as generally available with no beta header, with actions such as screenshot, zoom, clicks, type, key, scroll, and wait.
  • Google's Gemini API documents computer use on several models, including gemini-3.5-flash (this course's Gemini default), in a loop where the model returns a function call with coordinates and your code executes it and returns a new screenshot. Responses can include a safety_decision that requires you to ask the end user for confirmation.
  • OpenAI's Responses API documents a computer tool with its own guidance on isolation and confirmation for consequential actions.

When to reach for it: only when there is no API. For Brightlane, the billing system has an API, so a computer-use agent clicking through the billing admin panel would be slower, costlier (every screenshot is image tokens, Module 12), and far less reliable than get_invoice. It earns its place for legacy back-office screens with no API, or for testing your own web UI.

The risks are the ones this module keeps returning to, amplified. Anything on screen is untrusted input: a web page can contain text that looks like instructions. Anthropic's docs recommend a dedicated virtual machine or container with minimal privileges, keeping sensitive data such as logins away from the model, limiting internet access to an allowlist of domains, and asking a human to confirm decisions with real-world consequences such as financial transactions. That is the approval gate from Part 5, applied to clicks.

File and workspace manipulation

Coding agents, report builders, and data agents live in a filesystem: they read inputs, write drafts, run code, and read results. The Workspace class above is the minimum safe version: one root directory, every path resolved and checked, writes capped. Two habits matter more than the code:

  • Give the agent its own workspace, never your home directory or the repository root. Our demo blocked ../../supportdesk/llm.py, /etc/passwd, and a traversal hidden inside a relative path.
  • Treat files the agent reads as untrusted content, like retrieved documents (Module 7). A file can carry injected instructions just as a web page can.

Subagents and delegation

A subagent is an agent started by another agent, with its own fresh context, its own tools, and a narrow task, which returns only a short result. In our code, a subagent is just run_agent called from inside a tool function, with a different system prompt and a subset of tools. Part 7 builds exactly that. Delegation buys three things: a clean context for the subtask (no leftover history competing for attention), a smaller tool list (fewer definitions, easier tool selection, Module 6), and the option to run subtasks in parallel. It costs a hand-off: whatever the subagent needs has to be written into its brief, and whatever it learned has to fit in its report.