Topic 7: Module 1 lab
The lab combines the module in one script: the store, the N x M arithmetic, a measured comparison of the hand-written and generated tool schemas, the same server over three transports with timings and a consistency check, and one scripted model turn that runs its tool through MCP instead of a local dispatcher.
"""Module 1 lab: one script that exercises everything from the module.
1. Loads the notes with NoteStore.
2. Prints the N x M integration arithmetic.
3. Compares the hand-written native tool schema with the one MCP generates.
4. Reaches the same server in memory, over stdio, and over Streamable HTTP,
checks all three return the same hits, and times them.
5. Runs one scripted-model turn through MCP instead of a local dispatcher.
Run from the repository root: PYTHONPATH=. python examples/m01_lab.py
"""
import json
import logging
import statistics
import subprocess
import sys
import time
import anyio
import httpx2
from mcp import Client, StdioServerParameters
sys.path.insert(0, "examples")
from m01_first_server import mcp # noqa: E402
from m01_native_tools import SEARCH_TOOL, scripted_model # noqa: E402
from notes_assistant.llm import ChatReply # noqa: E402
from notes_assistant.store import NoteStore # noqa: E402
logging.getLogger().setLevel(logging.WARNING) # keep per-request HTTP logs out of the lab output
HTTP_PORT = 8011
HTTP_URL = f"http://127.0.0.1:{HTTP_PORT}/mcp"
CALLS = 50
QUERY = {"query": "nap study", "limit": 2}
def step1_store() -> None:
store = NoteStore("notes")
print(f"1. NoteStore: {len(store.list_notes())} notes, top hit for 'nap study' is {store.search('nap study')[0].note_id}")
def step2_arithmetic() -> None:
print("2. Integrations needed (hosts x tools versus hosts + tools):")
for hosts, tools in [(4, 6), (10, 50), (30, 400)]:
print(f" {hosts:>3} hosts, {tools:>3} tools: {hosts * tools:>6} bespoke adapters versus {hosts + tools:>4} MCP implementations")
def to_chat_tool(tool) -> dict:
"""Minimal MCP-to-Chat-Completions conversion (Module 6 builds the real one)."""
return {"type": "function", "function": {"name": tool.name, "description": tool.description, "parameters": tool.input_schema}}
async def step3_schema() -> None:
async with Client(mcp) as client:
generated = to_chat_tool((await client.list_tools()).tools[0])
for label, schema in [("hand-written native", SEARCH_TOOL), ("generated by MCP", generated)]:
chars = len(json.dumps(schema))
print(f"3. {label:<20} {chars} characters, about {chars // 4} tokens (estimate: characters / 4)")
async def time_target(label: str, target: object) -> list[str]:
start = time.perf_counter()
async with Client(target) as client:
connect_ms = 1000 * (time.perf_counter() - start)
first = await client.call_tool("search_notes", QUERY) # warm-up call, not timed
samples = []
for _ in range(CALLS):
t0 = time.perf_counter()
await client.call_tool("search_notes", QUERY)
samples.append(1000 * (time.perf_counter() - t0))
ids = [hit["note_id"] for hit in first.structured_content["result"]]
print(f" {label:<7} connect {connect_ms:7.1f} ms median call {statistics.median(samples):5.2f} ms hits {ids}")
return ids
async def wait_for_http() -> None:
for _ in range(50):
try:
async with httpx2.AsyncClient() as http:
await http.get(HTTP_URL)
return
except httpx2.TransportError:
await anyio.sleep(0.1)
raise RuntimeError("HTTP server did not start")
async def step4_transports() -> None:
print(f"4. Same server, three transports ({CALLS} timed calls each; timings vary run to run):")
stdio = StdioServerParameters(command=sys.executable, args=["examples/m01_first_server.py"], env={"PYTHONPATH": "."})
server = subprocess.Popen(
[sys.executable, "examples/m01_first_server.py", "--http", str(HTTP_PORT)],
env={"PYTHONPATH": "."}, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL,
)
try:
await wait_for_http()
results = [await time_target("memory", mcp), await time_target("stdio", stdio), await time_target("http", HTTP_URL)]
finally:
server.terminate()
server.wait()
print(f" all three transports agree: {results[0] == results[1] == results[2]}")
async def step5_scripted_turn() -> None:
async with Client(mcp) as client:
tools = [to_chat_tool(t) for t in (await client.list_tools()).tools]
messages = [{"role": "user", "content": "How many participants will the nap study have?"}]
reply: ChatReply = scripted_model(messages, tools)
call = reply.tool_calls[0]
result = await client.call_tool(call.name, call.arguments)
messages += [reply.as_message(), {"role": "tool", "tool_call_id": call.id, "content": json.dumps(result.structured_content["result"])}]
final = scripted_model(messages, tools)
print(f"5. Scripted stand-in model through MCP: {final.content}")
async def main() -> None:
step1_store()
step2_arithmetic()
await step3_schema()
await step4_transports()
await step5_scripted_turn()
if __name__ == "__main__":
anyio.run(main)Code explained
- In simple words: a single end-to-end rehearsal of everything in this module, with a stopwatch running.
- What happens:
step1_store()loads the notes and prints the top hit for "nap study" straight fromNoteStore, the baseline every transport must match.step2_arithmetic()prints N x M versus N + M for three team sizes.step3_schema()fetches the tool definition from the MCP server, converts it to the Chat Completions format with a five-lineto_chat_tool()(Module 6 builds the fullto_openai_toolwith namespacing), and compares its size with the hand-writtenSEARCH_TOOLfrom C.1. Tokens are estimated as characters divided by 4, which is a rough rule of thumb, not a tokenizer.step4_transports()starts the server over Streamable HTTP on port 8011 as a subprocess, waits until the port answers, then for each transport (in memory, stdio, HTTP) connects, makes one untimed warm-up call, times 50 calls, and records the median. It checks all three return the same note ids and always stops the HTTP subprocess infinally.step5_scripted_turn()repeats C.1's loop, but the tool list comes fromlist_tools()and the tool runs throughcall_tool(). The scripted stand-in model is the same one as in C.1; only the tool plumbing changed.
- Comes out: real output from
PYTHONPATH=. python examples/m01_lab.py(connect and call times vary run to run; the ranges over six runs are in the table after the output):text1. NoteStore: 8 notes, top hit for 'nap study' is lab-sync-2026-09-02 2. Integrations needed (hosts x tools versus hosts + tools): 4 hosts, 6 tools: 24 bespoke adapters versus 10 MCP implementations 10 hosts, 50 tools: 500 bespoke adapters versus 60 MCP implementations 30 hosts, 400 tools: 12000 bespoke adapters versus 430 MCP implementations 3. hand-written native 420 characters, about 105 tokens (estimate: characters / 4) 3. generated by MCP 372 characters, about 93 tokens (estimate: characters / 4) 4. Same server, three transports (50 timed calls each; timings vary run to run): memory connect 3.2 ms median call 1.07 ms hits ['lab-sync-2026-09-02', 'sleep-and-memory'] stdio connect 681.5 ms median call 2.01 ms hits ['lab-sync-2026-09-02', 'sleep-and-memory'] http connect 56.1 ms median call 3.55 ms hits ['lab-sync-2026-09-02', 'sleep-and-memory'] all three transports agree: True 5. Scripted stand-in model through MCP: (scripted) The answer should be in: lab-sync-2026-09-02, sleep-and-memory, spaced-repetitionHow to read the numbers honestly:
Measurement What it shows What it does not show Schema size, 420 versus 372 characters The generated schema is about 12 tokens smaller That MCP makes schemas smaller. The gap exists only because our hand-written version has parameter descriptions and the generated one does not yet. Add descriptions (Module 3) and they converge. For one tool, 12 tokens is noise; for 50 tools it adds up, which is Module 10's topic Median call over six runs: 0.95 to 1.3 ms in memory, 2.0 to 2.4 ms stdio, 3.5 to 5.7 ms HTTP Each process or network boundary adds a millisecond or two on one machine Remote latency. A real network adds tens of milliseconds, and a model call takes hundreds to thousands, so transport overhead is rarely what makes an assistant feel slow Connect over six runs: 1 to 11 ms memory, 54 to 110 ms HTTP, 680 to 820 ms stdio Launching a Python subprocess and importing the SDK dominates stdio start-up Steady-state cost. A host starts a stdio server once per session, not once per call All three transports agree The protocol, not the transport, defines behaviour Anything about answer quality; the model here is scripted The HTTP server is already running when its connect timer starts, so HTTP connect time is the discovery round trip plus client set-up, not server start-up. Fifty samples on one machine are enough to see the ordering (memory, then stdio, then HTTP) but not to trust differences under about half a millisecond, and the HTTP median moved by 2 ms between runs, which is noise. Rerun it a few times on your machine before quoting any number.