CourseModel Context Protocol · Module 6: Clients and Hosts · part 35 of 83
Part 35 · Module 6: Clients and Hosts

Topic 2: Client integration

30 min read·22 Sept 2026

Connecting over both transports

The SDK has one Client class. You give it one argument and it picks the transport from the argument's type: a URL string means Streamable HTTP, a StdioServerParameters means "start this subprocess", and a server object means an in-memory connection (for tests). Everything after async with is identical across the three, which is exactly what a host wants.

python
"""Connect to the notes server three ways and time the same call on each."""
from __future__ import annotations

import time

import anyio

from mcp import Client

from m06_common import NOTES_HTTP_URL, ROOT, notes_params
from notes_assistant.server import build_server
from notes_assistant.store import NoteStore


async def probe(label: str, target) -> None:
    started = time.perf_counter()
    async with Client(target) as client:
        connect_ms = (time.perf_counter() - started) * 1000
        tools = await client.list_tools()
        timings = []
        for _ in range(20):
            t0 = time.perf_counter()
            result = await client.call_tool("search_notes", {"query": "sleep memory", "limit": 2})
            timings.append((time.perf_counter() - t0) * 1000)
        timings.sort()
        hits = [h["note_id"] for h in result.structured_content["hits"]]
        name = client.server_info.name if client.server_info else "?"
        print(f"{label:<10} server={name} protocol={client.protocol_version}")
        print(f"{'':<10} tools={[t.name for t in tools.tools]} hits={hits}")
        print(f"{'':<10} connect={connect_ms:.0f} ms  call median={timings[10]:.1f} ms  p95={timings[18]:.1f} ms")


async def main() -> None:
    await probe("in-memory", build_server(NoteStore(ROOT / "notes")))
    await probe("stdio", notes_params())
    await probe("http", NOTES_HTTP_URL)


if __name__ == "__main__":
    anyio.run(main)

Code explained

  • In simple words: dial the same server three different ways and compare how long the phone takes to ring and how fast each question is answered.
  • What happens:
    • probe() enters async with Client(target). Entering is when the connection happens: for stdio the SDK starts the subprocess, for HTTP it sends the first request, and in both cases it negotiates the protocol version. We time that as connect.
    • It lists tools, then calls search_notes 20 times and sorts the timings, so timings[10] is the median and timings[18] is roughly the 95th percentile.
    • client.server_info and client.protocol_version are plain properties once the block is entered. server_info can be None for a 2026-era server that does not identify itself, hence the guard.
    • result.structured_content is the tool's return value as JSON (the SearchResult model from Module 5), so we read note ids without parsing text.
  • Comes out (timings vary from run to run; three runs on the same machine gave stdio connects of 685 to 976 ms and HTTP call medians of 3.9 to 7.3 ms):
    text
    in-memory  server=notes protocol=2026-07-28
               tools=['search_notes', 'create_note'] hits=['sleep-and-memory', 'reading-list']
               connect=1 ms  call median=1.2 ms  p95=2.0 ms
    stdio      server=notes protocol=2026-07-28
               tools=['search_notes', 'create_note'] hits=['sleep-and-memory', 'reading-list']
               connect=976 ms  call median=2.6 ms  p95=9.9 ms
    http       server=notes protocol=2026-07-28
               tools=['search_notes', 'create_note'] hits=['sleep-and-memory', 'reading-list']
               connect=153 ms  call median=6.3 ms  p95=12.8 ms

    All three return the same tools and the same ranking, which is the point: the transport is invisible above the Client. The costs differ. stdio pays almost a second once, to start a Python process and import the SDK, then answers in about 2 to 3 ms because a pipe is cheap. HTTP connects faster here because the server is already running, but each call carries HTTP overhead, so its median call is a few milliseconds slower. In-memory is fastest and is for tests only. The differences between single calls (1 to 7 ms) are small next to one LLM round trip (typically hundreds of milliseconds or more), so choose the transport for deployment reasons, not speed.

SituationUse thisWhy
A server that runs on the user's machine and reads local filesClient(StdioServerParameters(...))The host controls the process lifetime and passes credentials through env; nothing listens on a port
A shared or remote server, or one several hosts use at onceClient("https://.../mcp")The server runs independently and can be deployed, scaled, and protected with auth (Module 7)
Unit and integration testsClient(build_server(store))No process, no port, about 1 ms per call, and the same protocol layer
HTTP with custom headers, proxies, or timeoutsClient(streamable_http_client(url, http_client=httpx2.AsyncClient(...)))Client(url) builds a default HTTP client; bring your own when you need to configure it

The rest of this course is yours to keep

This course is bought on its own, once, and stays readable afterwards, including the parts added to it later.