CourseModel Context Protocol · Module 1: Foundations · part 3 of 83
Part 3 · Module 1: Foundations

Topic 3: The running project

13 min read·22 Sept 2026

B.1 Nare's notes folder

The data for the whole course is eight short Markdown notes in notes/. Each file is one note, named after its note id (a lowercase slug such as sleep-and-memory), with a small frontmatter header: a block between two --- lines holding the title, tags, and creation date.

bash
ls notes/
cat notes/lab-sync-2026-09-02.md

Code explained

  • In simple words: look inside the filing cabinet before building anything that searches it.
  • What happens: ls lists the eight note files; cat prints one of them, the meeting note from 2 September that most questions in this course will end up hitting.
  • Comes out:
    text
    coffee-and-focus.md
    deep-work.md
    lab-sync-2026-09-02.md
    pomodoro-technique.md
    reading-list.md
    sleep-and-memory.md
    spaced-repetition.md
    zettelkasten.md
    ---
    title: Lab sync, 2 September 2026
    tags: [meeting, sleep]
    created: 2026-09-02
    ---
    Attendees: Priya, Tomas, Nare.
    Decision: run the nap study with 24 participants instead of 16, budget approved by Priya.
    Action: Tomas books the sleep lab for the weeks of 5 and 12 October.
    Action: Nare drafts the consent form by 15 September.
    Risk: EEG caps are shared with another group, so booking conflicts are likely.

    Notice the answer to "how many participants?" is on line 2 of the body, not line 1. Keep that in mind; it matters in Part C.

B.2 NoteStore: business logic with no MCP in it

The first design decision of the project is also the most important: the code that knows about notes knows nothing about MCP. notes_assistant/store.py reads, searches, and writes notes. The MCP server you build later is a thin layer that calls it. That separation means the store can be unit tested without a protocol, reused by any script, and kept stable while the protocol changes underneath (and Part E shows it does change).

python
"""Business logic for the research notes assistant.

This module knows nothing about MCP. It reads and writes a folder of Markdown
notes, so it can be unit tested and reused by any server, host, or script.
"""
from __future__ import annotations

import re
from dataclasses import dataclass, field
from datetime import date
from pathlib import Path

NOTE_ID_PATTERN = re.compile(r"^[a-z0-9][a-z0-9-]{0,63}$")
WORD_PATTERN = re.compile(r"[a-z0-9]+")
MAX_TITLE_LENGTH = 120
MAX_BODY_LENGTH = 20_000


class NoteError(Exception):
    """Base class for every error the store raises on purpose."""


class NoteNotFound(NoteError):
    """Raised when a note id does not match any file."""


class NoteExists(NoteError):
    """Raised when creating a note whose id is already taken."""


class InvalidNote(NoteError):
    """Raised when a note id, title, body, or tag is not acceptable."""


@dataclass(frozen=True)
class Note:
    note_id: str
    title: str
    tags: tuple[str, ...]
    created: str
    body: str

    def snippet(self, length: int = 120) -> str:
        """Return the first line of the body, cut to `length` characters."""
        first_line = self.body.strip().splitlines()[0] if self.body.strip() else ""
        return first_line if len(first_line) <= length else first_line[: length - 3] + "..."


@dataclass(frozen=True)
class SearchHit:
    note_id: str
    title: str
    score: int
    snippet: str
    tags: tuple[str, ...] = field(default_factory=tuple)


def slugify(title: str) -> str:
    """Turn a title into a note id: lowercase words joined by hyphens."""
    words = WORD_PATTERN.findall(title.lower())
    return "-".join(words)[:64].strip("-")


def validate_note_id(note_id: str) -> str:
    """Reject ids that could escape the folder or are not simple slugs."""
    if not NOTE_ID_PATTERN.fullmatch(note_id):
        raise InvalidNote(
            f"Invalid note id {note_id!r}. Use lowercase letters, digits, and hyphens, "
            "for example 'sleep-and-memory'."
        )
    return note_id


def _parse(note_id: str, text: str) -> Note:
    """Split a Markdown file into frontmatter fields and body."""
    meta: dict[str, str] = {}
    body = text
    if text.startswith("---\n"):
        header, _, body = text[4:].partition("\n---\n")
        for line in header.splitlines():
            key, _, value = line.partition(":")
            meta[key.strip()] = value.strip()
    raw_tags = meta.get("tags", "").strip("[]")
    tags = tuple(t.strip() for t in raw_tags.split(",") if t.strip())
    return Note(
        note_id=note_id,
        title=meta.get("title", note_id),
        tags=tags,
        created=meta.get("created", ""),
        body=body.strip(),
    )


class NoteStore:
    """A folder of Markdown notes, one note per file, named <note_id>.md."""

    def __init__(self, root: Path | str) -> None:
        self.root = Path(root)
        if not self.root.is_dir():
            raise NoteError(f"Notes folder not found: {self.root}")

    def _path(self, note_id: str) -> Path:
        return self.root / f"{validate_note_id(note_id)}.md"

    def list_notes(self) -> list[Note]:
        """Every note, sorted by id so the order is deterministic."""
        return [self.get(p.stem) for p in sorted(self.root.glob("*.md"))]

    def get(self, note_id: str) -> Note:
        path = self._path(note_id)
        if not path.is_file():
            raise NoteNotFound(f"No note with id {note_id!r}.")
        return _parse(note_id, path.read_text(encoding="utf-8"))

    def search(self, query: str, limit: int = 5, tag: str | None = None) -> list[SearchHit]:
        """Rank notes by how often the query words appear (title words count 3 times)."""
        terms = WORD_PATTERN.findall(query.lower())
        if not terms:
            raise InvalidNote("Query must contain at least one letter or digit.")
        hits: list[SearchHit] = []
        for note in self.list_notes():
            if tag and tag not in note.tags:
                continue
            body_words = WORD_PATTERN.findall(note.body.lower())
            title_words = WORD_PATTERN.findall(note.title.lower())
            tag_words = [t.lower() for t in note.tags]
            score = sum(
                body_words.count(t) + 3 * title_words.count(t) + 2 * tag_words.count(t)
                for t in terms
            )
            if score > 0:
                hits.append(SearchHit(note.note_id, note.title, score, note.snippet(), note.tags))
        hits.sort(key=lambda h: (-h.score, h.note_id))
        return hits[:limit]

    def create(self, title: str, body: str, tags: list[str] | tuple[str, ...] = ()) -> Note:
        """Write a new note file. Never overwrites an existing note."""
        title = title.strip()
        if not title or len(title) > MAX_TITLE_LENGTH:
            raise InvalidNote(f"Title must be 1 to {MAX_TITLE_LENGTH} characters.")
        if not body.strip() or len(body) > MAX_BODY_LENGTH:
            raise InvalidNote(f"Body must be 1 to {MAX_BODY_LENGTH} characters.")
        clean_tags = tuple(slugify(t) for t in tags if slugify(t))
        note_id = slugify(title)
        if not note_id:
            raise InvalidNote("Title must contain at least one letter or digit.")
        path = self._path(note_id)
        if path.exists():
            raise NoteExists(f"A note with id {note_id!r} already exists.")
        header = (
            f"---\ntitle: {title}\ntags: [{', '.join(clean_tags)}]\n"
            f"created: {date.today().isoformat()}\n---\n"
        )
        path.write_text(header + body.strip() + "\n", encoding="utf-8")
        return self.get(note_id)

Code explained

  • In simple words: NoteStore is a librarian for one folder: it can list the shelf, fetch a note by its id, rank notes for a query, and file a new note without ever overwriting an old one.
  • What happens:
    • Constants. NOTE_ID_PATTERN allows 1 to 64 characters of lowercase letters, digits, and hyphens, starting with a letter or digit. WORD_PATTERN defines a "word" for search. MAX_TITLE_LENGTH and MAX_BODY_LENGTH cap input sizes so nobody can write a 50 MB note through a tool call.
    • NoteError, NoteNotFound, NoteExists, InvalidNote. A small exception family. Every error the store raises on purpose is a NoteError, so the MCP layer can later catch exactly these and turn them into messages the model can act on, while real bugs still surface as bugs.
    • Note. A frozen dataclass (immutable once built) holding one parsed note. snippet() returns the first line of the body, cut to 120 characters with ... when longer.
    • SearchHit. One search result: id, title, integer score, snippet, and tags. It carries only a snippet, not the whole body, which keeps search results small.
    • slugify(title). Turns "Nap study: consent form!" into nap-study-consent-form by lowercasing, keeping only letter and digit runs, and joining them with hyphens, capped at 64 characters.
    • validate_note_id(note_id). Rejects anything that is not a simple slug. This is the path traversal guard: ../secrets can never become a file path, because it fails the pattern before any file is touched.
    • _parse(note_id, text). Splits the file at the second --- line, reads key: value pairs from the header, parses the [a, b] tag list, and returns a Note. A file without frontmatter still parses, with the id as its title.
    • NoteStore.__init__. Stores the folder path and fails fast if it does not exist, so a wrong NOTES_DIR is caught at startup rather than on the first question.
    • NoteStore._path. The only place a note id becomes a file path, and it always validates first.
    • NoteStore.list_notes. Every note, sorted by id, so output order is the same on every machine.
    • NoteStore.get. Reads and parses one note or raises NoteNotFound.
    • NoteStore.search. Splits the query into words, then scores each note: every body occurrence counts 1, every title occurrence 3, every tag match 2. Optional tag filters notes first. Results are sorted by score (highest first) and then by id, so ties break predictably, and cut to limit. An empty query raises InvalidNote instead of returning everything. This is deliberately simple keyword scoring; it is transparent, which makes it perfect for teaching and for tests.
    • NoteStore.create. Validates title, body, and tags, derives the id with slugify, refuses to overwrite (NoteExists), writes the frontmatter and body, and returns the parsed note by reading it back.
  • Comes out: nothing yet; this file only defines things. The next example exercises every public method.

Here is the store in use. The script only reads from notes/, and it tries create() in a throwaway temporary folder so the sample notes stay untouched.

python
"""Module 1: a tour of NoteStore, the business logic every later module builds on.

Run from the repository root:  PYTHONPATH=. python examples/m01_store_tour.py
"""
from notes_assistant.store import InvalidNote, NoteNotFound, NoteStore

store = NoteStore("notes")

print("All notes:")
for note in store.list_notes():
    print(f"  {note.note_id:<24} {note.created}  tags={list(note.tags)}")

print("\nsearch('sleep memory', limit=3):")
for hit in store.search("sleep memory", limit=3):
    print(f"  score={hit.score:>2}  {hit.note_id:<24} {hit.snippet[:60]}")

print("\nsearch('sleep', tag='meeting'):")
for hit in store.search("sleep", tag="meeting"):
    print(f"  score={hit.score:>2}  {hit.note_id}")

note = store.get("lab-sync-2026-09-02")
print(f"\nget('lab-sync-2026-09-02').title = {note.title!r}")
print(f"body has {len(note.body.splitlines())} lines")

for bad in ["../secrets", "no-such-note"]:
    try:
        store.get(bad)
    except (InvalidNote, NoteNotFound) as exc:
        print(f"\nget({bad!r}) raised {type(exc).__name__}: {exc}")

print(f"\nsearch('quantum') returned {store.search('quantum')}")

# create() writes a file, so try it in a throwaway folder, never in notes/.
import tempfile  # noqa: E402

from notes_assistant.store import NoteExists  # noqa: E402

with tempfile.TemporaryDirectory() as scratch:
    scratch_store = NoteStore(scratch)
    created = scratch_store.create("Nap study consent form", "Draft due 15 September.", tags=["Meeting", "sleep"])
    print(f"\ncreated {created.note_id!r} with tags {list(created.tags)}")
    try:
        scratch_store.create("Nap study: consent form!", "Second try.")
    except NoteExists as exc:
        print(f"second create raised NoteExists: {exc}")

Code explained

  • In simple words: a test drive of the librarian, including two ways a request can go wrong and a search that finds nothing.
  • What happens: it lists every note with its date and tags; searches "sleep memory" and shows the top three scores; filters a search by the meeting tag; fetches one note by id; tries a path traversal id and a missing id; searches for a word no note contains; and finally creates a note in a temporary folder twice, with titles that slugify to the same id.
  • Comes out:
    text
    All notes:
      coffee-and-focus         2026-04-01  tags=['caffeine', 'focus', 'sleep']
      deep-work                2026-07-08  tags=['focus', 'productivity']
      lab-sync-2026-09-02      2026-09-02  tags=['meeting', 'sleep']
      pomodoro-technique       2026-04-15  tags=['focus', 'productivity']
      reading-list             2026-06-20  tags=['reading', 'sleep', 'memory']
      sleep-and-memory         2026-03-02  tags=['sleep', 'memory', 'neuroscience']
      spaced-repetition        2026-03-10  tags=['memory', 'learning']
      zettelkasten             2026-05-05  tags=['notes', 'productivity', 'learning']
    
    search('sleep memory', limit=3):
      score=15  sleep-and-memory         Slow-wave sleep replays the day's hippocampal activity and m
      score= 7  reading-list             Why We Sleep, Matthew Walker (2017): readable, but some clai
      score= 3  coffee-and-focus         Caffeine blocks adenosine receptors, so sleep pressure is ma
    
    search('sleep', tag='meeting'):
      score= 3  lab-sync-2026-09-02
    
    get('lab-sync-2026-09-02').title = 'Lab sync, 2 September 2026'
    body has 5 lines
    
    get('../secrets') raised InvalidNote: Invalid note id '../secrets'. Use lowercase letters, digits, and hyphens, for example 'sleep-and-memory'.
    
    get('no-such-note') raised NoteNotFound: No note with id 'no-such-note'.
    
    search('quantum') returned []
    
    created 'nap-study-consent-form' with tags ['meeting', 'sleep']
    second create raised NoteExists: A note with id 'nap-study-consent-form' already exists.

    Read the scores against the weights: sleep-and-memory scores 15 because "sleep" and "memory" appear in its title (3 each), its tags (2 each), and several times in its body. coffee-and-focus and spaced-repetition both score 3 for "sleep memory"; the tie is broken by id, so coffee-and-focus wins third place. The error messages are written for a reader who can fix the input (they say what a valid id looks like), which is exactly what a model will need when these errors reach it through MCP in later modules. Finally, tags=["Meeting", "sleep"] came back lowercased, because create() slugifies tags.

B.3 The chat() helper: one function, three providers

The host you build later needs a language model. The course supports three providers so everyone can use a free option: Groq (hosted, free tier, fast Llama models), Gemini (Google's hosted models, free tier), and Ollama (runs open models on your own machine, no key at all). All three expose an OpenAI-compatible Chat Completions endpoint, so one small helper talks to each of them through the openai client library.

python
"""One helper for three LLM providers: Groq, Gemini, and Ollama.

All three expose an OpenAI-compatible Chat Completions endpoint, so a single
client library (`openai`) talks to each of them. Pick a provider with the
LLM_PROVIDER environment variable and override the model with LLM_MODEL.
"""
from __future__ import annotations

import json
import os
from dataclasses import dataclass, field
from typing import Any

from openai import OpenAI

PROVIDERS: dict[str, dict[str, str]] = {
    "groq": {
        "base_url": "https://api.groq.com/openai/v1",
        "key_env": "GROQ_API_KEY",
        "default_model": "llama-3.3-70b-versatile",
    },
    "gemini": {
        "base_url": "https://generativelanguage.googleapis.com/v1beta/openai/",
        "key_env": "GEMINI_API_KEY",
        "default_model": "gemini-2.5-flash",
    },
    "ollama": {
        "base_url": "http://localhost:11434/v1",
        "key_env": "",
        "default_model": "qwen3:8b",
    },
}


@dataclass
class ToolCall:
    id: str
    name: str
    arguments: dict[str, Any]


@dataclass
class ChatReply:
    content: str | None
    tool_calls: list[ToolCall] = field(default_factory=list)

    def as_message(self) -> dict[str, Any]:
        """The assistant message to append to the conversation history."""
        message: dict[str, Any] = {"role": "assistant", "content": self.content or ""}
        if self.tool_calls:
            message["tool_calls"] = [
                {
                    "id": call.id,
                    "type": "function",
                    "function": {"name": call.name, "arguments": json.dumps(call.arguments)},
                }
                for call in self.tool_calls
            ]
        return message


def make_client(provider: str | None = None) -> tuple[OpenAI, str]:
    """Build an OpenAI-compatible client for the chosen provider and return it with the model name."""
    name = (provider or os.environ.get("LLM_PROVIDER", "groq")).lower()
    if name not in PROVIDERS:
        raise ValueError(f"Unknown LLM_PROVIDER {name!r}. Choose one of {sorted(PROVIDERS)}.")
    settings = PROVIDERS[name]
    api_key = os.environ.get(settings["key_env"], "") if settings["key_env"] else "ollama"
    if not api_key:
        raise RuntimeError(f"Set {settings['key_env']} in your environment to use {name}.")
    base_url = os.environ.get("OLLAMA_BASE_URL", settings["base_url"]) if name == "ollama" else settings["base_url"]
    model = os.environ.get("LLM_MODEL", settings["default_model"])
    return OpenAI(api_key=api_key, base_url=base_url), model


def chat(
    messages: list[dict[str, Any]],
    tools: list[dict[str, Any]] | None = None,
    provider: str | None = None,
    temperature: float = 0.0,
) -> ChatReply:
    """Send a conversation (and optional tool definitions) and return text or tool calls."""
    client, model = make_client(provider)
    kwargs: dict[str, Any] = {"model": model, "messages": messages, "temperature": temperature}
    if tools:
        kwargs["tools"] = tools
    response = client.chat.completions.create(**kwargs)
    message = response.choices[0].message
    calls = [
        ToolCall(id=c.id, name=c.function.name, arguments=json.loads(c.function.arguments or "{}"))
        for c in (message.tool_calls or [])
        if c.type == "function"
    ]
    return ChatReply(content=message.content, tool_calls=calls)

Code explained

  • In simple words: chat() is a universal power adapter: you plug in a conversation and a list of tools, and it hands back either text or a list of tool calls, whichever provider is on the other end.
  • What happens:
    • PROVIDERS. For each provider: the base URL of its OpenAI-compatible API, the environment variable holding the key, and a default model (llama-3.3-70b-versatile on Groq, gemini-2.5-flash on Gemini, qwen3:8b on Ollama). Keys never appear in code.
    • ToolCall. One request from the model to run a tool: the call id (needed to match the result back), the tool name, and the arguments already decoded from JSON into a dict.
    • ChatReply. The model's reply: optional text plus zero or more tool calls. as_message() converts it back into the assistant message format the API expects in the conversation history, re-encoding arguments as a JSON string, because the next request must include what the model said.
    • make_client(provider). Reads LLM_PROVIDER (default groq), rejects unknown names with a helpful list, finds the key (Ollama gets a dummy key because it needs none), honours OLLAMA_BASE_URL for a remote Ollama box, reads LLM_MODEL to override the default, and returns an OpenAI client pointed at the provider plus the model name. A missing key raises RuntimeError naming the exact variable to set.
    • chat(messages, tools, provider, temperature). Builds the request (temperature 0 by default so answers are as repeatable as the provider allows), adds tools only when there are some, sends it, and converts the first choice into a ChatReply, keeping only calls of type function.
  • Comes out: nothing on import. Without a key, calling chat() raises RuntimeError: Set GROQ_API_KEY in your environment to use groq., which you will see in Part C. To use a provider:
    SituationUse thisWhy
    Fastest start, hosted, free tierexport GROQ_API_KEY=... (default provider)Fast inference; the default model is good at tool calling
    You already have a Google accountexport LLM_PROVIDER=gemini GEMINI_API_KEY=...Free tier with generous limits
    No internet access, or data must stay on your machineexport LLM_PROVIDER=ollama after ollama pull qwen3:8bNothing leaves your machine; slower on a laptop

    No LLM key was available while writing this course, so any output that depends on a real model is marked as an illustrative sample run. Everything else you see is real output.