Topic 3: The running project
B.1 Nare's notes folder
The data for the whole course is eight short Markdown notes in notes/. Each file is one note, named after its note id (a lowercase slug such as sleep-and-memory), with a small frontmatter header: a block between two --- lines holding the title, tags, and creation date.
ls notes/
cat notes/lab-sync-2026-09-02.mdCode explained
- In simple words: look inside the filing cabinet before building anything that searches it.
- What happens:
lslists the eight note files;catprints one of them, the meeting note from 2 September that most questions in this course will end up hitting. - Comes out:text
coffee-and-focus.md deep-work.md lab-sync-2026-09-02.md pomodoro-technique.md reading-list.md sleep-and-memory.md spaced-repetition.md zettelkasten.md --- title: Lab sync, 2 September 2026 tags: [meeting, sleep] created: 2026-09-02 --- Attendees: Priya, Tomas, Nare. Decision: run the nap study with 24 participants instead of 16, budget approved by Priya. Action: Tomas books the sleep lab for the weeks of 5 and 12 October. Action: Nare drafts the consent form by 15 September. Risk: EEG caps are shared with another group, so booking conflicts are likely.Notice the answer to "how many participants?" is on line 2 of the body, not line 1. Keep that in mind; it matters in Part C.
B.2 NoteStore: business logic with no MCP in it
The first design decision of the project is also the most important: the code that knows about notes knows nothing about MCP. notes_assistant/store.py reads, searches, and writes notes. The MCP server you build later is a thin layer that calls it. That separation means the store can be unit tested without a protocol, reused by any script, and kept stable while the protocol changes underneath (and Part E shows it does change).
"""Business logic for the research notes assistant.
This module knows nothing about MCP. It reads and writes a folder of Markdown
notes, so it can be unit tested and reused by any server, host, or script.
"""
from __future__ import annotations
import re
from dataclasses import dataclass, field
from datetime import date
from pathlib import Path
NOTE_ID_PATTERN = re.compile(r"^[a-z0-9][a-z0-9-]{0,63}$")
WORD_PATTERN = re.compile(r"[a-z0-9]+")
MAX_TITLE_LENGTH = 120
MAX_BODY_LENGTH = 20_000
class NoteError(Exception):
"""Base class for every error the store raises on purpose."""
class NoteNotFound(NoteError):
"""Raised when a note id does not match any file."""
class NoteExists(NoteError):
"""Raised when creating a note whose id is already taken."""
class InvalidNote(NoteError):
"""Raised when a note id, title, body, or tag is not acceptable."""
@dataclass(frozen=True)
class Note:
note_id: str
title: str
tags: tuple[str, ...]
created: str
body: str
def snippet(self, length: int = 120) -> str:
"""Return the first line of the body, cut to `length` characters."""
first_line = self.body.strip().splitlines()[0] if self.body.strip() else ""
return first_line if len(first_line) <= length else first_line[: length - 3] + "..."
@dataclass(frozen=True)
class SearchHit:
note_id: str
title: str
score: int
snippet: str
tags: tuple[str, ...] = field(default_factory=tuple)
def slugify(title: str) -> str:
"""Turn a title into a note id: lowercase words joined by hyphens."""
words = WORD_PATTERN.findall(title.lower())
return "-".join(words)[:64].strip("-")
def validate_note_id(note_id: str) -> str:
"""Reject ids that could escape the folder or are not simple slugs."""
if not NOTE_ID_PATTERN.fullmatch(note_id):
raise InvalidNote(
f"Invalid note id {note_id!r}. Use lowercase letters, digits, and hyphens, "
"for example 'sleep-and-memory'."
)
return note_id
def _parse(note_id: str, text: str) -> Note:
"""Split a Markdown file into frontmatter fields and body."""
meta: dict[str, str] = {}
body = text
if text.startswith("---\n"):
header, _, body = text[4:].partition("\n---\n")
for line in header.splitlines():
key, _, value = line.partition(":")
meta[key.strip()] = value.strip()
raw_tags = meta.get("tags", "").strip("[]")
tags = tuple(t.strip() for t in raw_tags.split(",") if t.strip())
return Note(
note_id=note_id,
title=meta.get("title", note_id),
tags=tags,
created=meta.get("created", ""),
body=body.strip(),
)
class NoteStore:
"""A folder of Markdown notes, one note per file, named <note_id>.md."""
def __init__(self, root: Path | str) -> None:
self.root = Path(root)
if not self.root.is_dir():
raise NoteError(f"Notes folder not found: {self.root}")
def _path(self, note_id: str) -> Path:
return self.root / f"{validate_note_id(note_id)}.md"
def list_notes(self) -> list[Note]:
"""Every note, sorted by id so the order is deterministic."""
return [self.get(p.stem) for p in sorted(self.root.glob("*.md"))]
def get(self, note_id: str) -> Note:
path = self._path(note_id)
if not path.is_file():
raise NoteNotFound(f"No note with id {note_id!r}.")
return _parse(note_id, path.read_text(encoding="utf-8"))
def search(self, query: str, limit: int = 5, tag: str | None = None) -> list[SearchHit]:
"""Rank notes by how often the query words appear (title words count 3 times)."""
terms = WORD_PATTERN.findall(query.lower())
if not terms:
raise InvalidNote("Query must contain at least one letter or digit.")
hits: list[SearchHit] = []
for note in self.list_notes():
if tag and tag not in note.tags:
continue
body_words = WORD_PATTERN.findall(note.body.lower())
title_words = WORD_PATTERN.findall(note.title.lower())
tag_words = [t.lower() for t in note.tags]
score = sum(
body_words.count(t) + 3 * title_words.count(t) + 2 * tag_words.count(t)
for t in terms
)
if score > 0:
hits.append(SearchHit(note.note_id, note.title, score, note.snippet(), note.tags))
hits.sort(key=lambda h: (-h.score, h.note_id))
return hits[:limit]
def create(self, title: str, body: str, tags: list[str] | tuple[str, ...] = ()) -> Note:
"""Write a new note file. Never overwrites an existing note."""
title = title.strip()
if not title or len(title) > MAX_TITLE_LENGTH:
raise InvalidNote(f"Title must be 1 to {MAX_TITLE_LENGTH} characters.")
if not body.strip() or len(body) > MAX_BODY_LENGTH:
raise InvalidNote(f"Body must be 1 to {MAX_BODY_LENGTH} characters.")
clean_tags = tuple(slugify(t) for t in tags if slugify(t))
note_id = slugify(title)
if not note_id:
raise InvalidNote("Title must contain at least one letter or digit.")
path = self._path(note_id)
if path.exists():
raise NoteExists(f"A note with id {note_id!r} already exists.")
header = (
f"---\ntitle: {title}\ntags: [{', '.join(clean_tags)}]\n"
f"created: {date.today().isoformat()}\n---\n"
)
path.write_text(header + body.strip() + "\n", encoding="utf-8")
return self.get(note_id)Code explained
- In simple words:
NoteStoreis a librarian for one folder: it can list the shelf, fetch a note by its id, rank notes for a query, and file a new note without ever overwriting an old one. - What happens:
- Constants.
NOTE_ID_PATTERNallows 1 to 64 characters of lowercase letters, digits, and hyphens, starting with a letter or digit.WORD_PATTERNdefines a "word" for search.MAX_TITLE_LENGTHandMAX_BODY_LENGTHcap input sizes so nobody can write a 50 MB note through a tool call. NoteError,NoteNotFound,NoteExists,InvalidNote. A small exception family. Every error the store raises on purpose is aNoteError, so the MCP layer can later catch exactly these and turn them into messages the model can act on, while real bugs still surface as bugs.Note. A frozen dataclass (immutable once built) holding one parsed note.snippet()returns the first line of the body, cut to 120 characters with...when longer.SearchHit. One search result: id, title, integer score, snippet, and tags. It carries only a snippet, not the whole body, which keeps search results small.slugify(title). Turns "Nap study: consent form!" intonap-study-consent-formby lowercasing, keeping only letter and digit runs, and joining them with hyphens, capped at 64 characters.validate_note_id(note_id). Rejects anything that is not a simple slug. This is the path traversal guard:../secretscan never become a file path, because it fails the pattern before any file is touched._parse(note_id, text). Splits the file at the second---line, readskey: valuepairs from the header, parses the[a, b]tag list, and returns aNote. A file without frontmatter still parses, with the id as its title.NoteStore.__init__. Stores the folder path and fails fast if it does not exist, so a wrongNOTES_DIRis caught at startup rather than on the first question.NoteStore._path. The only place a note id becomes a file path, and it always validates first.NoteStore.list_notes. Every note, sorted by id, so output order is the same on every machine.NoteStore.get. Reads and parses one note or raisesNoteNotFound.NoteStore.search. Splits the query into words, then scores each note: every body occurrence counts 1, every title occurrence 3, every tag match 2. Optionaltagfilters notes first. Results are sorted by score (highest first) and then by id, so ties break predictably, and cut tolimit. An empty query raisesInvalidNoteinstead of returning everything. This is deliberately simple keyword scoring; it is transparent, which makes it perfect for teaching and for tests.NoteStore.create. Validates title, body, and tags, derives the id withslugify, refuses to overwrite (NoteExists), writes the frontmatter and body, and returns the parsed note by reading it back.
- Constants.
- Comes out: nothing yet; this file only defines things. The next example exercises every public method.
Here is the store in use. The script only reads from notes/, and it tries create() in a throwaway temporary folder so the sample notes stay untouched.
"""Module 1: a tour of NoteStore, the business logic every later module builds on.
Run from the repository root: PYTHONPATH=. python examples/m01_store_tour.py
"""
from notes_assistant.store import InvalidNote, NoteNotFound, NoteStore
store = NoteStore("notes")
print("All notes:")
for note in store.list_notes():
print(f" {note.note_id:<24} {note.created} tags={list(note.tags)}")
print("\nsearch('sleep memory', limit=3):")
for hit in store.search("sleep memory", limit=3):
print(f" score={hit.score:>2} {hit.note_id:<24} {hit.snippet[:60]}")
print("\nsearch('sleep', tag='meeting'):")
for hit in store.search("sleep", tag="meeting"):
print(f" score={hit.score:>2} {hit.note_id}")
note = store.get("lab-sync-2026-09-02")
print(f"\nget('lab-sync-2026-09-02').title = {note.title!r}")
print(f"body has {len(note.body.splitlines())} lines")
for bad in ["../secrets", "no-such-note"]:
try:
store.get(bad)
except (InvalidNote, NoteNotFound) as exc:
print(f"\nget({bad!r}) raised {type(exc).__name__}: {exc}")
print(f"\nsearch('quantum') returned {store.search('quantum')}")
# create() writes a file, so try it in a throwaway folder, never in notes/.
import tempfile # noqa: E402
from notes_assistant.store import NoteExists # noqa: E402
with tempfile.TemporaryDirectory() as scratch:
scratch_store = NoteStore(scratch)
created = scratch_store.create("Nap study consent form", "Draft due 15 September.", tags=["Meeting", "sleep"])
print(f"\ncreated {created.note_id!r} with tags {list(created.tags)}")
try:
scratch_store.create("Nap study: consent form!", "Second try.")
except NoteExists as exc:
print(f"second create raised NoteExists: {exc}")Code explained
- In simple words: a test drive of the librarian, including two ways a request can go wrong and a search that finds nothing.
- What happens: it lists every note with its date and tags; searches "sleep memory" and shows the top three scores; filters a search by the
meetingtag; fetches one note by id; tries a path traversal id and a missing id; searches for a word no note contains; and finally creates a note in a temporary folder twice, with titles that slugify to the same id. - Comes out:text
All notes: coffee-and-focus 2026-04-01 tags=['caffeine', 'focus', 'sleep'] deep-work 2026-07-08 tags=['focus', 'productivity'] lab-sync-2026-09-02 2026-09-02 tags=['meeting', 'sleep'] pomodoro-technique 2026-04-15 tags=['focus', 'productivity'] reading-list 2026-06-20 tags=['reading', 'sleep', 'memory'] sleep-and-memory 2026-03-02 tags=['sleep', 'memory', 'neuroscience'] spaced-repetition 2026-03-10 tags=['memory', 'learning'] zettelkasten 2026-05-05 tags=['notes', 'productivity', 'learning'] search('sleep memory', limit=3): score=15 sleep-and-memory Slow-wave sleep replays the day's hippocampal activity and m score= 7 reading-list Why We Sleep, Matthew Walker (2017): readable, but some clai score= 3 coffee-and-focus Caffeine blocks adenosine receptors, so sleep pressure is ma search('sleep', tag='meeting'): score= 3 lab-sync-2026-09-02 get('lab-sync-2026-09-02').title = 'Lab sync, 2 September 2026' body has 5 lines get('../secrets') raised InvalidNote: Invalid note id '../secrets'. Use lowercase letters, digits, and hyphens, for example 'sleep-and-memory'. get('no-such-note') raised NoteNotFound: No note with id 'no-such-note'. search('quantum') returned [] created 'nap-study-consent-form' with tags ['meeting', 'sleep'] second create raised NoteExists: A note with id 'nap-study-consent-form' already exists.Read the scores against the weights:
sleep-and-memoryscores 15 because "sleep" and "memory" appear in its title (3 each), its tags (2 each), and several times in its body.coffee-and-focusandspaced-repetitionboth score 3 for "sleep memory"; the tie is broken by id, socoffee-and-focuswins third place. The error messages are written for a reader who can fix the input (they say what a valid id looks like), which is exactly what a model will need when these errors reach it through MCP in later modules. Finally,tags=["Meeting", "sleep"]came back lowercased, becausecreate()slugifies tags.
B.3 The chat() helper: one function, three providers
The host you build later needs a language model. The course supports three providers so everyone can use a free option: Groq (hosted, free tier, fast Llama models), Gemini (Google's hosted models, free tier), and Ollama (runs open models on your own machine, no key at all). All three expose an OpenAI-compatible Chat Completions endpoint, so one small helper talks to each of them through the openai client library.
"""One helper for three LLM providers: Groq, Gemini, and Ollama.
All three expose an OpenAI-compatible Chat Completions endpoint, so a single
client library (`openai`) talks to each of them. Pick a provider with the
LLM_PROVIDER environment variable and override the model with LLM_MODEL.
"""
from __future__ import annotations
import json
import os
from dataclasses import dataclass, field
from typing import Any
from openai import OpenAI
PROVIDERS: dict[str, dict[str, str]] = {
"groq": {
"base_url": "https://api.groq.com/openai/v1",
"key_env": "GROQ_API_KEY",
"default_model": "llama-3.3-70b-versatile",
},
"gemini": {
"base_url": "https://generativelanguage.googleapis.com/v1beta/openai/",
"key_env": "GEMINI_API_KEY",
"default_model": "gemini-2.5-flash",
},
"ollama": {
"base_url": "http://localhost:11434/v1",
"key_env": "",
"default_model": "qwen3:8b",
},
}
@dataclass
class ToolCall:
id: str
name: str
arguments: dict[str, Any]
@dataclass
class ChatReply:
content: str | None
tool_calls: list[ToolCall] = field(default_factory=list)
def as_message(self) -> dict[str, Any]:
"""The assistant message to append to the conversation history."""
message: dict[str, Any] = {"role": "assistant", "content": self.content or ""}
if self.tool_calls:
message["tool_calls"] = [
{
"id": call.id,
"type": "function",
"function": {"name": call.name, "arguments": json.dumps(call.arguments)},
}
for call in self.tool_calls
]
return message
def make_client(provider: str | None = None) -> tuple[OpenAI, str]:
"""Build an OpenAI-compatible client for the chosen provider and return it with the model name."""
name = (provider or os.environ.get("LLM_PROVIDER", "groq")).lower()
if name not in PROVIDERS:
raise ValueError(f"Unknown LLM_PROVIDER {name!r}. Choose one of {sorted(PROVIDERS)}.")
settings = PROVIDERS[name]
api_key = os.environ.get(settings["key_env"], "") if settings["key_env"] else "ollama"
if not api_key:
raise RuntimeError(f"Set {settings['key_env']} in your environment to use {name}.")
base_url = os.environ.get("OLLAMA_BASE_URL", settings["base_url"]) if name == "ollama" else settings["base_url"]
model = os.environ.get("LLM_MODEL", settings["default_model"])
return OpenAI(api_key=api_key, base_url=base_url), model
def chat(
messages: list[dict[str, Any]],
tools: list[dict[str, Any]] | None = None,
provider: str | None = None,
temperature: float = 0.0,
) -> ChatReply:
"""Send a conversation (and optional tool definitions) and return text or tool calls."""
client, model = make_client(provider)
kwargs: dict[str, Any] = {"model": model, "messages": messages, "temperature": temperature}
if tools:
kwargs["tools"] = tools
response = client.chat.completions.create(**kwargs)
message = response.choices[0].message
calls = [
ToolCall(id=c.id, name=c.function.name, arguments=json.loads(c.function.arguments or "{}"))
for c in (message.tool_calls or [])
if c.type == "function"
]
return ChatReply(content=message.content, tool_calls=calls)Code explained
- In simple words:
chat()is a universal power adapter: you plug in a conversation and a list of tools, and it hands back either text or a list of tool calls, whichever provider is on the other end. - What happens:
PROVIDERS. For each provider: the base URL of its OpenAI-compatible API, the environment variable holding the key, and a default model (llama-3.3-70b-versatileon Groq,gemini-2.5-flashon Gemini,qwen3:8bon Ollama). Keys never appear in code.ToolCall. One request from the model to run a tool: the call id (needed to match the result back), the tool name, and the arguments already decoded from JSON into a dict.ChatReply. The model's reply: optional text plus zero or more tool calls.as_message()converts it back into the assistant message format the API expects in the conversation history, re-encoding arguments as a JSON string, because the next request must include what the model said.make_client(provider). ReadsLLM_PROVIDER(defaultgroq), rejects unknown names with a helpful list, finds the key (Ollama gets a dummy key because it needs none), honoursOLLAMA_BASE_URLfor a remote Ollama box, readsLLM_MODELto override the default, and returns anOpenAIclient pointed at the provider plus the model name. A missing key raisesRuntimeErrornaming the exact variable to set.chat(messages, tools, provider, temperature). Builds the request (temperature 0 by default so answers are as repeatable as the provider allows), addstoolsonly when there are some, sends it, and converts the first choice into aChatReply, keeping only calls of typefunction.
- Comes out: nothing on import. Without a key, calling
chat()raisesRuntimeError: Set GROQ_API_KEY in your environment to use groq., which you will see in Part C. To use a provider:Situation Use this Why Fastest start, hosted, free tier export GROQ_API_KEY=...(default provider)Fast inference; the default model is good at tool calling You already have a Google account export LLM_PROVIDER=gemini GEMINI_API_KEY=...Free tier with generous limits No internet access, or data must stay on your machine export LLM_PROVIDER=ollamaafterollama pull qwen3:8bNothing leaves your machine; slower on a laptop No LLM key was available while writing this course, so any output that depends on a real model is marked as an illustrative sample run. Everything else you see is real output.