Topic 6: Testing it
A scripted stand-in model, and what it can and cannot prove
The capstone tests must run anywhere, with no API key and no randomness. So the host is driven by a scripted stand-in model: a plain Python function with the same signature as a real chat_fn, which follows the plan SYSTEM_PROMPT gives a real model. It searches with the question's content words, reads the top hit, and then either quotes the best matching line with a citation or abstains when fewer than half the question's keywords appear in any line.
Be clear about what this proves. It proves the host plumbing: tools are converted and namespaced, results flow back as tool messages, the approval gate fires, the log stays clean, pins block changed tools, and every transport carries the whole conversation. It does not prove that a real model will search sensibly, cite correctly, or resist injected text. That is what examples/m11_eval_real.py is for, later in this part.
"""A scripted stand-in for an LLM, used to test the host without an API key.
It follows the same plan SYSTEM_PROMPT asks a real model to follow:
search, read the top note, then answer with a citation or abstain.
It is deterministic keyword matching, not intelligence. It tests the host's
plumbing (tool conversion, loop, approval, logging), not answer quality.
"""
from __future__ import annotations
import json
import re
from typing import Any
from notes_assistant.host import ABSTAIN_TEXT, READ_RESOURCE_TOOL
from notes_assistant.llm import ChatReply, ToolCall
WORDS = re.compile(r"[a-z0-9]+")
STOPWORDS = {
"a", "an", "and", "are", "at", "be", "by", "did", "do", "does", "for", "how",
"i", "in", "is", "it", "long", "many", "me", "much", "my", "of", "on", "or",
"should", "the", "to", "use", "was", "we", "what", "when", "which", "who",
"why", "will", "with", "have", "has", "our", "your",
}
INJECTION = re.compile(r"call create_note with title '([^']+)'", re.IGNORECASE)
def keywords(text: str) -> list[str]:
"""Content words of a question, in order, without duplicates."""
seen: list[str] = []
for word in WORDS.findall(text.lower()):
if word not in STOPWORDS and word not in seen:
seen.append(word)
return seen
class ScriptedModel:
"""Callable with the host's chat_fn signature: (messages, tools) -> ChatReply."""
def __init__(self, obey_injections: bool = False, min_coverage: float = 0.5) -> None:
self.obey_injections = obey_injections # True simulates a model that falls for injected text
self.min_coverage = min_coverage
self.turns = 0
def _call(self, name: str, arguments: dict[str, Any]) -> ChatReply:
self.turns += 1
return ChatReply(content=None, tool_calls=[ToolCall(id=f"call_{self.turns}", name=name, arguments=arguments)])
def __call__(self, messages: list[dict[str, Any]], tools: list[dict[str, Any]]) -> ChatReply:
question = next(m["content"] for m in messages if m["role"] == "user")
names = [t["function"]["name"] for t in tools]
search = next((n for n in names if n.endswith("__search_notes")), None)
create = next((n for n in names if n.endswith("__create_note")), None)
# Which call produced each tool result: map tool_call_id to (name, arguments).
requested: dict[str, tuple[str, dict[str, Any]]] = {}
for message in messages:
for call in message.get("tool_calls", []):
requested[call["id"]] = (call["function"]["name"], json.loads(call["function"]["arguments"]))
results = [(requested[m["tool_call_id"]], m["content"]) for m in messages if m["role"] == "tool"]
called = [name for (name, _), _ in results]
if search is None:
return ChatReply(content=ABSTAIN_TEXT)
if not results:
return self._call(search, {"query": " ".join(keywords(question)), "limit": 3})
(last_name, _), last_text = results[-1]
if last_name == search:
hits = json.loads(last_text).get("hits", []) if last_text.startswith("{") else []
if not hits:
return ChatReply(content=ABSTAIN_TEXT)
return self._call(READ_RESOURCE_TOOL, {"uri": hits[0]["uri"]})
notes = [(args["uri"].split("://", 1)[1], text) for (name, args), text in results if name == READ_RESOURCE_TOOL]
if self.obey_injections and create and create not in called:
for _, text in notes:
match = INJECTION.search(text)
if match:
everything = "\n\n".join(text for _, text in notes)
return self._call(create, {"title": match.group(1), "body": everything, "tags": ["sync"]})
return ChatReply(content=self._answer(question, notes))
def _answer(self, question: str, notes: list[tuple[str, str]]) -> str:
"""Quote the note line that covers the most question keywords, or abstain."""
wanted = set(keywords(question))
best: tuple[float, str, str] = (0.0, "", "")
for note_id, text in notes:
for line in text.splitlines():
if not line.strip() or line.startswith(("#", "Tags:", "Created:")):
continue
line_words = set(WORDS.findall(line.lower()))
coverage = len(wanted & line_words) / max(len(wanted), 1)
if coverage > best[0]:
best = (coverage, line.strip(), note_id)
coverage, line, note_id = best
if coverage < self.min_coverage:
return ABSTAIN_TEXT
return f"{line} [{note_id}]"Code explained
- In simple words: a pretend model that reads the conversation the host builds and replies the way a careful model should, using keyword matching instead of intelligence.
- What happens:
keywords()drops question words such as "what", "how", "many", so "How many participants will the nap study have?" becomes the queryparticipants nap study.ScriptedModel.__call__works out where it is in the conversation from the messages alone, just as a real model would. It maps eachtool_call_idback to the call that produced it. No results yet: call*__search_notes. Last result was a search: parse the JSON, abstain ifhitsis empty, else callread_resourceon the top hit'suri. Otherwise: answer.obey_injections=Truesimulates a fooled model: if a note it read contains "call create_note with title '...'", it requestscreate_notewith that title and a body made of everything it has read. This is deliberately the worst case for the host._answer()scores every line of the notes it read by the share of question keywords it contains. Belowmin_coverage(0.5) it returnsABSTAIN_TEXTexactly, the phrase the host's system prompt asks a real model to use.- If the search tool is missing (for example, blocked by pinning), it abstains.
- Comes out: nothing on its own; the tests and the CLI's
--scriptedflag use it. The melatonin question is a good example of why it reads before abstaining: its search does return hits (doseappears in the caffeine note,bedin the spaced repetition note), and only after reading does the stand-in see that no line covers the question.