CourseLarge Language Models · Module 5: Reasoning and Advanced Prompting · part 26 of 80
Part 26 · Module 5: Reasoning and Advanced Prompting

Part E: Programmatic prompting

30 min read·22 Sept 2026

Prompts as versioned code, not string literals

Module 4 stored prompts as versioned text files with fingerprints. This part goes one step further: the prompt becomes a data structure that code can edit. A PromptVersion holds the parts of the triage prompt (category keyword definitions, a fallback) as fields, renders the text from them, and names each version by a hash of the rendered text. Once a prompt is a value, a program can propose edits, score them against the eval set, keep the good ones, and log every result against the exact fingerprint that produced it.

Prompt search needs a scorer that runs thousands of times. With a real model that is thousands of calls, so the offline run uses the deterministic keyword_reader as the "model": it follows whatever keyword definitions the prompt contains. The search results below are real for that reader; the loop, the logging, the overfitting, and the fix are the same with a real model.

examples/m05_prompt_search.py

python
"""Part E: prompts as versioned code, automatic prompt search against the dev set, and meta-prompting.

The "model" that reads each triage prompt is a deterministic stand-in
(keyword_reader, through ScriptedLLM): it follows the prompt's category
definitions literally. That makes the search loop run for real here, with real
dev and test scores for THIS reader. With a key, pass --live and the same loop
scores candidates with supportdesk.llm.chat instead (48 calls per candidate).
"""
from __future__ import annotations

import hashlib
import json
import sys
import time
from dataclasses import dataclass, field, replace
from pathlib import Path

from examples.m05_helpers import TRIAGE_KEYWORDS, keyword_responder, parse_category, triage_messages, triage_prompt, wilson
from supportdesk.data import CATEGORIES, Ticket, load_tickets
from supportdesk.kb_search import tokenize
from supportdesk.llm import Usage
from supportdesk.pricing import cost_usd
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages

RUN_LOG = Path("runs/m05_prompt_search.jsonl")


@dataclass(frozen=True)
class PromptVersion:
    """A prompt as data: its parts are fields, the text is rendered, and the fingerprint names it."""
    name: str
    keywords: tuple[tuple[str, tuple[str, ...]], ...]
    fallback: str = "how_to"
    parent: str | None = None
    note: str = ""

    @classmethod
    def from_dict(cls, name: str, keywords: dict[str, list[str]], **kw) -> "PromptVersion":
        return cls(name, tuple((c, tuple(ws)) for c, ws in keywords.items()), **kw)

    def as_dict(self) -> dict[str, list[str]]:
        return {c: list(ws) for c, ws in self.keywords}

    @property
    def text(self) -> str:
        return triage_prompt(self.as_dict(), self.fallback)

    @property
    def fingerprint(self) -> str:
        return hashlib.sha256(self.text.encode()).hexdigest()[:12]


def validate(version: PromptVersion) -> list[str]:
    """Problems that make a candidate unusable, whoever wrote it (a person, the search, or a model)."""
    problems = []
    names = [c for c, _ in version.keywords]
    missing = [c for c in CATEGORIES if c not in names]
    if missing:
        problems.append(f"missing categories {missing}")
    if extra := [c for c in names if c not in CATEGORIES]:
        problems.append(f"unknown categories {extra}")
    if version.fallback not in CATEGORIES:
        problems.append(f"fallback {version.fallback!r} is not a category")
    if len(version.text) > 2000:
        problems.append("prompt longer than 2,000 characters")
    return problems


@dataclass
class Score:
    correct: int
    n: int
    calls: int
    failures: list[tuple[Ticket, str | None]] = field(default_factory=list)

    @property
    def acc(self) -> float:
        return self.correct / self.n


def score(chat_fn, version: PromptVersion, tickets: list[Ticket]) -> Score:
    """Run the prompt over labelled tickets and count exact category matches."""
    s = Score(0, len(tickets), 0)
    for t in tickets:
        predicted = parse_category(chat_fn(triage_messages(version.text, t.text), temperature=0.0).text)
        s.calls += 1
        if predicted == t.gold["category"]:
            s.correct += 1
        else:
            s.failures.append((t, predicted))
    return s


def log_run(version: PromptVersion, split: str, s: Score) -> None:
    RUN_LOG.parent.mkdir(exist_ok=True)
    with RUN_LOG.open("a") as f:
        f.write(json.dumps({"prompt": version.name, "fingerprint": version.fingerprint, "parent": version.parent,
                            "split": split, "correct": s.correct, "n": s.n, "note": version.note}) + "\n")


# Proposers: turn a version and its failures into candidate versions -----------------------

def propose_edits(version: PromptVersion, failures: list[tuple[Ticket, str | None]], words_per_ticket: int = 3) -> list[PromptVersion]:
    """Deterministic proposer: for each failure, add a word from the ticket to the gold category,
    or remove the keyword that pulled it into the wrong category."""
    kw = version.as_dict()
    used = {w for ws in kw.values() for w in ws}
    seen, out = set(), []
    for ticket, predicted in failures:
        gold = ticket.gold["category"]
        words = [w for w in tokenize(ticket.text) if len(w) >= 4 and w.isascii() and w not in used]
        for w in list(dict.fromkeys(words))[:words_per_ticket]:
            if (gold, "+", w) not in seen:
                seen.add((gold, "+", w))
                new = {c: ws + [w] if c == gold else ws for c, ws in kw.items()}
                out.append(PromptVersion.from_dict("", new, fallback=version.fallback, note=f"add '{w}' to {gold}"))
        if predicted in kw:
            for w in kw[predicted]:
                if w in ticket.text.lower() and (predicted, "-", w) not in seen:
                    seen.add((predicted, "-", w))
                    new = {c: [x for x in ws if x != w] if c == predicted else ws for c, ws in kw.items()}
                    out.append(PromptVersion.from_dict("", new, fallback=version.fallback, note=f"remove '{w}' from {predicted}"))
    for c in CATEGORIES:
        if c != version.fallback:
            out.append(replace(version, fallback=c, note=f"fallback {c}"))
    return out


def search(chat_fn, start: PromptVersion, dev: list[Ticket], rounds: int = 8, min_gain: int = 1,
           holdout: list[Ticket] | None = None) -> tuple[PromptVersion, list[str], int]:
    """Greedy hill climbing: score every candidate on dev, keep the best if it gains at least min_gain.

    With a holdout set, each kept version is also scored there; the search stops if the holdout
    score drops, and returns the EARLIEST version with the best holdout score (the simplest prompt
    that the unseen tickets support), not the last one.
    """
    best, best_score = start, score(chat_fn, start, dev)
    calls, history = best_score.calls, [f"round 0  {best_score.correct}/{len(dev)}  {start.name} (start)"]
    held = score(chat_fn, start, holdout).correct if holdout else None
    kept = [(held, 0, start)]
    if holdout:
        calls += len(holdout)
        history[0] += f"  holdout {held}/{len(holdout)}"
    for r in range(1, rounds + 1):
        candidates = [c for c in propose_edits(best, best_score.failures) if not validate(c)]
        scored = []
        for c in candidates:
            s = score(chat_fn, c, dev)
            calls += s.calls
            scored.append((s.correct, -len(c.text), c, s))    # ties go to the shorter prompt
        top_correct, _, top, top_score = max(scored, key=lambda x: x[:2])
        if top_correct - best_score.correct < min_gain:
            history.append(f"round {r}  no candidate gains {min_gain}+ of {len(candidates)} tried; stop")
            break
        line = f"round {r}  {top_correct}/{len(dev)}  {len(candidates):>3} candidates  kept: {top.note}"
        if holdout:
            h = score(chat_fn, top, holdout).correct
            calls += len(holdout)
            line += f"  holdout {h}/{len(holdout)}"
            if h < held:
                history.append(line + "  holdout dropped; stop and keep the previous version")
                break
            held = h
        best = replace(top, name=f"{start.name}.s{r}", parent=best.name)
        best_score = top_score
        kept.append((held, r, best))
        history.append(line)
    if holdout:
        top_held = max(h for h, _, _ in kept)
        choice = min((r, v) for h, r, v in kept if h == top_held)
        history.append(f"returning round {choice[0]} ({choice[1].name}): earliest version with the best holdout score {top_held}/{len(holdout)}")
        return choice[1], history, calls
    return best, history, calls


# Meta-prompting: a model proposes the next version ----------------------------------------

META = ("You improve a ticket-classification prompt. Below is the current prompt and tickets it got wrong, "
        "with the correct category. Rewrite ONLY the category lines, one per category, in the same format "
        "'- category: word, word, ...'. Keep all six categories. Return the lines and nothing else.\n\n"
        "<current_prompt>\n{prompt}\n</current_prompt>\n<failures>\n{failures}\n</failures>")


def meta_messages(version: PromptVersion, failures: list[tuple[Ticket, str | None]], k: int = 8) -> list[dict]:
    lines = "\n".join(f"- ticket: {t.subject!r} | predicted: {p} | correct: {t.gold['category']}" for t, p in failures[:k])
    return [{"role": "user", "content": META.format(prompt=version.text, failures=lines)}]


def parse_meta(text: str, base: PromptVersion, name: str) -> PromptVersion:
    """Read '- category: words' lines from a model reply into a new version (then validate it)."""
    kw: dict[str, list[str]] = {}
    for line in text.splitlines():
        if line.startswith("- ") and ":" in line:
            cat, words = line[2:].split(":", 1)
            kw[cat.strip()] = [w.strip().lower() for w in words.split(",") if w.strip()]
    return PromptVersion.from_dict(name, kw, fallback=base.fallback, parent=base.name, note="meta-prompt proposal")


def scripted_meta(messages, kwargs):
    """ScriptedLLM stand-in for the meta-prompt call (plumbing only): it adds each failure's first
    subject word to the correct category, and, to show why validation exists, drops feature_request."""
    content = messages[-1]["content"]
    current = dict(line[2:].split(": ", 1) for line in content.splitlines() if line.startswith("- ") and ": " in line
                   and not line.startswith("- ticket"))
    for line in content.split("<failures>\n", 1)[1].splitlines():
        if line.startswith("- ticket"):
            subject = line.split("'")[1].lower()
            correct = line.rsplit("correct: ", 1)[1]
            current[correct] += ", " + subject.split()[0]
    current.pop("feature_request")
    return "\n".join(f"- {c}: {w}" for c, w in current.items())


if __name__ == "__main__":
    live = "--live" in sys.argv
    if live:
        from supportdesk.llm import chat as reader
    else:
        reader = ScriptedLLM(responder=keyword_responder)
    dev, test = load_tickets("dev"), load_tickets("test")
    v1 = PromptVersion.from_dict("triage-kw-v1", TRIAGE_KEYWORDS, note="hand-written keyword definitions")

    print(f"{v1.name} fingerprint {v1.fingerprint}, {len(v1.text)} characters, problems: {validate(v1) or 'none'}")
    started = time.perf_counter()
    best, history, calls = search(reader, v1, dev)
    print("\n".join(history))
    elapsed = time.perf_counter() - started
    tokens = sum(count_messages(c["messages"]) for c in getattr(reader, "calls", []))  # stand-in keeps its calls
    usd = cost_usd(Usage(input_tokens=tokens), "openai/gpt-oss-120b")
    print(f"search: {calls} reader calls, {tokens:,} input tokens (o200k estimate), "
          f"{usd:.2f} USD of input at gpt-oss-120b prices if they were real calls; {elapsed:.1f} s\n")

    # Same search, but it only sees 32 dev tickets and early-stops on the other 16.
    early, history2, calls2 = search(reader, replace(v1, name="triage-kw-v1e"), dev[:32], holdout=dev[32:])
    print("\n".join(history2))
    print(f"search with holdout: {calls2} reader calls\n")

    print(f"{'version':<16} {'fingerprint':<13} {'dev (n=48)':<22} {'test (n=24)':<22}")
    for v in dict.fromkeys((v1, best, early)):
        d, t = score(reader, v, dev), score(reader, v, test)
        log_run(v, "dev", d)
        log_run(v, "test", t)
        cells = [f"{s.correct}/{s.n} ({wilson(s.correct, s.n)[0]:.2f}-{wilson(s.correct, s.n)[1]:.2f})" for s in (d, t)]
        print(f"{v.name:<16} {v.fingerprint:<13} {cells[0]:<22} {cells[1]:<22}")
    added = {c: [w for w in ws if w not in dict(v1.as_dict()).get(c, [])] for c, ws in best.as_dict().items()}
    print("words the search added:", {c: ws for c, ws in added.items() if ws})

    # Meta-prompting: ask a model to write the next version; validate before scoring.
    meta_llm = reader if live else ScriptedLLM(responder=scripted_meta)
    failures = score(reader, v1, dev).failures
    proposal = parse_meta(meta_llm(meta_messages(v1, failures), temperature=0.7).text, v1, "triage-kw-meta1")
    problems = validate(proposal)
    print(f"\nmeta-prompt proposal {proposal.fingerprint}: problems {problems or 'none'}")
    if problems and not live:
        repaired = PromptVersion.from_dict("triage-kw-meta1r", {**proposal.as_dict(), "feature_request": list(dict(v1.keywords)["feature_request"])},
                                           parent=v1.name, note="meta proposal with the dropped category restored")
        proposal = repaired
        print(f"restored the dropped category from the parent: problems {validate(proposal) or 'none'}")
    if not validate(proposal):
        d, t = score(reader, proposal, dev), score(reader, proposal, test)
        print(f"{proposal.name:<16} dev {d.correct}/{d.n}  test {t.correct}/{t.n}")

Code explained

  • In simple words: treat the prompt as a value, let code try small edits to it, keep what scores better on dev, then check on test whether the gains were real.
  • What happens:
    • PromptVersion is a frozen dataclass: keywords and fallback are the editable parts, text renders them with the helpers' triage_prompt, and fingerprint is the first 12 hex digits of the SHA-256 of the text. Two versions with different names but the same text share a fingerprint, which is how you know a "new" version changed nothing.
    • validate rejects candidates that are missing a category, invent one, have a bad fallback, or are too long, whoever wrote them.
    • score runs a version over labelled tickets through any chat-shaped function and keeps the failures; log_run appends one JSON line per scored version and split to runs/m05_prompt_search.jsonl.
    • propose_edits is the deterministic proposer: for each failure, add one of the ticket's words to the correct category, remove the keyword that pulled the ticket into the wrong category, or change the fallback.
    • search is greedy hill climbing: score every valid candidate on dev, keep the best if it gains at least min_gain tickets (ties go to the shorter prompt), repeat for rounds. With holdout, it also scores each kept version on tickets the search never optimizes on, stops if that score drops, and returns the earliest version with the best holdout score.
    • META, meta_messages, parse_meta, and scripted_meta are the meta-prompting path, covered two sections down.
  • Comes out: nothing when imported (the tests import it); running it prints the search log and scores shown in the next section.

Automatic prompt optimization, and optimizing against an eval set

Automatic prompt optimization means a program, not a person, proposes prompt changes and keeps the ones that score better on a labelled set. Published systems use a model as the proposer: APE (Zhou et al., "Large Language Models Are Human-Level Prompt Engineers", 2022) generated instructions that matched or beat human-written ones on 19 of 24 tasks, and OPRO (Yang et al., "Large Language Models as Optimizers", ICLR 2024) fed scored past prompts back to a model and reported prompts beating human-designed ones by up to 8 percent on GSM8K and up to 50 percent on Big-Bench Hard tasks. DSPy (Khattab et al., 2023; version 3.3.1 on PyPI as of September 2026) turns this into a framework: you declare the steps of a pipeline and a metric, and an optimizer (its current optimizers include BootstrapFewShot, COPRO, MIPROv2, SIMBA, and GEPA) writes the instructions and picks the examples. The DSPy paper reports self-bootstrapped pipelines beating standard few-shot prompting by over 25 percent with GPT-3.5 and 65 percent with Llama-2-13b-chat.

All of them share the same danger: an optimizer maximizes the number you give it, and with 48 dev tickets, that number is noisy. Run the search:

bash
PYTHONPATH=. python examples/m05_prompt_search.py

Code explained

  • In simple words: run the search twice (on all of dev, then on part of dev with a holdout), score the results on test once, and try a meta-prompt proposal.
  • What happens: the first search optimizes on all 48 dev tickets for up to 8 rounds. The second optimizes on the first 32 dev tickets and watches the other 16. Then every resulting version is scored on dev and on the 24 test tickets, each result is logged with its fingerprint, and the meta-prompt path runs once.
  • Comes out:
    • The plain search climbs from 32/48 to 41/48 on dev in 8 rounds: 12,720 reader calls, about a second offline. On test the same version drops from 18/24 to 16/24. Look at the words it added: "team" for billing, "left", "data", and "legal" for account access, "teams" for feature requests. Most of them fix one particular dev ticket and say nothing about the category in general; "team" shows up in tickets of every category, and "passwort" (German) in the holdout run is the same story. The test drop of 2 is within noise, but the dev gain of 9 is clearly not a real improvement of that size: it is overfitting to 48 tickets.
    • The holdout search improves its 32 tickets from 22 to 28 while the 16 held-out tickets stay at 10/16 every round. No edit helped tickets it had not seen, so it returns round 0, the original prompt, with the same fingerprint a3e369c1e09e. That is the correct answer: the search found nothing that generalizes.
    • The meta-prompt proposal fails validation (it dropped a category) and is shown in the next section.

With a real model, the cost of search is the number to plan for. The plain search scored 12,720 prompt-and-ticket pairs; the script measures those prompts at 2,341,948 input tokens with the o200k counter, which would be about 0.35 USD of input on gpt-oss-120b, plus output and any hidden reasoning tokens, plus hours of wall-clock time at real latency. Real optimizers are more frugal (they propose a handful of candidates per round and score on minibatches), but the discipline is the same: optimize on one set, stop on a second, report on a third that the search never touched, and log fingerprints.

SituationUse thisWhy
Fewer than about 100 labelled examplesManual edits, one at a time, with Module 4's paired comparisonSearch will overfit a small set, as measured here
Hundreds of labelled examples and a clear metricAutomatic search with a holdout for early stopping, final test onceThe search finds edits a person would not try; the holdout keeps it honest
A multi-step pipeline with a metric at the endDSPy or a similar frameworkIt optimizes instructions and examples for every step jointly
No labelled dataDo not optimize; build the eval set first (Module 4, Part B)Without a metric, an optimizer optimizes nothing you care about

Meta-prompting: using a model to write prompts

Meta-prompting here means asking a model to write or improve a prompt: you give it the current prompt and examples of its failures, and it returns a revised version. (The term is also used for Suzgun and Kalai's "Meta-Prompting" (2024), where one model orchestrates copies of itself as experts; that is closer to Module 8's agents.) Provider consoles offer the same idea as prompt generators and improvers.

A model-written prompt is a candidate, not a result. It goes through the same gate as any other candidate: validate, score on dev, confirm on holdout. The script's meta-prompt path does exactly that: meta_messages builds a request with the current prompt and up to eight failures, parse_meta reads the reply back into a PromptVersion, and validate checks it before anything is scored. Offline, scripted_meta stands in for the model (plumbing only). It adds each failure's first subject word to the right category and, to show why the gate exists, drops the feature_request line, which is a realistic failure: models rewriting a list often lose an item.

In the output above, the proposal fails validation with missing categories ['feature_request']. After restoring the dropped category from the parent, it scores 36/48 on dev and 17/24 on test, against the original's 32/48 and 18/24: another dev gain that does not show up on test. With a key, --live sends the same meta-prompt to the real model; expect it to propose sensible synonyms, some over-specific words, and occasionally a dropped or renamed category. The validator and the holdout, not your impression of the prompt, decide whether it ships.

SituationUse thisWhy
Writing a first draft for a new taskMeta-prompt to draft, then edit by handFast start; a model knows common prompt structure
Fixing specific failuresMeta-prompt with the failures, then validate and scoreThe failures give the model information to act on
Unattended optimization loopsMeta-prompt as the proposer inside a search with holdout early stoppingThat is OPRO's design; the holdout guards against overfitting
Security or policy wordingA person writes itModel rewrites can drop constraints silently

Module Lab

The lab runs the module's pieces as one pipeline over one week: route every ticket, decide the refund scenarios with best-of-n and the billing-record verifier (escalating when nothing verifies), map every ticket to a category and count in code, and write a run record that names the prompt fingerprints behind the numbers. Offline, every model call goes to a ScriptedLLM stand-in; the refund sampler follows the simulated error model from the best-of-n section, and the map step uses the keyword reader, so the routing and map numbers are real for those rule stand-ins and the refund accuracy is the simulation's.

examples/m05_lab.py:

python
"""Module 5 lab: one pipeline over the week's tickets and the 20 refund scenarios.

1. Route all 72 tickets (router prompt, conditional paths).
2. Refund decisions: sample up to N answers with the scaffold prompt, accept the first one the
   billing-record verifier confirms, escalate if none (best-of-n with early stopping).
3. Map-reduce the week into report counts (map: one triage call per ticket; reduce: code counts).
4. Print calls, tokens, dollars, and the prompt fingerprints that produced the run.

Offline, every model call goes to a ScriptedLLM stand-in (plumbing only): the refund sampler follows
the simulated error model from m05_best_of_n.py, so its accuracy is the simulation's, not a model's.
With a key:  LLM_PROVIDER=groq GROQ_API_KEY=... PYTHONPATH=. python examples/m05_lab.py --live
"""
import hashlib
import json
import random
import sys
from collections import Counter

from examples.m05_helpers import (CASES, LABELS, SCAFFOLD, TRIAGE_KEYWORDS, case_from_messages, decide,
                                  keyword_responder, parse_category, parse_decision, prompt_scaffold,
                                  shortcut_baseline, triage_messages, triage_prompt, wilson)
from examples.m05_router import ROUTER, gold_route, route, rule_router
from supportdesk.data import load_tickets
from supportdesk.llm import Usage
from supportdesk.pricing import cost_usd
from supportdesk.stand_in import ScriptedLLM

MODEL = "openai/gpt-oss-120b"
MAX_SAMPLES = 5
rng = random.Random(7)


def simulated_sampler(messages, kwargs):
    """Stand-in: right with p=0.9 on easy cases, 0.45 on trap cases; traps usually give the shortcut answer."""
    case = case_from_messages(messages)
    trap = shortcut_baseline(case.text) != case.gold
    if rng.random() < (0.45 if trap else 0.9):
        label = case.gold
    elif trap and rng.random() < 0.8:
        label = shortcut_baseline(case.text)
    else:
        label = rng.choice([x for x in LABELS if x != case.gold])
    return f"(scripted stand-in reasoning)\nDECISION: {label}"


def verified_decision(chat_fn, case, usage: Usage) -> tuple[str, int]:
    """Sample until the billing-record check accepts an answer; escalate after MAX_SAMPLES."""
    for i in range(1, MAX_SAMPLES + 1):
        reply = chat_fn(prompt_scaffold(case), temperature=0.8, max_tokens=2000)
        usage.input_tokens += reply.usage.input_tokens
        usage.output_tokens += reply.usage.output_tokens
        label = parse_decision(reply.text)
        if label is not None and label == decide(case.facts):   # the external check
            return label, i
    return "escalate", MAX_SAMPLES


def fingerprint(text: str) -> str:
    return hashlib.sha256(text.encode()).hexdigest()[:12]


if __name__ == "__main__":
    live = "--live" in sys.argv
    if live:
        from supportdesk.llm import chat
        router_llm = sampler = chat
    else:
        router_llm, sampler = ScriptedLLM(responder=rule_router), ScriptedLLM(responder=simulated_sampler)
    print(f"prompts: router {fingerprint(ROUTER)}  refund scaffold {fingerprint(SCAFFOLD)}  "
          f"({'live model' if live else 'scripted stand-ins'})\n")

    # 1. Route the week.
    tickets = load_tickets()
    routes = [route(router_llm, t) for t in tickets]
    agree = sum(r == gold_route(t) for r, t in zip(routes, tickets))
    print(f"1. routing: {dict(Counter(routes))}; agrees with hand labels on {agree}/{len(tickets)}")

    # 2. Refund decisions with a verifier.
    usage = Usage()
    outcomes = [verified_decision(sampler, case, usage) for case in CASES]
    decided = [(case, label) for case, (label, _) in zip(CASES, outcomes) if label != "escalate"]
    calls = sum(n for _, n in outcomes)
    wrong = sum(label != case.gold for case, label in decided)
    lo, hi = wilson(len(decided), len(CASES))
    print(f"2. refunds: decided {len(decided)}/{len(CASES)} (95% CI {lo:.2f}-{hi:.2f}), wrong among decided {wrong}, "
          f"escalated {len(CASES) - len(decided)}; {calls} calls ({calls / len(CASES):.2f} per case)")
    print(f"   tokens in/out {usage.input_tokens}/{usage.output_tokens}; "
          f"USD {cost_usd(usage, MODEL):.4f} at {MODEL} prices")

    # 3. Map step: classify every ticket (keyword stand-in offline); reduce step: count in code.
    mapper = chat if live else ScriptedLLM(responder=keyword_responder)
    triage_text = triage_prompt(TRIAGE_KEYWORDS)
    predicted = [parse_category(mapper(triage_messages(triage_text, t.text), temperature=0.0).text) for t in tickets]
    counts = Counter(predicted)
    gold_counts = Counter(t.gold["category"] for t in tickets)
    right = sum(p == t.gold["category"] for p, t in zip(predicted, tickets))
    print(f"3. week: {len(tickets)} map calls; predicted {dict(counts.most_common())}")
    print(f"   gold      {dict(gold_counts.most_common())}; per-ticket agreement {right}/{len(tickets)}")

    # 4. Run record.
    record = {"router": fingerprint(ROUTER), "scaffold": fingerprint(SCAFFOLD), "routing_agree": agree,
              "refund_decided": len(decided), "refund_wrong": wrong, "refund_calls": calls, "map_agree": right,
              "usd": round(cost_usd(usage, MODEL), 6), "live": live}
    print("4. run record:", json.dumps(record))

Code explained

  • In simple words: the whole module in one run: route, decide with a check, summarize the week, and record what produced the numbers.
  • What happens:
    • The header prints the fingerprints of the router prompt and the refund scaffold, so the run record can be tied to exact prompt text (Part E's idea, applied to the pipeline).
    • Step 1 reuses route and rule_router from the router example, with the safe fallback to human.
    • Step 2's verified_decision samples the scaffold prompt at temperature 0.8 up to MAX_SAMPLES = 5 times and accepts the first answer that decide confirms; if none passes, the case is escalated. Tokens are summed from each call's usage and priced with cost_usd.
    • Step 3 is map-reduce: one triage call per ticket (the keyword reader offline), counted with Counter, compared with the gold counts.
    • Step 4 prints a JSON run record you can append to runs/.
    • --live swaps every stand-in for supportdesk.llm.chat.
  • Comes out: routing matches the router example (67/72 agreement). All 20 refund scenarios are decided with no wrong decisions, using 31 calls (1.55 per case) and about 0.0016 USD at gpt-oss-120b prices; wrong answers are impossible here because only verified answers are accepted, and the escalation count is the number to watch with a real model. The map step's category counts differ visibly from gold (29 how_to predicted vs 18 real) even though per-ticket agreement is 50/72: counting errors do not cancel out, so a weekly report built on a weak map step misstates the trends. Improve the map step (Part E's search, a real model, or Module 6's structured output) before trusting its counts.

To run it with a real model: LLM_PROVIDER=groq GROQ_API_KEY=... PYTHONPATH=. python examples/m05_lab.py --live. Compare three things against the offline run: the escalation count, the calls per refund case, and the map agreement. If escalations are high, read which step failed (the scaffold's ROLE or DAYS_SINCE lines usually tell you) before raising MAX_SAMPLES.

Project Milestone

After this module, the Brightlane repository contains:

  • examples/m05_helpers.py: the refund policy as a rule function, 20 labelled scenarios in four languages, four prompt styles, parse_decision, evaluate with Wilson intervals, and two classical baselines.
  • examples/m05_cot_prompts.py, m05_effort_dial.py: prompt-style and reasoning-effort harnesses with --live modes and a cost curve.
  • examples/m05_voting_math.py, m05_tinylm_vote.py, m05_best_of_n.py, m05_ensemble.py: voting math, a real TinyLM vote, verifier vs scorer, and a measured ensemble.
  • examples/m05_least_to_most.py, m05_chain.py, m05_router.py, m05_map_reduce.py: decomposition with typed hand-offs, conditional routing, and a counted map-reduce.
  • examples/m05_self_correct.py: a critique loop with a ceiling and a verifier gate.
  • examples/m05_prompt_search.py: PromptVersion, validation, hill-climbing search with holdout early stopping, and a meta-prompt proposer; results logged to runs/m05_prompt_search.jsonl.
  • examples/m05_lab.py and tests/test_m05_reasoning.py.

The tests pin down the behavior that the rest of the course relies on:

tests/test_m05_reasoning.py:

python
"""Tests for Module 5: the refund rule, parsers, voting math, chains, routing, self-correction, prompt search."""
import random
from collections import Counter
from datetime import date

import pytest
from pydantic import ValidationError

from examples.m05_chain import RefundFacts, run_chain
from examples.m05_helpers import (CASES, Facts, PROMPTS, TODAY, case_from_messages, decide, evaluate,
                                  parse_decision, wilson)
from examples.m05_least_to_most import least_to_most, settled, stand_in
from examples.m05_prompt_search import PromptVersion, score, search, validate
from examples.m05_router import route
from examples.m05_self_correct import critique_and_revise, make_stand_in
from examples.m05_voting_math import majority_exact
from examples.m05_helpers import TRIAGE_KEYWORDS, keyword_responder
from supportdesk.data import load_tickets
from supportdesk.stand_in import ScriptedLLM


def test_rule_order_and_14_day_boundary():
    assert decide(Facts("member", "annual", TODAY, True)) == "needs_owner"       # role first
    assert decide(Facts("owner", "annual", date(2026, 1, 1), True)) == "refund_duplicate"
    assert decide(Facts("owner", "annual", date(2026, 9, 7), False)) == "refund_full"    # 14 days: inside
    assert decide(Facts("owner", "annual", date(2026, 9, 6), False)) == "no_refund"      # 15 days: outside
    assert decide(Facts("billing_admin", "monthly", TODAY, False)) == "no_refund"


def test_gold_distribution_is_stable():
    assert len(CASES) == 20
    assert Counter(c.gold for c in CASES) == {"no_refund": 8, "refund_full": 6, "refund_duplicate": 3, "needs_owner": 3}


def test_parse_decision_takes_last_known_label():
    assert parse_decision("DECISION: no_refund\n...\nDECISION: refund_full") == "refund_full"
    assert parse_decision("**DECISION:** needs_owner") == "needs_owner"
    assert parse_decision("Decision - Refund Full") is None
    assert parse_decision("DECISION: maybe") is None


def test_wilson_contains_point_estimate():
    lo, hi = wilson(14, 20)
    assert lo < 0.7 < hi and 0 <= lo and hi <= 1


def test_majority_math():
    assert majority_exact(0.7, 5) == pytest.approx(0.83692, abs=1e-5)
    assert all(majority_exact(0.5, n) == pytest.approx(0.5) for n in (1, 2, 3, 8, 15))
    assert majority_exact(0.3, 9) < 0.3  # voting amplifies a minority-correct model's errors


def test_harness_with_oracle_scores_all():
    def oracle(messages, kwargs):
        return f"DECISION: {decide(case_from_messages(messages).facts)}"
    result = evaluate(ScriptedLLM(responder=oracle), PROMPTS["direct"])
    assert (result.correct, result.unparsed) == (20, 0)


def test_chain_escalates_on_missing_facts():
    null = ScriptedLLM(replies=['{"role": null, "billing": null, "charged_on": null, "duplicate": null}'])
    trace = run_chain(CASES[0], null, ScriptedLLM(replies=[]))
    assert trace["outcome"] == "escalate" and trace["stage"] == "extract" and trace["calls"] == 1
    with pytest.raises(ValidationError):
        RefundFacts.model_validate_json('{"role": "owner", "billing": "annual", "charged_on": "3 September", "duplicate": false}')


def test_router_sends_unexpected_output_to_a_human():
    ticket = load_tickets()[0]
    assert route(ScriptedLLM(replies=['{"route": "refund_everything"}']), ticket) == "human"
    assert route(ScriptedLLM(replies=["not json"]), ticket) == "human"


def test_least_to_most_stops_once_settled():
    llm = ScriptedLLM(responder=stand_in)
    solved, calls = least_to_most(llm, CASES[3].text, settled)   # R04: a member, so question 1 settles it
    assert calls == 2 and solved[0][1] == "no"


def test_verifier_never_lets_a_correct_answer_become_wrong():
    llm = make_stand_in(random.Random(0))
    for _ in range(20):
        for case in CASES:
            history, _ = critique_and_revise(llm, case, 4, use_verifier=True)
            first_right = next((i for i, h in enumerate(history) if h == case.gold), None)
            if first_right is not None:
                assert all(h == case.gold for h in history[first_right:])


def test_without_verifier_correct_answers_can_flip():
    llm = make_stand_in(random.Random(0))
    flips = 0
    for case in CASES * 10:
        history, _ = critique_and_revise(llm, case, 1, use_verifier=False)
        flips += history[0] == case.gold and history[1] != case.gold
    assert flips > 0


def test_prompt_versions_and_search():
    v1 = PromptVersion.from_dict("v1", TRIAGE_KEYWORDS)
    assert v1.fingerprint == PromptVersion.from_dict("renamed", TRIAGE_KEYWORDS).fingerprint  # text, not name
    broken = PromptVersion.from_dict("bad", {k: v for k, v in TRIAGE_KEYWORDS.items() if k != "bug"})
    assert validate(broken) and not validate(v1)
    reader = ScriptedLLM(responder=keyword_responder)
    dev = load_tickets("dev")
    best, _, _ = search(reader, v1, dev, rounds=2)
    assert score(reader, best, dev).correct >= score(reader, v1, dev).correct + 2

Code explained

  • In simple words: twelve fast checks that the rule, the parsers, the math, and the loop guards behave as the module claims.
  • What happens: the rule tests pin the order of checks and the 14-day boundary (day 14 inside, day 15 outside). The parser tests pin "last known label wins" and "unknown means unparsed". The voting test pins the binomial value 0.83692 and the p = 0.5 and p = 0.3 behavior. The chain, router, and least-to-most tests check the safety paths: escalate on missing facts, route unknown output to a person, stop once settled. The two self-correction tests are the module's central claim in executable form: with the verifier a right answer never becomes wrong; without it, it does. The last test checks that fingerprints depend on text, not names, that validation catches a missing category, and that two rounds of search gain at least two dev tickets.
  • Comes out: run PYTHONPATH=. python -m pytest -q tests/test_m05_reasoning.py:

The honest state of the project: the refund chain decides correctly whenever its facts validate, and every other number in this module is either a classical baseline, TinyLM, a simulation with stated assumptions, or ScriptedLLM plumbing. Your first tasks with a key are the three --live runs: m05_cot_prompts.py (which prompt style works for your model), m05_effort_dial.py (what effort level is worth paying for), and m05_lab.py (escalation rate and calls per case). Maya's question for the next review: what share of refund requests can the chain decide with a verified answer, and what does each cost?

Interview Questions

1. When does chain-of-thought help, and when is it a waste? It helps when the answer depends on several facts combined by rules, dates, or arithmetic, because each written step becomes context for the next and the model gets more computation before committing. Sprague et al. (ICLR 2025) found the gains concentrated on math and logic, with little effect elsewhere. For recognition tasks like ticket category it mostly adds output tokens: in this module, free-form reasoning was priced at about 3.6 times a direct answer on gpt-oss-120b. Measure on your dev set, and on a reasoning model do not add "think step by step" at all.

2. Zero-shot CoT, demonstrated CoT, or a structured scaffold: how do you choose? Zero-shot ("think step by step") is the cheapest to write and good for exploring a task. Demonstrations teach a specific order of checks at a fixed input cost per call (138 extra tokens here). A scaffold asks for named fields, which is the cheapest reasoning to generate and the only one where each step is machine-checkable against a record. For a production decision like refunds, I would use a scaffold and compare its fields to the billing system.

3. Explain self-consistency and when majority voting makes things worse. Sample several reasoning paths at temperature above zero and take the most common final answer. With independent samples and per-sample accuracy p above 0.5, accuracy rises with n (0.7 becomes 0.837 at n = 5). It gets worse when p is below 0.5 (0.3 falls to 0.010 at n = 31) or when one tempting wrong answer is more likely than the right one: the vote then converges on the wrong answer. Correlated errors cap the gain at the share of questions the model usually gets right. TinyLM showed it: consistent wrong answers survived the vote.

4. Best-of-n with a verifier vs majority vote: what is the difference in practice? A vote needs only comparable answers, but it cannot beat the model's typical answer. A verifier recognizes the right answer, so accuracy approaches coverage (the chance at least one sample is right), and with early stopping you pay for few samples on easy cases. In the simulation, voting stalled near 0.76 while verify-and-stop reached 0.978 at n = 5 with 1.56 calls per case, and escalated the rest instead of guessing. The catch is that you need a verifier: tests, a lookup, a schema, a rule.

5. You have three triage models at 69, 68, and 60 percent. Will voting them give you 80? Probably not, and you can tell before trying by measuring error overlap. In this module, three such classifiers had pairwise error correlations of +0.25 to +0.47 and the ensemble scored 51/72 against 50/72 for the best member, within noise. The useful number is the ceiling (at least one member right on 62/72): it tells you a smarter combiner, such as routing by language or confidence, has room that voting cannot use.

6. How do you decide how much test-time compute to spend? Build the cost curve: for each setting (samples, verifier, effort level), measure accuracy on the dev set and compute dollars per 1,000 requests and per 1,000 correct answers with real prices. Then spend compute where it pays: a cheap first pass for everyone and more only where a check fails or confidence is low. In this module, majority-of-15 cost 15 times majority-of-1 for a few points; verify-and-stop cost 1.6 calls for most of the possible gain. Prompt caching lowers repeated-prompt costs but does not change the shape.

7. Why put the refund policy in code instead of in the prompt? Because the policy is a set of ordered rules over typed facts, and code applies ordered rules exactly, for free, every time. The model's job is the part code cannot do: read a ticket in five languages into typed facts, and write a friendly reply. Splitting it this way also gives you a checkpoint: when a decision is wrong you can see whether extraction or the rule failed. The chain in this module escalated 6 of 20 cases at extraction and made zero wrong decisions on the rest.

8. How do you design a router, and what should it do with output it does not understand? List a small number of routes with one-line definitions, ask for JSON, and validate the choice. Anything unparseable or unknown goes to the safest route, which for a support desk is a person. Count calls and cost per path, and bias toward the cheap mistake: a question sent to a person costs minutes, a refund request answered by a generic article costs a customer. In this module the router missed no refund requests and got 67/72 overall.

9. When is map-reduce better than one long-context call? When the input exceeds the window, when accuracy degrades in the middle of long prompts, or when you need exact counts, because models miscount and code does not. At small scale a single call can be cheaper: for 72 tickets it used 2,126 input tokens against about 6,000 for map plus reduce. At 5,000 tickets the single call would be about 145,000 tokens while each map call stays near 71.

10. A teammate adds a "review your answer and fix any mistakes" step and reports it helps. What do you ask? What information does the critique have that the first answer did not? Without an external signal, self-correction tends to flip right answers to wrong: Huang et al. (ICLR 2024) measured GPT-3.5 dropping from 75.8 to 38.1 percent on CommonSenseQA after one self-correction round, and in this module's simulation accuracy fell from 0.71 to 0.43 over four rounds. Ask whether the gain was measured on a held-out set, whether the stop condition used gold labels, and whether a verifier could gate the revisions instead.

11. Your automatic prompt search raised dev accuracy from 67 to 85 percent. Ship it? Not until it is checked on data the search never saw. In this module a search took dev from 32/48 to 41/48 by adding words like "team" and "legal" to category definitions; test accuracy went from 18/24 to 16/24. With a holdout used for early stopping, the search correctly returned the original prompt. Optimize on one set, stop on a second, report once on a third, and log the fingerprint of every version.

12. How would you use a model to write prompts safely? Treat its output as a candidate: give it the current prompt and concrete failures, parse the reply into a structured version, validate it (all categories present, constraints kept, length limit), score it on dev, and confirm on a holdout before shipping. Models rewriting lists lose items; in this module's scripted demo the validator caught a dropped category before any scoring. Never let a model rewrite security or policy wording unreviewed.

Other Tools and Providers

Tool or providerWhat it offersWhen to choose it over this module's approach
DSPy (3.3.1)Declarative pipelines with optimizers such as MIPROv2, GEPA, SIMBA, COPRO, BootstrapFewShotYou have a metric and hundreds of labelled examples and want instructions and examples optimized per step
TextGradTreats natural-language feedback as "gradients" to optimize prompts and outputsYou want optimization driven by written critiques rather than scores
promptfooTest suites and side-by-side comparisons across prompts and providersYou want a ready harness and CI checks for prompt versions
Reasoning models: openai/gpt-oss-120b on Groq, Gemini thinking models, Qwen3 in thinking mode on OllamaBuilt-in reasoning with effort or budget controlsThe task is multi-step and volume or latency allows it; start with a direct prompt
Provider batch APIsAsynchronous processing at a discount (batch_discount in pricing.py)Map steps and prompt search that do not need answers in seconds
LangGraph, LlamaIndex workflowsGraph or workflow frameworks for chains, routers, and loops with statePipelines grow beyond a few functions and you want persistence and tracing
Provider prompt generators and improvers (Anthropic Console, OpenAI playground)Model-written prompt drafts and revisionsMeta-prompting for a first draft; still validate and score
Unit tests, JSON Schema, SQL lookupsExact verifiersAny time an answer can be checked; they turn best-of-n and revision loops from guesswork into gates

Coming Up in Module 6

This module validated JSON by hand at every hand-off and escalated when a model wrote "3 September" instead of an ISO date. Module 6, Structured Outputs and Tool Use, makes that the provider's job: schema-constrained output with the canonical Triage and DraftReply schemas, parsing and repair strategies, and tool calling, so the refund chain's "look up the billing record" becomes a tool the model can call instead of facts we hand it.