CourseLarge Language Models · Module 4: Prompt Engineering Fundamentals · part 17 of 80
Part 17 · Module 4: Prompt Engineering Fundamentals

Part B: Build the test set before you tune

18 min read·22 Sept 2026

Why the order matters

If you write a prompt, try it on a few tickets, tweak it, and try again, you are tuning on whatever tickets you happened to look at. The prompt gets better at those tickets and you have no idea whether it got better at tickets in general. The fix is an order of work you do not break:

Workflow

The Brightlane dataset already has a dev split (48 tickets) for tuning and a test split (24 tickets) for the final check. What it does not have is a gold needs_human label, even though the Triage schema in supportdesk/schemas.py defines one. So step one is labeling.

Step 1: fill in the missing label, from the definition

The schema's definition is: "True if an agent must act (refund, account change, legal, security) or the help center cannot answer it." The dataset marks each ticket answerable (the help center can answer it), which covers the second half. For the first half, the tickets that are answerable but still need an agent to do something are listed by hand, with a reason for each.

json
{
  "definition": "True when an agent must act (refund, account change, legal, security) or the help center cannot answer the ticket. Same wording as supportdesk/schemas.py Triage.needs_human.",
  "rule": "needs_human = (not gold.answerable) or (ticket id in agent_action)",
  "agent_action": {
    "T-1001": "refund of a duplicate charge",
    "T-1004": "refund inside the 14-day window",
    "T-1011": "all users of an enterprise locked out: security and access incident",
    "T-1024": "GDPR erasure request: legal",
    "T-1027": "SLA credit: money owed",
    "T-1031": "refund of a duplicate charge (Spanish)",
    "T-1041": "past invoices are reissued by support on request"
  },
  "annotator": "one annotator (the module author), labeled before any prompt was written; treat as noisy"
}

Code explained

  • In simple words: a small label file that turns a written definition into a rule plus a short list of exceptions, each with its reason.
  • What happens: gold_labels in m04_prompts.py computes needs_human = (not answerable) or (id in agent_action). Tickets such as a locked account are not listed because the help-center answer is "wait 15 minutes; support cannot unlock it sooner", which needs no agent action. The annotator field records that one person labeled these before writing any prompt.
  • Comes out: 11 of 48 dev tickets are needs_human: true (6 unanswerable ones plus 5 answerable ones from the list; the other two listed tickets are in test, where 6 of 24 are true). With one annotator, a few of these calls are arguable; that noise caps how high any prompt can score on this label, which matters for ceiling detection in Part E.

Labeling first also forces you to decide what the label means. If you cannot write down why T-1011 (all users locked out of SSO) needs a human, the model will not guess it either.

Step 2: freeze the sets

A frozen test set is one whose ticket ids and gold labels are recorded with a hash, so any later change (a relabel, a ticket added, a different split) is detected instead of silently changing your scores.

examples/m04_eval.py

python
"""Module 4: a test-set harness for triage prompts.

Order of work this file enforces:
  1. freeze   write evals/triage_<split>_manifest.json (ids + hash of gold labels) BEFORE tuning
  2. run      run one prompt version over the frozen split, save every prediction under runs/
  3. compare  per-label accuracy with 95% Wilson intervals, paired bootstrap between versions

Backends:
  rules   keyword baseline wrapped in ScriptedLLM (a real, dumb classifier; not a model)
  copy    ScriptedLLM that copies the majority label of the few-shot examples in the prompt
  mixed   ScriptedLLM replaying well-formed and malformed replies (plumbing test only)
  llm     supportdesk.llm.chat with your provider (needs a key or a local Ollama)

Examples:
  PYTHONPATH=. python examples/m04_eval.py freeze --split dev
  PYTHONPATH=. python examples/m04_eval.py run --version v2 --backend rules
  LLM_PROVIDER=groq PYTHONPATH=. python examples/m04_eval.py run --version v3 --backend llm --shots balanced --k 6
  PYTHONPATH=. python examples/m04_eval.py compare runs/a.jsonl runs/b.jsonl
"""
from __future__ import annotations

import argparse
import hashlib
import html
import json
import math
import random
import re
import time
from collections import Counter
from dataclasses import dataclass, field
from pathlib import Path

from supportdesk.data import load_tickets
from supportdesk.stand_in import ScriptedLLM

from examples.m04_prompts import (EVAL_DIR, LABELS, ROOT, ExampleSelector, arrange, gold_labels,
                                  load_prompt, parse_triage)

RUN_DIR = ROOT / "runs"


# Step 1: freeze the test set before tuning -------------------------------------------

def manifest_path(split: str) -> Path:
    return EVAL_DIR / f"triage_{split}_manifest.json"


def labels_digest(split: str) -> tuple[list[str], str]:
    tickets = load_tickets(split)
    payload = json.dumps([[t.id, gold_labels(t)] for t in tickets], sort_keys=True)
    return [t.id for t in tickets], hashlib.sha256(payload.encode()).hexdigest()


def freeze_eval_set(split: str) -> dict:
    """Record exactly which tickets and gold labels the prompt will be judged on."""
    path = manifest_path(split)
    if path.exists():
        raise FileExistsError(f"{path.name} already exists; a frozen set is never rewritten silently.")
    ids, digest = labels_digest(split)
    manifest = {"split": split, "n": len(ids), "ids": ids, "labels_sha256": digest,
                "frozen_at": time.strftime("%Y-%m-%d %H:%M:%S")}
    path.write_text(json.dumps(manifest, indent=2) + "\n", encoding="utf-8")
    return manifest


def verify_manifest(split: str) -> dict:
    path = manifest_path(split)
    if not path.exists():
        raise RuntimeError(f"No frozen eval set for {split!r}. Run: python examples/m04_eval.py freeze --split {split}")
    manifest = json.loads(path.read_text(encoding="utf-8"))
    ids, digest = labels_digest(split)
    if ids != manifest["ids"] or digest != manifest["labels_sha256"]:
        raise RuntimeError(f"The {split} set changed after it was frozen. Scores are no longer comparable.")
    return manifest


# Statistics ----------------------------------------------------------------------------

def wilson(k: int, n: int, z: float = 1.96) -> tuple[float, float]:
    """95% Wilson score interval for k successes out of n. Behaves well for small n."""
    if n == 0:
        return 0.0, 0.0
    p = k / n
    centre = (p + z * z / (2 * n)) / (1 + z * z / n)
    half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
    return max(0.0, centre - half), min(1.0, centre + half)


def paired_bootstrap(a: list[bool], b: list[bool], iters: int = 5000, seed: int = 0) -> tuple[float, float, float]:
    """Accuracy difference b - a on the SAME tickets, with a 95% bootstrap interval."""
    rng = random.Random(seed)
    n = len(a)
    diffs = []
    for _ in range(iters):
        idx = [rng.randrange(n) for _ in range(n)]
        diffs.append(sum(b[i] - a[i] for i in idx) / n)
    diffs.sort()
    return (sum(b) - sum(a)) / n, diffs[int(0.025 * iters)], diffs[int(0.975 * iters) - 1]


# Backends: everything below has the same signature as supportdesk.llm.chat ---------------

TICKET_RE = re.compile(r"<ticket[^>]*>\n(.*?)\n</ticket>\s*$", re.S)
LABELS_RE = re.compile(r"<labels>(.*?)</labels>")


def last_ticket_text(messages: list[dict]) -> str:
    """The ticket being classified is the final <ticket> block of the last user message.

    Search from the LAST '<ticket' only: a regex run from the start would begin at
    the first example's tag and swallow every example (a real bug found in Part E).
    """
    content = messages[-1]["content"]
    start = content.rfind("<ticket")
    match = TICKET_RE.search(content, max(start, 0))
    return html.unescape(match.group(1)) if match else ""


def keyword_triage(text: str) -> dict:
    """A deliberately simple English keyword baseline. Rules were written while reading dev tickets."""
    t = text.lower()
    has = lambda *words: any(w in t for w in words)  # noqa: E731
    if has("cancel") or (has("refund") and has("annual", "renew", "months")):
        category = "cancellation"
    elif has("password", "locked", "2fa", "saml", "sso users", "log in", "login", "sign in", "owner",
             "gdpr", "region", "stored", "backups", "reset link"):
        category = "account_access"
    elif has("please add", "integration", "when will", "roadmap", "do you have", "custom domain", "sync"):
        category = "feature_request"
    elif has("charge", "invoice", "refund", "price", "cost", "pay ", "vat", "tax", "discount", "declined", "credit"):
        category = "billing"
    elif has("not loading", "stopped", "fails", "error", "500", "isn't firing", "not firing", "lost my changes",
             "bug", "no push", "collapse"):
        category = "bug"
    else:
        category = "how_to"
    if has("anyone", "all our users", "outage", "attack", "didn't request", "unlock it now"):
        priority = "urgent"
    elif has("charged", "twice", "lost my phone", "can't access", "compliance", "deal-breaker", "garbage",
             "by mistake", "important"):
        priority = "high"
    elif has("?") and not has("stopped", "fails", "error", "missing", "not ", "isn't"):
        priority = "low"
    else:
        priority = "normal"
    needs_human = has("refund", "twice", "gdpr", "legal", "attack", "reissue", "declined", "don't recognize",
                      "garbage", "left company", "owner left", "sla credit", "owed credits")
    return {"category": category, "priority": priority, "needs_human": needs_human}


def rules_backend() -> ScriptedLLM:
    return ScriptedLLM(responder=lambda messages, kw: json.dumps(keyword_triage(last_ticket_text(messages))),
                       model="keyword-rules")


def copy_backend() -> ScriptedLLM:
    """Majority label of the examples shown in the prompt; ties go to the example nearest the ticket."""
    def respond(messages: list[dict], kwargs: dict) -> str:
        shown = [json.loads(html.unescape(s)) for s in LABELS_RE.findall(messages[-1]["content"])]
        if not shown:
            return "I need examples to copy from."
        out = {}
        for label in LABELS:
            counts = Counter(ex[label] for ex in shown)
            top = max(counts.values())
            out[label] = next(ex[label] for ex in reversed(shown) if counts[ex[label]] == top)
        return json.dumps(out)
    return ScriptedLLM(responder=respond, model="copy-examples")


MIXED_REPLIES = [
    '{"category": "billing", "priority": "high", "needs_human": true}',
    '```json\n{"category": "how_to", "priority": "low", "needs_human": false}\n```',
    'Sure! Here is the triage:\n{"category": "bug", "priority": "normal", "needs_human": "false"}',
    '{"category": "Billing Issue", "priority": "high", "needs_human": true}',
    '{"category": "account_access", "priority": "high", "needs_human": tr',
    "I think this is probably a billing question with high priority.",
]


def mixed_backend() -> ScriptedLLM:
    """Cycles through well-formed and malformed replies to test the parser and failure counting."""
    cycle = iter(MIXED_REPLIES * 100)
    return ScriptedLLM(responder=lambda messages, kw: next(cycle), model="mixed-replies")


def make_backend(name: str):
    if name == "rules":
        return rules_backend()
    if name == "copy":
        return copy_backend()
    if name == "mixed":
        return mixed_backend()
    if name == "llm":
        from supportdesk.llm import chat
        return chat
    raise ValueError(f"Unknown backend {name!r}")


# Step 2: run a prompt version over the frozen split -------------------------------------

@dataclass
class Run:
    meta: dict
    records: list[dict] = field(default_factory=list)

    def correct(self, label: str) -> list[bool]:
        """Per-ticket correctness; a parse failure counts as wrong."""
        return [bool(r["pred"]) and r["pred"][label] == r["gold"][label] for r in self.records]


def run_eval(chat_fn, version: str, split: str = "dev", shots: str = "none", k: int = 0,
             most_similar: str = "last", backend: str = "custom", allow_test: bool = False,
             seed: int = 0, **chat_kwargs) -> Run:
    if split == "test" and not allow_test:
        raise RuntimeError("The test split is for the final check only. Pass allow_test=True (CLI: --final) once.")
    manifest = verify_manifest(split)
    template = load_prompt(version)
    pool = load_tickets("dev")  # few-shot examples only ever come from dev, never from test
    selector = ExampleSelector(pool)
    by_id = {t.id: t for t in load_tickets(split)}
    run = Run(meta={"version": version, "fingerprint": template.fingerprint, "split": split, "n": manifest["n"],
                    "shots": shots, "k": k, "most_similar": most_similar, "backend": backend, "seed": seed,
                    "chat_kwargs": chat_kwargs})
    for ticket_id in manifest["ids"]:
        ticket = by_id[ticket_id]
        examples = arrange(selector.select(ticket, shots, k, seed), most_similar)
        messages = template.render(ticket, examples)
        result = chat_fn(messages, **chat_kwargs)
        pred, error = parse_triage(result.text)
        run.records.append({"id": ticket.id, "language": ticket.language, "gold": gold_labels(ticket),
                            "pred": pred, "error": error, "raw": result.text[:300],
                            "input_tokens": result.usage.input_tokens, "output_tokens": result.usage.output_tokens,
                            "latency_ms": result.latency_ms, "examples": [e.id for e in examples]})
    return run


def save_run(run: Run, name: str | None = None) -> Path:
    RUN_DIR.mkdir(exist_ok=True)
    m = run.meta
    name = name or f"{m['version']}-{m['backend']}-{m['shots']}{m['k']}-{m['split']}"
    path = RUN_DIR / f"{name}.jsonl"
    with path.open("w", encoding="utf-8") as f:
        f.write(json.dumps({"meta": m}) + "\n")
        for r in run.records:
            f.write(json.dumps(r, ensure_ascii=False) + "\n")
    return path


def load_run(path: Path | str) -> Run:
    lines = Path(path).read_text(encoding="utf-8").splitlines()
    return Run(meta=json.loads(lines[0])["meta"], records=[json.loads(x) for x in lines[1:]])


# Step 3: report and compare ---------------------------------------------------------------

def summarize(run: Run) -> str:
    m = run.meta
    n = len(run.records)
    failures = [r for r in run.records if r["error"]]
    tokens = sum(r["input_tokens"] for r in run.records) / max(n, 1)
    lines = [f"{m['version']} ({m['fingerprint']}) backend={m['backend']} shots={m['shots']} k={m['k']} "
             f"split={m['split']} n={n}",
             f"  parse failures: {len(failures)}/{n}   mean input tokens: {tokens:.0f}"]
    for label in LABELS:
        ok = run.correct(label)
        lo, hi = wilson(sum(ok), n)
        lines.append(f"  {label:<12} {sum(ok):>3}/{n}  acc {sum(ok) / n:.3f}  95% CI [{lo:.3f}, {hi:.3f}]")
    return "\n".join(lines)


def compare(a: Run, b: Run) -> str:
    if [r["id"] for r in a.records] != [r["id"] for r in b.records]:
        raise ValueError("Runs cover different tickets; paired comparison needs the same frozen set.")
    lines = [f"{b.meta['version']}/{b.meta['backend']} minus {a.meta['version']}/{a.meta['backend']} "
             f"(paired, n={len(a.records)})"]
    for label in LABELS:
        diff, lo, hi = paired_bootstrap(a.correct(label), b.correct(label))
        verdict = "within noise" if lo <= 0 <= hi else ("better" if diff > 0 else "worse")
        lines.append(f"  {label:<12} {diff:+.3f}  95% CI [{lo:+.3f}, {hi:+.3f}]  {verdict}")
    return "\n".join(lines)


def plateau(runs: list[Run], label: str = "category", window: int = 3) -> tuple[bool, str]:
    """Ceiling check: are the last `window` versions all within noise of the best one so far?"""
    if len(runs) < window:
        return False, f"need at least {window} runs"
    scores = [sum(r.correct(label)) / len(r.records) for r in runs]
    best = max(range(len(runs)), key=lambda i: scores[i])
    notes = []
    for i in range(len(runs) - window, len(runs)):
        if i == best:
            continue
        diff, lo, hi = paired_bootstrap(runs[i].correct(label), runs[best].correct(label))
        notes.append((runs[i].meta["version"], round(diff, 3), round(lo, 3), round(hi, 3)))
        if lo > 0:
            return False, f"best ({runs[best].meta['version']}) is clearly above {runs[i].meta['version']}: keep iterating"
    return True, f"last {window} versions within noise of best ({runs[best].meta['version']}): {notes}"


# Command line -----------------------------------------------------------------------------

def main() -> None:
    parser = argparse.ArgumentParser(description="Triage prompt harness (Module 4)")
    sub = parser.add_subparsers(dest="cmd", required=True)
    f = sub.add_parser("freeze")
    f.add_argument("--split", default="dev")
    r = sub.add_parser("run")
    r.add_argument("--version", required=True)
    r.add_argument("--backend", default="rules", choices=["rules", "copy", "mixed", "llm"])
    r.add_argument("--split", default="dev")
    r.add_argument("--shots", default="none", choices=["none", "random", "similar", "balanced"])
    r.add_argument("--k", type=int, default=0)
    r.add_argument("--most-similar", default="last", choices=["first", "last"])
    r.add_argument("--seed", type=int, default=0, help="seed for --shots random")
    r.add_argument("--final", action="store_true", help="allow the test split (use once, at the end)")
    r.add_argument("--max-tokens", type=int, default=None)
    r.add_argument("--reasoning-effort", default=None, help="for reasoning models, e.g. low (llm backend only)")
    c = sub.add_parser("compare")
    c.add_argument("runs", nargs="+")
    args = parser.parse_args()

    if args.cmd == "freeze":
        m = freeze_eval_set(args.split)
        print(f"froze {m['split']}: n={m['n']} labels_sha256={m['labels_sha256'][:16]}")
    elif args.cmd == "run":
        kwargs = {"max_tokens": args.max_tokens} if args.max_tokens else {}
        if args.reasoning_effort:
            kwargs["reasoning_effort"] = args.reasoning_effort
        run = run_eval(make_backend(args.backend), args.version, args.split, args.shots, args.k,
                       args.most_similar, args.backend, allow_test=args.final, seed=args.seed, **kwargs)
        path = save_run(run)
        print(summarize(run))
        print(f"  saved {path.relative_to(ROOT)}")
    else:
        runs = [load_run(p) for p in args.runs]
        for run in runs:
            print(summarize(run))
        for a, b in zip(runs, runs[1:]):
            print(compare(a, b))


if __name__ == "__main__":
    main()

Code explained

  • In simple words: the harness: it freezes the test set, runs a prompt version over every ticket through any chat function, saves every prediction, and reports accuracy with honest error bars.
  • What happens:
    • freeze_eval_set(split) writes evals/triage_<split>_manifest.json with the ticket ids and a SHA-256 hash of all gold labels, and refuses to overwrite an existing manifest. labels_digest computes the hash. verify_manifest(split) recomputes it before every run and stops if anything changed.
    • wilson(k, n) is the Wilson score interval, a 95% confidence interval for a proportion that behaves sensibly at small n (it never goes below 0 or above 1, unlike the textbook plus-or-minus formula). With n=48, it is wide: that is the point.
    • paired_bootstrap(a, b) compares two runs on the same tickets. It resamples tickets with replacement 5,000 times and takes the middle 95% of the accuracy differences. Paired means each resample uses the same tickets for both runs, so a hard ticket is hard for both; this makes the comparison much more sensitive than comparing two separate intervals.
    • last_ticket_text(messages) pulls the ticket back out of the rendered prompt and unescapes it. The comment explains a bug the first version had; Part E walks through how it was found.
    • keyword_triage(text) is the keyword baseline: plain English substring rules for all three labels. It is not a model and was written while reading dev tickets, so its dev score is optimistic.
    • rules_backend, copy_backend, and mixed_backend wrap behaviors in ScriptedLLM so they have exactly the call signature of supportdesk.llm.chat. The harness cannot tell them from a real model, which is what makes them good plumbing tests. copy_backend reads the <labels> of the few-shot examples in the prompt and answers with the majority label (ties go to the example nearest the ticket), so it measures how informative the examples are. mixed_backend cycles through the malformed replies from Part A. make_backend("llm") returns the real chat function.
    • Run holds metadata and per-ticket records; correct(label) counts a parse failure as wrong, because a triage system that cannot read its own output has failed that ticket.
    • run_eval(...) verifies the manifest, loads the version, selects and arranges examples for each ticket (from dev only, never from test), renders, calls the chat function, parses, and records gold, prediction, error, raw reply, tokens, latency, and which examples were shown. It refuses the test split unless allow_test=True.
    • save_run and load_run write and read one JSONL file per run. summarize prints parse failures, mean input tokens, and per-label accuracy with Wilson intervals. compare prints paired differences with a verdict of better, worse, or within noise. plateau is the ceiling check from Part E.
    • main() exposes freeze, run, and compare on the command line.
  • Comes out: see the runs below.

Now freeze both splits, and watch the two guards work.

bash
PYTHONPATH=. python examples/m04_eval.py freeze --split dev
PYTHONPATH=. python examples/m04_eval.py freeze --split test
PYTHONPATH=. python examples/m04_eval.py freeze --split dev                          # second time: refused
PYTHONPATH=. python examples/m04_eval.py run --version v1 --backend rules --split test  # test: refused

Code explained

  • In simple words: record exactly what the prompt will be judged on, then show that the harness will not let you quietly redo it or peek at the test set.
  • What happens: the first two commands write the manifests. The third tries to freeze dev again and hits FileExistsError. The fourth tries to score on test without --final and hits RuntimeError. (Only the last line of each traceback is shown.)
  • Comes out: the dev and test hashes, then the two refusals. If you ever need to relabel, delete the manifest on purpose, relabel, refreeze, and rerun every version you want to compare, because old scores are no longer comparable.
text
froze dev: n=48 labels_sha256=8174913f3e0aaabc
froze test: n=24 labels_sha256=0ddc5b10e2dd83a1
FileExistsError: triage_dev_manifest.json already exists; a frozen set is never rewritten silently.
RuntimeError: The test split is for the final check only. Pass allow_test=True (CLI: --final) once.

Step 3: a baseline and a control

Before you measure a prompt, measure something dumb. The keyword baseline gives the floor a model has to beat, and because it ignores the instructions, it doubles as a control: any change in its score between prompt versions can only come from a bug in the plumbing.

bash
PYTHONPATH=. python examples/m04_eval.py run --version v1 --backend rules
PYTHONPATH=. python examples/m04_eval.py run --version v2 --backend rules
PYTHONPATH=. python examples/m04_eval.py compare runs/v1-rules-none0-dev.jsonl runs/v2-rules-none0-dev.jsonl

Code explained

  • In simple words: score the keyword rules through the full harness with prompts v1 and v2, then compare the two runs ticket by ticket.
  • What happens: each run renders all 48 dev tickets with the given version, sends them to the rules backend, parses, scores, and saves a run file. compare prints both summaries and then the paired difference (only the difference is shown below the two summaries).
  • Comes out: 38/48 on category (0.792, CI 0.657 to 0.883), 33/48 on priority, 43/48 on needs_human. The two versions score identically, as a control should. Notice the interval width: with n=48, a true accuracy anywhere from about 66% to 88% is consistent with 38 correct. v2 costs 265 input tokens per ticket on average against 166 for v1; with the rules backend that extra text buys nothing, and whether it buys something from a real model is exactly what the harness is for.
text
v1 (16af0ee83319) backend=rules shots=none k=0 split=dev n=48
  parse failures: 0/48   mean input tokens: 166
  category      38/48  acc 0.792  95% CI [0.657, 0.883]
  priority      33/48  acc 0.688  95% CI [0.547, 0.801]
  needs_human   43/48  acc 0.896  95% CI [0.778, 0.955]
  saved runs/v1-rules-none0-dev.jsonl
v2 (6dca0c9eb0b6) backend=rules shots=none k=0 split=dev n=48
  parse failures: 0/48   mean input tokens: 265
  category      38/48  acc 0.792  95% CI [0.657, 0.883]
  priority      33/48  acc 0.688  95% CI [0.547, 0.801]
  needs_human   43/48  acc 0.896  95% CI [0.778, 0.955]
  saved runs/v2-rules-none0-dev.jsonl
v2/rules minus v1/rules (paired, n=48)
  category     +0.000  95% CI [+0.000, +0.000]  within noise
  priority     +0.000  95% CI [+0.000, +0.000]  within noise
  needs_human  +0.000  95% CI [+0.000, +0.000]  within noise

Read the failures before you trust the number. Every non-English dev ticket (es, de, ja, hi) gets the wrong category from the rules, because the keywords are English. Priority errors are mostly "normal" tickets phrased as questions, which the rules call "low". A model's advantage will likely show up exactly there.

Step 4: test the plumbing with deliberately bad replies

The mixed backend returns the Part A reply corpus in rotation. The goal is not a score; it is to prove that parse failures are counted, labeled, and do not crash anything.

bash
PYTHONPATH=. python examples/m04_eval.py run --version v1 --backend mixed
PYTHONPATH=. python -c "
from collections import Counter
from examples.m04_eval import load_run
run = load_run('runs/v1-mixed-none0-dev.jsonl')
print(Counter(r['error'] for r in run.records if r['error']))"

Code explained

  • In simple words: run the harness against a fake model that answers in broken ways half the time, and count why each reply failed.
  • What happens: the six scripted replies repeat 8 times over 48 tickets. Three of the six fail to parse, so 24 failures are expected. The one-liner reads the saved run and tallies the error messages.
  • Comes out: exactly 24/48 parse failures, split 16 "no JSON object found" (the truncated and prose replies) and 8 "bad category" (the invented label). The accuracies are meaningless, because the replies ignore the tickets. This is ScriptedLLM output, not model output.
text
v1 (16af0ee83319) backend=mixed shots=none k=0 split=dev n=48
  parse failures: 24/48   mean input tokens: 166
  category       4/48  acc 0.083  95% CI [0.033, 0.196]
  priority       6/48  acc 0.125  95% CI [0.059, 0.247]
  needs_human   12/48  acc 0.250  95% CI [0.149, 0.388]
  saved runs/v1-mixed-none0-dev.jsonl
Counter({'no JSON object found': 16, "bad category 'billing issue'": 8})

Step 5: the same harness with a real model

With a key (or a local Ollama), the only change is the backend. This is the command to run for your own numbers:

bash
export LLM_PROVIDER=groq GROQ_API_KEY=your-key-here     # or: LLM_PROVIDER=gemini GEMINI_API_KEY=...; or LLM_PROVIDER=ollama
PYTHONPATH=. python examples/m04_eval.py run --version v1 --backend llm --reasoning-effort low
PYTHONPATH=. python examples/m04_eval.py run --version v2 --backend llm --reasoning-effort low
PYTHONPATH=. python examples/m04_eval.py run --version v3 --backend llm --shots balanced --k 6 --reasoning-effort low
PYTHONPATH=. python examples/m04_eval.py compare runs/v1-llm-none0-dev.jsonl runs/v2-llm-none0-dev.jsonl runs/v3-llm-balanced6-dev.jsonl

Code explained

  • In simple words: the same three versions, now answered by a real model through supportdesk.llm.chat, then compared pairwise.
  • What happens: --backend llm passes the real chat function to run_eval. --reasoning-effort low is forwarded for reasoning models such as Groq's default openai/gpt-oss-120b (Module 3); leave it out for models that do not accept it. Temperature is 0 by default in chat. Each run makes 48 calls and records real token usage and latency. compare prints all three summaries and the v1-to-v2 and v2-to-v3 paired differences.
  • Comes out: in this build there is no key, so the first call stops with RuntimeError: Set GROQ_API_KEY in your environment to use groq. The block below is an illustrative sample run (not captured in this build; produced for teaching). Your output will differ, including the numbers, and you should not quote them. It shows the shape to expect and one pattern worth looking for: a real model usually beats the keyword rules on non-English tickets, and gains between versions often land inside the noise band at n=48.
text
v2/llm minus v1/llm (paired, n=48)
  category     +0.021  95% CI [-0.042, +0.083]  within noise
  priority     +0.104  95% CI [+0.021, +0.188]  better
  needs_human  +0.063  95% CI [-0.021, +0.146]  within noise