CourseLarge Language Models · Module 10: Evaluation · part 51 of 80
Part 51 · Module 10: Evaluation

Part A: Why evaluation is the bottleneck

11 min read·22 Sept 2026

By the end of this module, you'll have:

  • A working eval harness for the Brightlane assistant (examples/m10_evals.py): 91 golden cases (the 72 tickets plus 19 curated hard cases in five languages), a system-under-test interface, deterministic checks for schema, category, citation, reply language, and forbidden content (roadmap dates, claimed refunds or unlocks), per-case JSON, and a summary with Wilson confidence intervals.
  • Real measurements of three systems: a keyword baseline (41.8% of cases pass), the same baseline after error analysis (60.4%), and the noise floor of two stochastic systems (TinyLM sampling and a seeded noisy stand-in, both about 2 points of standard deviation run to run).
  • The statistics to read those numbers honestly: a power calculation that says detecting a 5-point difference needs hundreds of cases, and paired comparisons (exact McNemar and a paired bootstrap) that show the v2 gain is real overall but not yet shown on the held-out split.
  • LLM-as-judge prompts (pointwise, pairwise, reference-based) that run on any provider through supportdesk.llm.chat, a swap harness for position and verbosity bias, and a calibration against 30 hand labels with Cohen's kappa.
  • System-level evals: component metrics against end-to-end, failure attribution that blames retrieval before tools before output, trajectory checks, and a multi-turn eval that catches a truncated-history bug.
  • A pytest CI gate on quality, cost, and latency with a pinned baseline file, which you will see block a real change and then pass, plus online signals, a simulated A/B test with real statistics, and a loop that turns a production incident into a regression case.

Prerequisites: Modules 1 to 8. You need supportdesk.llm.chat and ScriptedLLM (Module 1), token counting and pricing.py (Module 2), sampling and seeds (Module 3), the dev/test discipline from Module 4, the Triage and DraftReply schemas (Module 6), KBSearch (Module 7), and the agent loop and trajectory idea (Module 8). Working Python, no statistics background: every formula is explained and computed in code.

Where we are: Module 9 asked when to change the model itself, and its verification step insisted on a held-out evaluation designed before any training. That rule applies to every change, not only fine-tuning. So far the course has measured things one example at a time, with small ad hoc test sets. This module builds the evaluation system the rest of the course runs on: the golden set, the checks, the statistics, the judges, and the CI gate that decides whether a prompt, model, or code change ships.

How this module is organized

PartWhat it covers
Part A: Why evaluation is the bottleneckNon-determinism, subjectivity, moving targets; what vibes-based development costs, measured; writing the eval before the feature
Part B: Constructing evalsSourcing real inputs and curating hard cases, success criteria, deterministic checks, the harness, golden datasets and regression suites
Part C: Noise, sample size, and comparing systemsWilson intervals, the noise floor, power calculations, error analysis before optimization, paired comparisons
Part D: Human evaluationWhen humans are unavoidable, rating design, agreement (Cohen's kappa), adjudication
Part E: LLM-as-judgeRubrics and scales, pointwise, pairwise, and reference-based judges, bias and the swap test, calibration against human labels, judge cost
Part F: System-level evaluationComponent vs end-to-end, failure attribution, agent trajectories, multi-turn conversations
Part G: Operating evalsCI gates on quality, cost, and latency; online signals; A/B tests; feeding production failures back

Part D comes before Part E on purpose: you cannot trust a judge until you have human labels to check it against.

A word on what is real here. No LLM API key was available when this module was built, so every number in "Comes out" was produced by code that runs without one: a keyword baseline, TinyLM (the 1.07M-parameter model from Module 1), KBSearch, seeded simulations, and ScriptedLLM stand-ins. Each is labelled. Every harness also runs against a real model through supportdesk.llm.chat, and the text gives you the exact command. Where a sample of real-model output appears, it is marked as illustrative.

All examples run in order from the root of your supportdesk copy. Setup once:

bash
cd supportdesk
source .venv/bin/activate          # or however you activated the course venv in Module 1
pip install scikit-learn==1.9.1    # only used to cross-check our kappa in the tests
export PYTHONPATH=.
mkdir -p evals runs
python -c "import torch; torch.set_num_threads(1); import supportdesk.kb_search, supportdesk.pricing, supportdesk.schemas; print('ok')"

Code explained

  • In simple words: step into the project, add one pinned library, make the package importable, and create folders for eval data (evals/) and run results (runs/).
  • What happens: PYTHONPATH=. lets examples/ scripts import both supportdesk.* and each other (examples.m10_stats). scikit-learn is optional; the tests use it only to check our own Cohen's kappa. The last line imports the canonical modules this module builds on. Scripts that use TinyLM call torch.set_num_threads(1) themselves, which keeps runs reproducible and polite on a busy machine.
  • Comes out: ok. If you see ModuleNotFoundError: supportdesk, you are not in the repository root or forgot the export.

The module creates these files. Data files you type in are shown in full where they first appear; the rest are generated by the scripts.

FileWhat it is
examples/m10_stats.pyWilson interval, exact McNemar, paired bootstrap, power, two-proportion test, Cohen's kappa, percentile
examples/m10_evals.pyThe harness: cases, system-under-test interface, checks, runs, summaries, comparisons, noise floor, error analysis, CLI
examples/m10_why.py, m10_power.pyPart A and Part C demonstrations
examples/m10_human.py, m10_judge.pyHuman labels and agreement; judge prompts, bias harness, calibration, cost
examples/m10_system.py, m10_ops.py, m10_lab.pySystem-level evals; online signals, A/B, feedback; the lab
evals/hard_cases.jsonl, labels.jsonl, gate.jsonCurated cases, 30 labelled drafts, gate thresholds (typed in)
evals/golden.jsonl, drafts.jsonl, regressions.jsonl, baseline.jsonGenerated: the frozen golden set, drafts to label, the regression suite, the pinned baseline
tests/test_m10_evals.py, tests/test_m10_gate.pyUnit tests for the harness; the CI quality gate

Part A: Why evaluation is the bottleneck

Three reasons LLM features are hard to test

Ordinary software has a spec and a unit test: add(2, 2) returns 4 or it is broken. An LLM feature breaks that model in three ways.

  • Non-determinism. The same input can give different outputs. Module 3 showed why: sampling draws from a probability distribution, and even at temperature 0 hosted providers do not guarantee identical output across runs. One good answer proves little.
  • Subjectivity. Many outputs have no single right answer. Two good replies to "can I get a refund?" can use different words, and two careful agents can disagree about whether a reply is good enough to send. Correctness becomes a judgment, and judgments need a written standard.
  • Moving targets. The provider updates the model, the help center changes, the product adds a plan, the ticket mix shifts toward a new language. A system that passed last month can fail this month without a single line of your code changing.

Put together: you cannot look at an LLM feature and know whether it works. You have to measure it on many inputs, repeatedly, against written criteria. That measurement, not the prompt or the model, is usually what limits how fast a team can improve an LLM product. Teams with a trusted eval can try ten ideas a day and keep the two that help. Teams without one argue.

Vibes-based development, measured

Vibes-based development means judging a change by trying a few inputs by hand and deciding it "looks good". It is how almost every LLM feature starts, and it fails in a predictable way: you try the inputs you thought of, which are the easy ones. The script below measures both problems on real systems. It samples TinyLM five times on two questions, then scores the course's keyword baseline (built in Part B) on three hand-picked tickets and on the full golden set.

python
"""Module 10, Part A: why evaluation is the bottleneck, measured.

    PYTHONPATH=. python examples/m10_why.py

1. Non-determinism: TinyLM (a real 1M-parameter model) answers the same ticket
   differently on every sampled run.
2. Vibes: three tickets you would try by hand look perfect; the full suite does not.
"""
from __future__ import annotations

import torch

from examples.m10_evals import GOLDEN, build_golden, load_cases, make_baseline, run_eval
from supportdesk import tinylm

torch.set_num_threads(1)
model, tokenizer = tinylm.load()
questions = {"seen in training": "can I get a refund on my annual Team plan?",
             "a real ticket (T-1004)": "We renewed our annual Business plan 5 days ago by mistake. Can we get our money back?"}
for label, question in questions.items():
    print(f"{label}: five sampled runs at temperature 0.7")
    for seed in range(5):
        params = tinylm.SamplingParams(max_new_tokens=40, temperature=0.7, seed=seed, stop=["\n"])
        text = tinylm.generate(model, tokenizer, f"Customer (Ana): Hi, {question}\nAgent (Lena):", params).text
        print(f"  seed {seed}: {text.strip()[:88]}")

if not GOLDEN.exists():
    build_golden()          # the golden file is explained in Part B
system = make_baseline(1)
cases = load_cases()
by_id = {c.id: c for c in cases}
hand_picked = [by_id[i] for i in ("T-1004", "T-1011", "T-1019")]
vibes = run_eval(system, hand_picked, "vibes", save=False)["summary"]
full = run_eval(system, cases, "full", save=False)["summary"]
print(f"\nhand-picked tickets: {vibes['passed']}/{vibes['n']} pass")
print(f"full golden set:     {full['passed']}/{full['n']} pass ({full['pass_rate']:.1%}, "
      f"95% CI {full['ci95'][0]:.1%} to {full['ci95'][1]:.1%})")

Code explained

  • In simple words: a two-part experiment: sample the same model repeatedly, then compare a hand-picked check with the full suite.
  • What happens: the first loop uses tinylm.generate with SamplingParams(max_new_tokens=40, temperature=0.7, seed=seed, stop=["\n"]) in the chat format TinyLM was trained on (Customer (Ana): ... then Agent (Lena):). The second part imports the harness from Part B and runs the same system on 3 cases and on 91.
  • Comes out: see the run below.
bash
python examples/m10_why.py

Code explained

  • In simple words: the same model asked the same thing five times, then the difference between "I tried three tickets and they worked" and "I ran all 91".
  • What happens: the script loads TinyLM, samples 40 tokens at temperature 0.7 with seeds 0 to 4 for a question that appears almost verbatim in its training corpus and for real ticket T-1004. Then it builds the golden set if it is missing (Part B explains it) and runs the version 1 keyword baseline on three tickets a developer might try by hand (a refund, an export question, a login question) and on all 91 cases, with run_eval from the harness.
  • Comes out: (TinyLM text is real model output; the script takes a few seconds)

What does vibes-based development cost? Three things, all visible later in this module. You ship regressions you never saw: the version 3 change in Part G looks harmless and breaks three unanswerable cases. You chase noise: TinyLM's pass rate moves between 11.0% and 17.6% with no change at all (Part C), so a "6-point win" from one run can be nothing. And you cannot answer the question every manager eventually asks, "did the new model make it better?", with anything but an anecdote.

Build the eval before the feature

The discipline this module teaches is simple to state: write the eval before you write the feature. Before changing the prompt, decide which cases the change should fix, which must not break, and what number counts as success. Then make the change and let the eval decide. This is test-driven development with statistics.

SituationUse thisWhy
First week of a feature, no data yet20 to 50 hand-written cases plus the deterministic checks from Part BCatches format and policy failures immediately; cheap to write
Real traffic existsSample real inputs into the golden set, label them, add curated hard casesYour users' inputs are the distribution that matters
Comparing two prompts or modelsPaired run on the same frozen cases, with McNemar or a paired bootstrapPairing removes case difficulty from the comparison (Part C)
Quality is a judgment (tone, helpfulness)A rubric, human labels, then a calibrated LLM judgeA judge you have not checked against humans is a guess (Parts D and E)
Every change after launchA CI gate plus a regression suiteNothing that was fixed breaks again silently (Part G)