Part 2: Measure before you train
The task, and the shared setup file
The task for most of this module: given a ticket, output one of Brightlane's six categories (billing, cancellation, account_access, bug, how_to, feature_request). The model sees this prompt and must produce the completion after Category::
Ticket: Charged twice this month
Hi, my card was charged 288 USD twice on 3 September for the Team plan (invoice INV-2026-004512). Please refund the duplicate.
Category: billing
Code explained
- In simple words: one training example: the ticket as the prompt, the label as the answer.
- What happens: the prompt is the subject and body after
Ticket:, thenCategory:. The completion is a space, one label word, and a newline. TinyLM's context is 128 tokens, so ticket text is cut to 108 tokens to leave room. - Comes out: nothing; this is the data format. Part 4 explains why the completion uses a short word like
cancelinstead of the category namecancellation.
Every script in this module imports one setup file, so the prompt format, the way we score a prediction, and the training loop are identical everywhere. Read it once now; the rest of the module refers back to it.
he training loop are identical everywhere. Read it once now; the rest of the module refers back to it.
examples/m09_setup.py (click to expand)
PYTHONCopy
"""Shared setup for Module 9: the triage task format, evaluation, and a small SFT training loop.
Every other m09 example imports from here, so the prompt format, the
label scorer, and the metric are identical everywhere.
"""
from __future__ import annotations
import math
import random
import time
from pathlib import Path
import torch
from supportdesk.data import CATEGORIES, Ticket
from supportdesk.tinylm import SamplingParams, TinyGPT, generate, loss_on
torch.set_num_threads(1) # one thread: reproducible, and plenty for a 1M-parameter model
ROOT = Path(__file__).resolve().parents[1]
MODELS = ROOT / "models"
DATA_OUT = ROOT / "data" / "m09"
PAD_ID = 0 # <|endoftext|> doubles as padding; padded positions are masked out of the loss
TICKET_TOKENS = 108 # ticket text is cut to this many tokens so prompt + answer fit in 128
def seed_everything(seed: int = 0) -> None:
random.seed(seed)
torch.manual_seed(seed)
# The task format ---------------------------------------------------------------
def ticket_prompt(tokenizer, subject: str, body: str) -> str:
"""'Ticket: <subject>\\n<body>\\nCategory:' with the ticket text cut to fit the context."""
ids = tokenizer.encode(f"Ticket: {subject}\n{body}").ids[:TICKET_TOKENS]
return tokenizer.decode(ids) + "\nCategory:"
def prompt_for(tokenizer, ticket: Ticket) -> str:
return ticket_prompt(tokenizer, ticket.subject, ticket.body)
# Each category is written as ONE word the base model saw often in pretraining (see Part 4).
LABEL_WORDS = {"billing": "billing", "cancellation": "cancel", "account_access": "account",
"bug": "error", "how_to": "how", "feature_request": "request"}
RAW_NAMES = {c: c for c in CATEGORIES} # the category names themselves, for comparison
def answer_for(category: str, words: dict = LABEL_WORDS) -> str:
"""The completion the model must learn: a space, the label word, a newline."""
return f" {words[category]}\n"
# Decoding ------------------------------------------------------------------------
def category_constraint(tokenizer, words: dict = LABEL_WORDS):
"""Constrained decoding: only token paths that spell one of the six label words are allowed.
Kept to demonstrate a trap (Part 2): with greedy decoding, TinyLM's generate() falls back to the
LOWEST allowed token id whenever the model's top token is not allowed, so it is not a classifier.
"""
paths = [tokenizer.encode(" " + w).ids for w in words.values()]
def allowed(generated: list[int]) -> set[int]:
n = len(generated)
return {p[n] for p in paths if len(p) > n and p[:n] == generated}
return allowed
@torch.no_grad()
def label_logprobs(model: TinyGPT, tokenizer, prompt: str, words: dict = LABEL_WORDS) -> dict[str, float]:
"""log P(" <label word>" | prompt) for every category: the sum over the label's tokens."""
p = tokenizer.encode(prompt).ids
labels = {c: tokenizer.encode(" " + w).ids for c, w in words.items()}
width = max(len(l) for l in labels.values())
p = p[-(model.cfg.context - width):]
x = torch.full((len(labels), len(p) + width - 1), PAD_ID)
for row, l in enumerate(labels.values()):
seq = p + l
x[row, :len(seq) - 1] = torch.tensor(seq[:-1])
logp = torch.log_softmax(model(x), dim=-1)
scores = {}
for row, (c, l) in enumerate(labels.items()):
positions = torch.arange(len(p) - 1, len(p) - 1 + len(l))
scores[c] = logp[row, positions, torch.tensor(l)].sum().item()
return scores
def classify(model: TinyGPT, tokenizer, prompt: str, words: dict = LABEL_WORDS) -> str:
"""The category whose label word the model finds most likely after the prompt."""
scores = label_logprobs(model, tokenizer, prompt, words)
return max(scores, key=scores.get)
def free_label(model: TinyGPT, tokenizer, prompt: str, words: dict = LABEL_WORDS) -> str:
"""Plain greedy generation up to a newline; text that is not a label word comes back as-is."""
g = generate(model, tokenizer, prompt, SamplingParams(max_new_tokens=8, temperature=0, stop=["\n"]))
text = g.text.strip()
return {w: c for c, w in words.items()}.get(text, text)
def constrained_greedy_label(model: TinyGPT, tokenizer, prompt: str, words: dict = LABEL_WORDS) -> str:
"""Greedy generate() with the allowed-token mask (shows the lowest-id fallback trap)."""
g = generate(model, tokenizer, prompt, SamplingParams(max_new_tokens=8, temperature=0),
allowed=category_constraint(tokenizer, words))
return {w: c for c, w in words.items()}.get(g.text.strip(), g.text.strip())
def wilson(k: int, n: int, z: float = 1.96) -> tuple[float, float]:
"""95% Wilson score interval for a proportion k/n (behaves well for small n)."""
if n == 0:
return (0.0, 1.0)
p = k / n
centre = (p + z * z / (2 * n)) / (1 + z * z / n)
half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
return (max(0.0, centre - half), min(1.0, centre + half))
def evaluate(model: TinyGPT, tokenizer, tickets: list[Ticket], mode: str = "score", prompt_fn=None,
words: dict = LABEL_WORDS) -> dict:
"""Accuracy on `tickets`, with n, a 95% CI, and every prediction.
mode "score" = log-probability scoring over the six labels (the method used everywhere),
"free" = plain greedy generation, "greedy_mask" = generate() with an allowed-token mask.
"""
prompt_fn = prompt_fn or prompt_for
decide = {"score": classify, "free": free_label, "greedy_mask": constrained_greedy_label}[mode]
preds = [decide(model, tokenizer, prompt_fn(tokenizer, t), words) for t in tickets]
gold = [t.gold["category"] for t in tickets]
k = sum(p == g for p, g in zip(preds, gold))
lo, hi = wilson(k, len(tickets))
return {"accuracy": k / len(tickets), "correct": k, "n": len(tickets), "ci95": (round(lo, 3), round(hi, 3)),
"valid_label_rate": sum(p in CATEGORIES for p in preds) / len(preds), "preds": preds}
def fmt(result: dict) -> str:
lo, hi = result["ci95"]
return f"{result['correct']}/{result['n']} = {result['accuracy']:.0%} (95% CI {lo:.0%} to {hi:.0%})"
# Supervised fine-tuning ------------------------------------------------------------
def encode_example(tokenizer, prompt: str, answer: str, context: int = 128) -> tuple[list[int], list[int]]:
"""Token ids for prompt+answer and a mask that is 1 only on answer tokens (loss on the answer only)."""
p = tokenizer.encode(prompt).ids
a = tokenizer.encode(answer).ids
ids = (p + a)[-(context + 1):]
mask = ([0] * len(p) + [1] * len(a))[-(context + 1):]
return ids, mask
def make_batch(rows: list[tuple[list[int], list[int]]]) -> tuple[torch.Tensor, torch.Tensor, torch.Tensor]:
"""Right-pad a list of (ids, mask) into x, y, and a loss mask aligned with y."""
width = max(len(ids) for ids, _ in rows) - 1
x = torch.full((len(rows), width), PAD_ID)
y = torch.full((len(rows), width), PAD_ID)
m = torch.zeros((len(rows), width))
for i, (ids, mask) in enumerate(rows):
n = len(ids) - 1
x[i, :n] = torch.tensor(ids[:-1])
y[i, :n] = torch.tensor(ids[1:])
m[i, :n] = torch.tensor(mask[1:], dtype=torch.float)
return x, y, m
def sft_train(model: TinyGPT, tokenizer, pairs: list[tuple[str, str]], *, epochs: int = 5, lr: float = 1e-3,
batch_size: int = 16, seed: int = 0, log_every: int = 0, on_epoch=None) -> dict:
"""Train on (prompt, answer) pairs with loss on answer tokens only. Trains whatever has requires_grad.
`on_epoch(epoch)` is called after each epoch (in eval mode), for example to score a validation set.
"""
seed_everything(seed)
rows = [encode_example(tokenizer, p, a) for p, a in pairs]
params = [p for p in model.parameters() if p.requires_grad]
opt = torch.optim.AdamW(params, lr=lr, weight_decay=0.0)
rng = random.Random(seed)
model.train()
started, step, losses = time.perf_counter(), 0, []
for epoch in range(epochs):
order = list(range(len(rows)))
rng.shuffle(order)
for i in range(0, len(order), batch_size):
x, y, m = make_batch([rows[j] for j in order[i:i + batch_size]])
loss = loss_on(model, x, y, m)
opt.zero_grad()
loss.backward()
torch.nn.utils.clip_grad_norm_(params, 1.0)
opt.step()
step += 1
losses.append(loss.item())
if log_every and step % log_every == 0:
print(f" step {step:4d} loss {loss.item():.3f}")
if on_epoch is not None:
model.eval()
paused = time.perf_counter()
on_epoch(epoch + 1)
started += time.perf_counter() - paused # evaluation time does not count as training time
model.train()
model.eval()
return {"steps": step, "seconds": round(time.perf_counter() - started, 1), "final_loss": round(sum(losses[-5:]) / len(losses[-5:]), 4)}
@torch.no_grad()
def answer_loss(model: TinyGPT, tokenizer, pairs: list[tuple[str, str]]) -> float:
"""Average loss on answer tokens for held-out pairs: a smoother signal than accuracy on 12 items."""
x, y, m = make_batch([encode_example(tokenizer, p, a) for p, a in pairs])
return round(loss_on(model, x, y, m).item(), 4)
def corpus_validation_text(tokenizer) -> str:
"""The same last 5% of corpus tokens that TinyLM's pretraining held out for validation."""
ids = tokenizer.encode((ROOT / "data" / "corpus.txt").read_text(encoding="utf-8")).ids
return tokenizer.decode(ids[int(len(ids) * 0.95):])
def count_trainable(model: torch.nn.Module) -> int:
return sum(p.numel() for p in model.parameters() if p.requires_grad)
# Loading the dataset built by m09_data.py ----------------------------------------------
class HeldOut:
"""A validation row that looks like a Ticket to evaluate() (subject, body, gold)."""
def __init__(self, row: dict) -> None:
self.subject, self.body, self.gold = row["subject"], row["body"], {"category": row["category"]}
def load_rows(name: str) -> list[dict]:
import json
with (DATA_OUT / f"{name}.jsonl").open(encoding="utf-8") as f:
return [json.loads(line) for line in f]
def to_pairs(tokenizer, rows: list[dict]) -> list[tuple[str, str]]:
return [(ticket_prompt(tokenizer, r["subject"], r["body"]), answer_for(r["category"])) for r in rows]
Code explained
- In simple words: the rulebook for the whole module: how a ticket becomes a prompt, how we read the model's answer, how we score it, and how we train.
- What happens:
- Top of file:
torch.set_num_threads(1)keeps runs reproducible and polite on a shared machine.PAD_ID = 0is the<|endoftext|>token, reused as padding.TICKET_TOKENS = 108keeps prompt plus answer inside TinyLM's 128-token window. ticket_promptandprompt_for: buildTicket: <subject>\n<body>\nCategory:, cutting the ticket (not the instruction) when it is too long.LABEL_WORDSandanswer_for: each category is written as one common word (cancel,account,error,how,request,billing).RAW_NAMESkeeps the real names for comparison.category_constraintandconstrained_greedy_label: greedy generation with an allowed-token mask. Kept only to show a trap in the next section.label_logprobs: the scoring method used everywhere. For each of the six labels it appends the label's tokens to the prompt, runs all six sequences as one batch, and adds up the log-probabilities of the label tokens. A log-probability is the logarithm of the probability the model gives a token (Module 3); adding them gives the log-probability of the whole label. Closer to 0 means more likely.classify: the label with the highest log-probability. The model never "writes" anything; we ask it how likely each allowed answer is and take the best one. This is the standard way to use a language model as a classifier, and it always returns a valid label.free_label: plain greedy generation, to see what the model writes on its own.wilson: a 95% confidence interval for accuracy, the range the true accuracy plausibly lies in given only n test items. The Wilson version behaves well for small n. With 24 test tickets the intervals are wide, and you will see that constantly.evaluate: runs one of the three decision modes over a list of tickets and returns accuracy, n, the interval, and every prediction.encode_exampleandmake_batch: turn (prompt, answer) pairs into padded tensors plus a loss mask that is 1 only on answer tokens. The model is trained to produce the answer, not to reproduce the ticket.sft_train: the supervised fine-tuning loop. Shuffle, batch, compute the masked cross-entropy loss (how surprised the model is by the right answer, Module 1), backpropagate, clip gradients, step the AdamW optimizer. It trains whatever parameters haverequires_grad=True, which is how the same loop does full fine-tuning and LoRA. An epoch is one pass over the training data; the learning rate (lr) is how big each update step is.answer_loss,corpus_validation_text,count_trainable,HeldOut,load_rows,to_pairs: small helpers for validation loss, the corpus text TinyLM's pretraining held out, counting trainable parameters, and loading the datasets Part 3 builds.
- Top of file:
- Comes out: nothing on import. Every later script uses it.
Freeze the eval plan, then measure the prompted baseline
The single most important habit in this module: decide how you will judge the fine-tune before you train it. If you pick the metric, the test set, or the success bar after seeing results, you will (without meaning to) pick the ones that make your model look good. So the first script writes the plan to disk and fingerprints the test set, and only then measures anything.
examples/m09_baseline.py
"""Before any training: freeze the evaluation plan, then measure what prompting the base model can do."""
import hashlib
import json
from collections import Counter
from examples.m09_setup import (DATA_OUT, LABEL_WORDS, RAW_NAMES, evaluate, fmt, label_logprobs, prompt_for,
wilson)
from supportdesk.data import CATEGORIES, load_tickets
from supportdesk.tinylm import load, next_token_distribution
test = load_tickets("test")
dev = load_tickets("dev")
# 1. Write the eval plan down first, and fingerprint the test set so nobody can quietly change it.
plan = {
"task": "ticket -> one of six categories",
"metric": "exact-match accuracy of the label with the highest log-probability (six candidates)",
"held_out_split": "test",
"n": len(test),
"test_ids_sha256": hashlib.sha256(",".join(t.id for t in test).encode()).hexdigest()[:16],
"training_data_may_use": "dev split only (48 tickets) plus data generated from it",
"ship_rule": "fine-tune must beat the best cheap baseline by more than the noise (paired exact McNemar "
"p < 0.05) AND keep corpus perplexity within 10% of the base model",
"also_report": ["95% Wilson interval", "per-category errors", "perplexity on corpus validation text and KB"],
}
DATA_OUT.mkdir(parents=True, exist_ok=True)
(DATA_OUT / "eval_plan.json").write_text(json.dumps(plan, indent=2) + "\n")
print("eval plan frozen:", plan["test_ids_sha256"], "n =", plan["n"])
# 2. Floors that need no model at all.
majority = Counter(t.gold["category"] for t in dev).most_common(1)[0][0]
k = sum(t.gold["category"] == majority for t in test)
lo, hi = wilson(k, len(test))
print(f"chance (1 of 6): {1 / len(CATEGORIES):.0%}")
print(f"always '{majority}' (dev majority): {k}/{len(test)} = {k / len(test):.0%} (95% CI {lo:.0%} to {hi:.0%})")
print("test label counts:", dict(Counter(t.gold["category"] for t in test)))
# 3. The base model, asked the obvious way: generate after "Category:".
model, tok = load()
free = evaluate(model, tok, test, mode="free")
print("base, free generation:", fmt(free), f"valid label rate {free['valid_label_rate']:.0%}")
for t, p in list(zip(test, free["preds"]))[:3]:
print(f" {t.id} gold={t.gold['category']:<15} model wrote {p!r}")
# 4. A trap: greedy generate() with an allowed-token mask.
masked = evaluate(model, tok, test, mode="greedy_mask")
print("base, greedy + allowed-token mask:", fmt(masked), dict(Counter(masked["preds"])))
first = test[0]
top = next_token_distribution(model, tok, prompt_for(tok, first), top=3)
ids = {w: tok.encode(" " + w).ids[0] for w in LABEL_WORDS.values()}
print(f" {first.id}: top next tokens {top}; label token ids {ids}")
# 5. The fix: score each of the six labels by log-probability and take the best.
for name, words in (("category names", RAW_NAMES), ("one-word labels", LABEL_WORDS)):
scored = evaluate(model, tok, test, words=words)
print(f"base, log-prob scoring, {name}:", fmt(scored), dict(Counter(scored["preds"])))
scores = label_logprobs(model, tok, prompt_for(tok, first))
print(f" {first.id} (gold {first.gold['category']}):", {c: round(v, 2) for c, v in scores.items()})
# 6. Prompting harder: list the labels, or show one short example per category (few-shot).
def listed_prompt(tokenizer, ticket):
return "Labels: " + ", ".join(LABEL_WORDS.values()) + "\n" + prompt_for(tokenizer, ticket)
SHOTS = {c: next(t for t in dev if t.gold["category"] == c and t.language == "en") for c in CATEGORIES}
def short(tokenizer, text, n):
return tokenizer.decode(tokenizer.encode(text).ids[:n])
def few_shot_prompt(tokenizer, ticket):
"""Six worked examples (subject cut to 6 tokens) plus as much of the ticket as fits in 124 tokens."""
shots = "".join(f"Ticket: {short(tokenizer, s.subject, 6)}\nCategory: {LABEL_WORDS[c]}\n" for c, s in SHOTS.items())
n = 40
while True: # cutting multi-byte text mid-character can re-encode longer, so re-measure
text = shots + short(tokenizer, f"Ticket: {ticket.subject}\n{ticket.body}", n) + "\nCategory:"
if len(tokenizer.encode(text).ids) <= 124:
return text
n -= 4
print("few-shot prompt, longest:", max(len(tok.encode(few_shot_prompt(tok, t)).ids) for t in test), "tokens")
for name, fn in (("labels listed", listed_prompt), ("6-shot", few_shot_prompt)):
r = evaluate(model, tok, test, prompt_fn=fn)
print(f"base, {name}, log-prob scoring:", fmt(r), dict(Counter(r["preds"])))
# 7. Calibration: subtract each label's score on a content-free ticket, removing the model's label prior.
null = label_logprobs(model, tok, "Ticket: N/A\nCategory:")
preds = []
for t in test:
s = label_logprobs(model, tok, prompt_for(tok, t))
preds.append(max(s, key=lambda c: s[c] - null[c]))
k = sum(p == t.gold["category"] for p, t in zip(preds, test))
lo, hi = wilson(k, len(test))
print(f"base, calibrated log-prob scoring: {k}/{len(test)} = {k / len(test):.0%} (95% CI {lo:.0%} to {hi:.0%})",
dict(Counter(preds)))
Code explained
- In simple words: write down the rules of the game, then see how far prompting the untouched base model gets.
- What happens:
planrecords the task, the metric, the held-out split, and the ship rule: the fine-tune must beat the best cheap baseline by more than noise (an exact McNemar test, explained in Part 7, with p < 0.05) and keep perplexity on the original corpus within 10%.test_ids_sha256is a hash of the 24 test ticket ids; the lab refuses to run if it changes.- Two floors that need no model: chance (1 in 6) and always predicting the most common dev category.
- Free generation: let the base model write after
Category:. - The trap: greedy
generate()with an allowed-token mask (only the six label tokens allowed). It prints the model's real top-3 next tokens and the label token ids. - The fix: log-probability scoring, once with the raw category names and once with the one-word labels.
- Harder prompting: list the allowed labels first, or show six short worked examples (few-shot, Module 4). The few-shot prompt is trimmed until it fits in 124 tokens; cutting Japanese or Hindi text mid-character can re-encode longer, so the loop re-measures.
- Calibration: subtract each label's score on a content-free ticket (
N/A), which removes the model's built-in preference for some labels (the "calibrate before use" trick from Zhao et al., 2021). Run it withpython -m examples.m09_baseline.
- Comes out:
eval plan frozen: 5a95381c9e48536b n = 24
chance (1 of 6): 17%
always 'how_to' (dev majority): 6/24 = 25% (95% CI 12% to 45%)
test label counts: {'cancellation': 2, 'billing': 5, 'account_access': 6, 'how_to': 6, 'bug': 3, 'feature_request': 2}
base, free generation: 0/24 = 0% (95% CI 0% to 14%) valid label rate 0%
T-1003 gold=cancellation model wrote 'Team costs 12 USD per user per user'
T-1006 gold=billing model wrote 'Team costs 12 USD per user per user'
T-1009 gold=account_access model wrote 'use one of your identity provider; re'
base, greedy + allowed-token mask: 6/24 = 25% (95% CI 12% to 45%) {'how_to': 24}
T-1003: top next tokens [(' Team', 0.4286), (' use', 0.103), (' Gantt', 0.0687)]; label token ids {'billing': 540, 'cancel': 566, 'account': 576, 'error': 1017, 'how': 417, 'request': 636}
base, log-prob scoring, category names: 5/24 = 21% (95% CI 9% to 40%) {'billing': 24}
base, log-prob scoring, one-word labels: 3/24 = 12% (95% CI 4% to 31%) {'cancellation': 11, 'billing': 13}
T-1003 (gold cancellation): {'billing': -7.91, 'cancellation': -7.17, 'account_access': -9.98, 'bug': -10.16, 'how_to': -9.09, 'feature_request': -10.4}
few-shot prompt, longest: 124 tokens
base, labels listed, log-prob scoring: 5/24 = 21% (95% CI 9% to 40%) {'cancellation': 7, 'billing': 17}
base, 6-shot, log-prob scoring: 4/24 = 17% (95% CI 7% to 36%) {'cancellation': 16, 'bug': 8}
base, calibrated log-prob scoring: 7/24 = 29% (95% CI 15% to 49%) {'cancellation': 1, 'account_access': 19, 'feature_request': 2, 'billing': 2}
Read it top to bottom:
- The floors: always answering
how_togets 6/24 (25%, interval 12% to 45%). Any system has to beat that. - Free generation scores 0/24 with a 0% valid-label rate. TinyLM has never seen
Category:followed by a label, so it continues with a memorized agent reply ("Team costs 12 USD per user..."). This is the "base models complete text, they do not follow instructions" point from Module 1, measured. - The trap. The masked greedy decoder says
how_tofor all 24 tickets and scores exactly the majority floor, which looks like a real result. It is not. The model's top next token for T-1003 isTeam(43%), which is not allowed. When the top token is masked out, TinyLM'sgenerate()falls back to a uniform choice among the allowed tokens and greedy picks the first, which is the lowest token id:howis id 417, the smallest of the six. So "prediction" here means "whichever label has the lowest id". Diagnose this kind of thing by printing the top tokens before trusting any constrained decoder, and use log-probability scoring for classification instead. - Log-probability scoring is honest, and the honest answer is that prompting TinyLM does not work: 3/24 to 5/24 with various label spellings and prompts, at or below the floor. The predictions collapse onto one or two labels (
billing24 times), which is the label prior talking, not the ticket. Calibration spreads the predictions a little and reaches 7/24 (29%), still inside the majority floor's interval.
So the prompted-baseline number to beat with TinyLM is about 25%, the same as "always say how_to". With a real model the prompted baseline would be far higher; two sections below show how to measure it.
The cheap alternatives, measured before any training
The other half of the baseline is the systems that need no language model at all. If one of these meets the target, you are done.
examples/m09_cheap_baselines.py
"""The alternatives a fine-tune must beat: keyword rules, TF-IDF + logistic regression, and retrieval."""
from collections import Counter
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from examples.m09_setup import wilson
from supportdesk.data import load_tickets
from supportdesk.kb_search import KBSearch
dev, test = load_tickets("dev"), load_tickets("test")
gold = [t.gold["category"] for t in test]
# 1. Keyword rules, written by reading the dev tickets only (first match wins, order matters).
RULES = [
("feature_request", ["when will", "roadmap", "integration", "do you have", "custom domain"]),
("cancellation", ["cancel", "money back", "stop now", "refund for", "reembolso", "解約"]),
("account_access", ["password", "locked", "login", "signing in", "2fa", "sso users", "owner", "region",
"passwort", "data stored", "dpa", "パスワード"]),
("bug", ["not loading", "stopped", "failing", "error", "isn't firing", "lost my changes", "outage",
"benachrichtigungen"]),
("billing", ["charge", "invoice", "price", "cost", "pay", "tax", "discount", "cobr", "कीमत", "downgrade"]),
]
# Guard against peeking: every keyword must occur in at least one dev ticket.
for _category, _words in RULES:
for _w in _words:
assert any(_w in t.text.lower() for t in dev), f"keyword {_w!r} does not come from dev"
def keyword_label(text: str) -> str:
text = text.lower()
for category, words in RULES:
if any(w in text for w in words):
return category
return "how_to" # the most common dev category is the fallback
# 2. A classic text classifier on character n-grams (works across languages), trained on dev only.
clf = make_pipeline(TfidfVectorizer(analyzer="char_wb", ngram_range=(2, 4), sublinear_tf=True),
LogisticRegression(max_iter=2000, C=10.0))
clf.fit([t.text for t in dev], [t.gold["category"] for t in dev])
# 3. Retrieval: find the closest help article, then use the category dev tickets for that article usually have.
search = KBSearch()
by_article = {}
for t in dev:
by_article.setdefault(t.gold["kb_article"], Counter())[t.gold["category"]] += 1
def retrieval_label(text: str) -> str:
hits = search.search(text, k=1)
if not hits or hits[0].article_id not in by_article:
return "how_to"
return by_article[hits[0].article_id].most_common(1)[0][0]
def report(name, preds):
k = sum(p == g for p, g in zip(preds, gold))
lo, hi = wilson(k, len(gold))
print(f"{name:<28} {k:>2}/{len(gold)} = {k / len(gold):.0%} (95% CI {lo:.0%} to {hi:.0%})")
return preds
if __name__ == "__main__":
report("keyword rules", [keyword_label(t.text) for t in test])
report("tf-idf + logistic regression", list(clf.predict([t.text for t in test])))
report("retrieval (KB top-1)", [retrieval_label(t.text) for t in test])
Code explained
- In simple words: three non-LLM classifiers, all built from the dev split only, scored on the same 24 test tickets.
- What happens:
- Keyword rules: an ordered list of (category, keywords); the first category with a matching keyword wins, else
how_to. The rules were written by reading the dev tickets only, and a guard asserts every keyword appears in at least one dev ticket, so nobody can quietly add a word they saw in the test set. A few keywords are Spanish, German, Japanese, and Hindi because the dev tickets are. - TF-IDF plus logistic regression: a classic text classifier over character 2-to-4-grams (which works across languages without a tokenizer), trained on the 48 dev tickets.
- Retrieval: find the closest KB article with Module 7's BM25 search, then predict the category that dev tickets pointing to that article usually have. Run it with
python -m examples.m09_cheap_baselines.
- Keyword rules: an ordered list of (category, keywords); the first category with a matching keyword wins, else
- Comes out:
keyword rules 15/24 = 62% (95% CI 43% to 79%)
tf-idf + logistic regression 12/24 = 50% (95% CI 31% to 69%)
retrieval (KB top-1) 12/24 = 50% (95% CI 31% to 69%)
Keyword rules get 15/24 (62%). TF-IDF and retrieval get 12/24. The three intervals overlap heavily, so these are not clearly different from each other, but all three are clearly above the prompted TinyLM (about 25%). The rules are a few dozen keywords and run in microseconds. That is the bar the fine-tune has to clear, together with whatever a strong hosted model scores prompted zero-shot.
The prompting baseline with a hosted model
TinyLM's 25% is the floor for this model. In a real project the question is "does a fine-tuned small model beat simply prompting a strong model?", so you must also measure the strong model prompted zero-shot. This harness does that through supportdesk.llm.chat, prices the run, and (Part 3) turns the strong model's labels into distillation data with provenance.
examples/m09_hosted_baseline.py
"""The prompting baseline with a real hosted model, and distillation data with provenance.
Run with a key: python -m examples.m09_hosted_baseline (uses LLM_PROVIDER, default groq)
Run without one: python -m examples.m09_hosted_baseline --stand-in (ScriptedLLM: tests the plumbing only)
"""
import argparse
import json
import re
import time
from datetime import date
from examples.m09_cheap_baselines import keyword_label
from examples.m09_setup import DATA_OUT, wilson
from supportdesk.data import CATEGORIES, load_tickets
from supportdesk.llm import chat, resolve
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.stand_in import ScriptedLLM
SYSTEM = ("You triage support tickets for Brightlane, a project-management SaaS. Reply with exactly one "
"category and nothing else: " + ", ".join(CATEGORIES) + ".")
def parse_label(text: str) -> str | None:
"""The first category name in the reply, or None if there is none."""
found = re.findall("|".join(CATEGORIES), (text or "").lower())
return found[0] if found else None
def run(llm, tickets, **kwargs):
preds, usages = [], []
for t in tickets:
result = llm([{"role": "system", "content": SYSTEM}, {"role": "user", "content": t.text}],
temperature=0.0, max_tokens=400, **kwargs)
preds.append(parse_label(result.text))
usages.append(result.usage)
return preds, usages
if __name__ == "__main__":
args = argparse.ArgumentParser()
args.add_argument("--stand-in", action="store_true")
stand_in = args.parse_args().stand_in
if stand_in: # NOT a model: answers with the keyword rules so the harness can be exercised offline
llm, model = ScriptedLLM(responder=lambda messages, kw: keyword_label(messages[-1]["content"])), None
extra = {}
else:
provider, model = resolve()
llm, extra = chat, {"provider": provider, "model": model}
test = load_tickets("test")
preds, usages = run(llm, test, **extra)
k = sum(p == t.gold["category"] for p, t in zip(preds, test))
lo, hi = wilson(k, len(test))
print(f"{'stand-in (keyword rules)' if stand_in else model}: {k}/{len(test)} = {k / len(test):.0%} "
f"(95% CI {lo:.0%} to {hi:.0%}); unparseable replies: {preds.count(None)}")
tokens_in, tokens_out = sum(u.input_tokens for u in usages), sum(u.output_tokens for u in usages)
print(f"tokens: {tokens_in} in, {tokens_out} out for {len(test)} tickets")
if model in PRICES:
dollars = sum(cost_usd(u, model) for u in usages)
print(f"cost: {dollars:.5f} USD, so about {dollars / len(test) * 10_000:.2f} USD per 10,000 tickets")
# Distillation: label the dev tickets with the stronger model, keep provenance on every row.
if not stand_in:
dev = load_tickets("dev")
teacher, _ = run(llm, dev, **extra)
rows = [{"id": t.id, "subject": t.subject, "body": t.body, "category": p, "gold": t.gold["category"],
"teacher": model, "labelled_on": date.today().isoformat(),
"terms_checked": "record the provider terms URL and version you checked"}
for t, p in zip(dev, teacher)]
agree = sum(r["category"] == r["gold"] for r in rows)
out = DATA_OUT / f"distilled_{int(time.time())}.jsonl"
out.write_text("".join(json.dumps(r, ensure_ascii=False) + "\n" for r in rows))
print(f"teacher agrees with the human label on {agree}/{len(rows)} dev tickets; wrote {out.name}")
Code explained
- In simple words: ask a real hosted model to triage the 24 test tickets zero-shot, score it with the same interval, and price the run.
- What happens:
SYSTEMlists the six categories and asks for one.parse_labeltakes the first category name in the reply (a reasoning model may add words; anything unparseable is counted as wrong, not silently dropped).runsends one chat request per ticket withtemperature=0.0andmax_tokens=400(headroom for a reasoning model's hidden tokens). With a key, the script also labels the 48 dev tickets with the model and writesdata/m09/distilled_<time>.jsonl, each row stamped with the teacher model, the date, and a reminder to record which terms you checked. With--stand-in, aScriptedLLManswers with the keyword rules from the previous section. That exercises the plumbing (parsing, scoring, token counting) without any model. - Comes out: no API key was available while this module was built, so here is the stand-in run (
python -m examples.m09_hosted_baseline --stand-in). This is ScriptedLLM output, not model output: the 15/24 is the keyword rules, and the token counts are the stand-in's own estimate.
stand-in (keyword rules): 15/24 = 62% (95% CI 43% to 79%); unparseable replies: 0
tokens: 1745 in, 45 out for 24 tickets
With a key (export GROQ_API_KEY=... then python -m examples.m09_hosted_baseline), you get the real zero-shot accuracy of openai/gpt-oss-120b, its token usage, and its cost. Put that number in your eval plan next to the others. Do not assume it; 24 tickets in five languages is exactly the kind of set where a strong model can still miss a few.