Part E: Adaptation Track
Scope. Find a narrow, well-labelled Brightlane task where prompting stops improving, build a fine-tuning dataset, adapt a model, and prove (or disprove) lift over the prompted baseline while measuring what the model lost elsewhere. General-capability regression means the tuned model got worse at things it could do before; you measure it on text or tasks unrelated to the fine-tune. Out of scope: training a model from scratch, and any claim of lift without a paired test on a frozen test split.
Milestones
| Week | Adaptation track |
|---|---|
| 1 | Choose the task and metric; measure the prompted baseline with 2 or 3 prompt variants to show the plateau; freeze test |
| 2 | Build the dataset (sources, labels, dedup against test, a label-quality check); pick the regression probes |
| 3 | First fine-tune; 50+ real failures of the prompted and tuned models tagged by cause |
| 4 | Recipe selection on a validation slice only (learning rate, steps, replay, data mix); log every run |
| 5 | Final retrain, one test run, paired tests, regression report, serving cost and latency of the tuned model |
| 6 | Evidence pack with the decision: ship, do not ship, or collect more data |
Starter: a TinyLM fine-tune against its prompted baseline
The task is ticket category. The prompted baseline scores each of the six labels after a prompt that lists them and picks the most probable, using the mean log-probability per label token. The starter scores labels rather than generating them because greedy generate() with an allowed mask falls back to the lowest allowed token id whenever the top token is disallowed, which would silently bias answers toward one label. The fine-tune updates every weight on the dev tickets, with loss only on the label tokens. Two recipes compete on a validation slice of dev, under a regression budget: held-out perplexity on TinyLM's own pretraining text may rise at most 10%. Then the winner is retrained on all of dev and scored on test once. The whole run takes about a minute on one CPU thread.
examples/cap_adapt_starter.py
# examples/cap_adapt_starter.py
"""Adaptation track starter: does fine-tuning beat the prompted baseline?
Task: ticket category (6 classes). Model: TinyLM (1.07M parameters).
1. Prompted baseline: after a prompt that lists the categories, score every
label by its mean log-probability per token and pick the best (no training).
2. Choose a recipe WITHOUT touching test: split the 48 dev tickets into 32 for
training and 16 for validation, try each recipe, keep the best validation
accuracy among recipes whose general-capability regression stays in budget.
3. Retrain the chosen recipe on all 48 dev tickets, score the 24 test tickets
once, and compare with the prompted and keyword baselines (paired sign test).
4. General-capability regression: perplexity on held-out support dialogue the
base model was pretrained on, before and after.
Full fine-tuning (every weight), loss on the label tokens only. Saves to
models/cap-triage/<recipe>, never over models/tinylm-base. About 1 to 2 minutes on 1 CPU thread.
"""
from __future__ import annotations
import sys
import time
from pathlib import Path
import torch
from supportdesk.data import CATEGORIES, DATA_DIR, Ticket, load_tickets
from supportdesk.tinylm import batches, load, loss_on, perplexity, save
sys.path.insert(0, str(Path(__file__).resolve().parent))
from cap_baseline import keyword_triage, sign_test_p, wilson # noqa: E402
torch.set_num_threads(1)
OUT_DIR = Path("models/cap-triage")
HEADER = "Categories: billing, cancellation, account_access, bug, how_to, feature_request.\n"
STEPS, BATCH = 150, 8
RECIPES = {"no_replay": {"lr": 1e-3, "replay": 0.0}, "replay": {"lr": 1e-3, "replay": 1.0}}
MAX_PPL_INCREASE = 0.10 # regression budget: held-out perplexity may rise at most 10%
def prompt_for(text: str, tok, room: int) -> list[int]:
"""Header + ticket (cut from the end if needed) + 'Category:' within `room` tokens."""
head, tail = tok.encode(HEADER + "Ticket: ").ids, tok.encode("\nCategory:").ids
body = tok.encode(text.replace("\n", " ")).ids[: room - len(head) - len(tail)]
return head + body + tail
def label_ids(tok) -> dict[str, list[int]]:
return {c: tok.encode(" " + c + "\n").ids for c in CATEGORIES}
@torch.no_grad()
def label_scores(model, tok, text: str) -> dict[str, float]:
"""Mean log-probability per token of each label after the prompt.
Dividing by the label length matters: summed log-probabilities favour short
labels, and prompted TinyLM then answers 'billing' (2 tokens) for every ticket.
"""
labels = label_ids(tok)
room = model.cfg.context - max(len(v) for v in labels.values())
p = prompt_for(text, tok, room)
scores = {}
for cat, ids in labels.items():
logp = torch.log_softmax(model(torch.tensor([p + ids])[:, :-1])[0], dim=-1)
scores[cat] = sum(logp[len(p) - 1 + i, t].item() for i, t in enumerate(ids)) / len(ids)
return scores
def classify(model, tok, text: str) -> str:
scores = label_scores(model, tok, text)
return max(scores, key=scores.get)
def predictions(model, tok, tickets: list[Ticket]) -> list[str]:
model.eval()
return [classify(model, tok, t.text) for t in tickets]
def training_batch(model, tok, tickets, gen: torch.Generator):
"""Random batch of (x, y, mask); the mask keeps loss on the label tokens only."""
labels = label_ids(tok)
room = model.cfg.context - max(len(v) for v in labels.values())
pick = torch.randint(0, len(tickets), (BATCH,), generator=gen).tolist()
seqs = [(prompt_for(tickets[i].text, tok, room), labels[tickets[i].gold["category"]]) for i in pick]
width = max(len(p) + len(lab) for p, lab in seqs) - 1
x = torch.zeros(BATCH, width, dtype=torch.long)
y = torch.zeros(BATCH, width, dtype=torch.long)
m = torch.zeros(BATCH, width)
for row, (p, lab) in enumerate(seqs):
ids = p + lab
x[row, : len(ids) - 1] = torch.tensor(ids[:-1])
y[row, : len(ids) - 1] = torch.tensor(ids[1:])
m[row, len(p) - 1: len(ids) - 1] = 1.0
return x, y, m
def corpus_split() -> tuple[str, str]:
"""Pretraining corpus as (train part, held-out 6,000 characters from the last 5%)."""
text = (DATA_DIR / "corpus.txt").read_text(encoding="utf-8")
cut = int(len(text) * 0.95)
return text[:cut], text[cut:][:6000]
def finetune(train: list[Ticket], lr: float, replay: float, replay_ids: torch.Tensor):
"""Full fine-tuning from the base model; replay adds `replay` x the loss on pretraining text."""
torch.manual_seed(0)
model, tok = load()
gen = torch.Generator().manual_seed(0)
replay_iter = batches(replay_ids, model.cfg.context, 4, torch.Generator().manual_seed(1))
opt = torch.optim.AdamW(model.parameters(), lr=lr, weight_decay=0.0)
model.train()
for _ in range(STEPS):
x, y, m = training_batch(model, tok, train, gen)
loss = loss_on(model, x, y, m)
if replay:
rx, ry = next(replay_iter)
loss = loss + replay * loss_on(model, rx, ry)
opt.zero_grad()
loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
opt.step()
model.eval()
return model, tok
def correct(preds: list[str], tickets: list[Ticket]) -> list[bool]:
return [p == t.gold["category"] for p, t in zip(preds, tickets)]
def main() -> None:
started = time.perf_counter()
dev, test = load_tickets("dev"), load_tickets("test")
fit, val = [t for i, t in enumerate(dev) if i % 3], [t for i, t in enumerate(dev) if i % 3 == 0]
corpus_train, heldout = corpus_split()
base, tok = load()
replay_ids = torch.tensor(tok.encode(corpus_train).ids)
ppl0 = perplexity(base, tok, heldout)
n = len(test)
kw = [keyword_triage(t.text)[0] == t.gold["category"] for t in test]
prompted_preds = predictions(base, tok, test)
prompted = correct(prompted_preds, test)
print(f"keyword baseline test {sum(kw):2}/{n} 95% CI {wilson(sum(kw), n)}")
print(f"prompted TinyLM test {sum(prompted):2}/{n} 95% CI {wilson(sum(prompted), n)} "
f"held-out ppl {ppl0:.3f} (labels predicted: {sorted(set(prompted_preds))})")
print(f"\nrecipe selection: train on {len(fit)} dev tickets, validate on {len(val)} (test untouched)")
best = None
for name, cfg in RECIPES.items():
model, _ = finetune(fit, cfg["lr"], cfg["replay"], replay_ids)
k = sum(correct(predictions(model, tok, val), val))
ppl = perplexity(model, tok, heldout)
in_budget = ppl / ppl0 - 1 <= MAX_PPL_INCREASE
print(f" {name:10} lr {cfg['lr']} replay {cfg['replay']}: val {k:2}/{len(val)} "
f"ppl {ppl0:.3f} -> {ppl:.3f} ({(ppl / ppl0 - 1) * 100:+.0f}%) "
f"{'within' if in_budget else 'OVER'} regression budget")
if in_budget and (best is None or k > best[1]):
best = (name, k)
if best is None:
print("no recipe stayed within the regression budget; stop and rethink the data or the recipe")
return
name, cfg = best[0], RECIPES[best[0]]
print(f"\nchosen recipe: {name}; retrain on all {len(dev)} dev tickets, then score test once")
model, _ = finetune(dev, cfg["lr"], cfg["replay"], replay_ids)
tuned = correct(predictions(model, tok, test), test)
ppl1 = perplexity(model, tok, heldout)
print(f"fine-tuned TinyLM test {sum(tuned):2}/{n} 95% CI {wilson(sum(tuned), n)} "
f"held-out ppl {ppl0:.3f} -> {ppl1:.3f} ({(ppl1 / ppl0 - 1) * 100:+.0f}%)")
for label, other in (("prompted TinyLM", prompted), ("keyword baseline", kw)):
wins = sum(a and not b for a, b in zip(tuned, other))
losses = sum(b and not a for a, b in zip(tuned, other))
print(f" vs {label:16} wins {wins:2} losses {losses:2} one-sided sign test p = {sign_test_p(wins, losses):.3f}")
save(model, tok, OUT_DIR / name, meta={"task": "category", "steps": STEPS, "batch": BATCH, **cfg,
"train": "dev split (48)", "base": "tinylm-base"})
print(f"saved to {OUT_DIR / name}; wall time {time.perf_counter() - started:.0f} s")
if __name__ == "__main__":
main()
Code explained
- In simple words: tutor a tiny model on 48 labelled tickets, pick the tutoring method with a practice quiz, and only then sit the real exam; also check it has not forgotten how to talk.
- What happens:
prompt_forbuilds header, ticket, and "Category:" as token ids, cutting the ticket so the prompt plus the longest label fits the 128-token context.label_idsencodes each label as " label\n" (between 2 tokens forbillingand 7 forfeature_request).label_scoresruns the prompt plus each label through the model and averages the label tokens' log-probabilities. Averaging matters: with sums, prompted TinyLM pickedbilling, the shortest label, for all 24 test tickets (5 correct), because every extra token costs log-probability.classifyandpredictionstake the best-scoring label for each ticket.training_batchsamples 8 tickets and builds inputs, targets, and a mask that is 1 only on the label tokens.corpus_splitreturns the pretraining corpus minus its last 5%, and 6,000 characters of that held-out tail for the regression probe.finetunestarts from the base model every time, trains 150 steps with AdamW and gradient clipping, and with replay adds the ordinary next-token loss on a batch of pretraining text each step (Module 9's defense against forgetting).mainscores both baselines, runs each recipe on 32 dev tickets and validates on the other 16, keeps the best validation accuracy among recipes within the regression budget, retrains it on all 48, scores test once, runs paired sign tests against the prompted and keyword baselines, and saves tomodels/cap-triage/<recipe>.
- Comes out: with seeds fixed and one thread, the numbers repeat on this machine; they may differ slightly elsewhere. Wall time varies with load.text
keyword baseline test 16/24 95% CI (0.467, 0.82) prompted TinyLM test 6/24 95% CI (0.12, 0.449) held-out ppl 1.929 (labels predicted: ['billing', 'cancellation', 'how_to']) recipe selection: train on 32 dev tickets, validate on 16 (test untouched) no_replay lr 0.001 replay 0.0: val 3/16 ppl 1.929 -> 24.922 (+1192%) OVER regression budget replay lr 0.001 replay 1.0: val 6/16 ppl 1.929 -> 2.011 (+4%) within regression budget chosen recipe: replay; retrain on all 48 dev tickets, then score test once fine-tuned TinyLM test 3/24 95% CI (0.043, 0.31) held-out ppl 1.929 -> 2.023 (+5%) vs prompted TinyLM wins 3 losses 6 one-sided sign test p = 0.910 vs keyword baseline wins 2 losses 15 one-sided sign test p = 1.000 saved to models/cap-triage/replay; wall time 55 s
This is a negative result, and it is the correct one.
- The regression budget did its job. Without replay, held-out perplexity rose from 1.929 to 24.922 (about 13 times): the model forgot its pretraining text while learning labels. With replay the rise was 4% on the 32-ticket run and 5% on the final one, inside the budget.
- The fine-tune did not beat the prompted baseline. 3 of 24 against 6 of 24, with 3 wins and 6 losses; p = 0.91 for "better than prompted". Both are far below the keyword rules at 16 of 24.
- Why: the model memorizes. In extra runs made while writing this (on the validation slice only, never on test), every recipe reached 32 of 32 on its own training tickets while validation stayed at 2 to 7 of 16: 300 steps gave 6, lighter replay 5, a doubled learning rate 2, and adding the 100 synthetic tickets to training gave 5 (150 steps) and 7 (300 steps). A 1-million-parameter model pretrained on 0.7 MB of templated text has too little general language to generalize from 32 to 48 examples. Module 9 reached the same "do not ship" conclusion from a different angle.
- What it proves about the method: recipe selection never saw test, the regression probe caught catastrophic forgetting, and the paired test refused to call noise a win. Point the same script at a real model and those three pieces are your evidence.
For the real capstone, use a small open-weight model with LoRA (Hugging Face PEFT and TRL on a GPU, or a managed fine-tuning service; Module 9 covers both), a dataset of a few hundred to a few thousand labelled tickets, and regression probes that matter to Brightlane: a general instruction-following set, the other triage fields, and multilingual tickets. Keep the structure of this starter: validation-only selection, a regression budget written down in week 1, one test run, paired tests.
| Situation | Use this | Why |
|---|---|---|
| Prompted accuracy still rises with better prompts or examples | Keep prompting (Modules 4 and 5) | Fine-tuning a moving target wastes data and time |
| Prompting plateaus, labels are plentiful, output format is fixed | LoRA fine-tune of a small model, with replay or a data mix | Cheaper serving and higher accuracy on the narrow task |
| Fewer than about 100 labels | Few-shot prompting or a classical classifier; collect labels from reviewed outputs | The TinyLM result above: small data gets memorized |
| The model must keep broad skills | Regression probes plus replay, and a budget that blocks promotion | Forgetting is silent unless you measure it |
Required evidence
- The prompted baseline across prompt variants (the plateau) and the strongest cheap baseline, with intervals.
- The dataset card: sources, size, label process, dedup against test, known label noise.
- Every training run in a log, with the recipe chosen on validation only; the final single test run with paired tests.
- The regression report against the budget; serving cost and latency of tuned versus prompted; the B2 report and B4 account (negative results are expected here).
Grading rubric
| Criterion | Weight | Excellent | Adequate | Missing |
|---|---|---|---|---|
| Plateau and baselines | 15% | Several prompt variants show the ceiling; strongest cheap baseline included | One prompt | No prompted baseline |
| Dataset | 20% | Documented, deduplicated against test, label quality checked | Documented | Undocumented |
| Experimental hygiene | 25% | Selection on validation only, one test run, paired tests, all runs logged | Test used once, no paired test | Tuned on test |
| Regression measurement | 20% | Budget set in advance, probes relevant to the product, replay or mix compared | Perplexity only | None |
| Error analysis, cost, what did not work | 20% | 50+ failures by cause; serving cost measured; honest negative results | Partial | Missing |