Part D: Frontier
The models you build on will change several times during the life of the Brightlane assistant. This part is about reading that change clearly: which trends are measured, what they mean for a support desk, and how to tell capability from marketing. Every trend number below comes from a cited source; follow the links, because these figures are revised often.
D1. Longer contexts and cheaper inference, and what they unlock
Two trends are well documented by Epoch AI, an independent research group that tracks AI progress:
- Context windows. In "LLMs now accept longer inputs, and the best models can use them more effectively" (Burnham and Adamczewski, 25 June 2025, epoch.ai/data-insights/context-windows), the longest context windows grew about 30x per year since mid-2023, and on two long-context benchmarks (Fiction.liveBench and MRCR) the input length at which top models reach 80% accuracy rose by over 250x in nine months. Advertised length and usable length both grew, but they are not the same number (Module 2's effective context). For a current example, Google's developer documentation lists a 1 million token input limit for
gemini-3.5-flash(ai.google.dev). - Price. In "LLM inference prices have fallen rapidly but unequally across tasks" (Cottier, Snodin, Owen, and Adamczewski, 12 March 2025, epoch.ai/data-insights/llm-inference-price-trends), the price of reaching GPT-4's level on PhD-level science questions fell about 40x per year, with the rate ranging from 9x to 900x per year depending on the task and threshold. The authors note that the fastest drops happened in the most recent year, so it is less clear they will persist. Read the unit carefully: this is the price of a fixed level of capability, not the price of the newest frontier model, which has not fallen at that rate.
What does that mean for Brightlane? The next script turns the trends into Brightlane arithmetic and also prices one thing longer contexts unlock: skipping retrieval entirely.
# examples/m14_frontier_math.py
"""Frontier: turning published trend numbers and a release announcement into Brightlane arithmetic.
Inputs are cited in the module text: Epoch AI price trends (9x to 900x per year
for a FIXED capability level), METR time-horizon doubling times, and the Gemini
3.8 Flash announcement (2 September 2026) with its introductory price.
"""
from __future__ import annotations
import math
from supportdesk.data import load_articles
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, Price, cost_usd
from supportdesk.tokens import count_tokens
def cost(price: Price, usage: Usage) -> float:
"""Same formula as supportdesk.pricing.cost_usd, for a price not (yet) in the PRICES table."""
return (usage.input_tokens * price.input + usage.output_tokens * price.output) / 1_000_000
if __name__ == "__main__":
kb_tokens = sum(count_tokens(a.title + "\n" + a.body) for a in load_articles())
print(f"whole help center: {kb_tokens} tokens; per 100k tickets with the whole help center in every prompt:")
for model in ("openai/gpt-oss-120b", "gemini-3.5-flash"):
plain = cost_usd(Usage(input_tokens=kb_tokens + 150, output_tokens=350), model) * 1e5
cached = cost_usd(Usage(input_tokens=kb_tokens + 150, cached_tokens=kb_tokens, output_tokens=350), model) * 1e5
print(f" {model:22} {plain:8.2f} USD, with the help center cached {cached:8.2f} USD")
monthly_now = 194.85 # gemini-3.5-flash triage bill for 100k tickets, from m14_triage_at_scale.py
print("same-capability price decline applied to a 194.85 USD/month bill")
for per_year in (9, 40, 900):
print(f" {per_year:>3}x per year: after 12 months {monthly_now / per_year:8.2f} USD, "
f"after 6 months {monthly_now / math.sqrt(per_year):8.2f} USD")
print("\nMETR 50% time horizon: months until a horizon grows from 14.5 h to a 40 h work week")
for label, days in (("196-day doubling (2019 to 2025 trend)", 196), ("89-day doubling (since 2024, TH1.1)", 89)):
doublings = math.log2(40 / 14.5)
print(f" {label}: {doublings:.2f} doublings = {doublings * days / 30.4:.1f} months "
f"(a 2x measurement error moves this by {days / 30.4:.1f} months)")
print("\nper-task cost of one triage call: 141 input tokens, 193 output tokens (incl. 150 assumed reasoning)")
old = PRICES["gemini-3.5-flash"]
intro, later = Price(0.75, 3.75, 0.0), Price(1.50, 7.50, 0.0) # Gemini 3.8 Flash, from the announcement
base = Usage(input_tokens=141, output_tokens=193)
print(f" gemini-3.5-flash {cost(old, base) * 1e5:7.2f} USD per 100k")
for name, price in (("3.8 Flash intro price", intro), ("3.8 Flash from Jan 2027", later)):
# How many times more output tokens can the new model spend before it costs more per task?
ratio = (cost(old, base) - 141 * price.input / 1e6) / (193 * price.output / 1e6)
print(f" {name:24} {cost(price, base) * 1e5:7.2f} USD per 100k at equal tokens; "
f"breaks even at {ratio:.2f}x the output tokens")
Code explained
- In simple words: a calculator that converts headline trend numbers and a price sheet into this project's monthly bill and planning dates.
- What happens:
- It counts the whole help center in tokens and prices putting all of it in every prompt, with and without prompt caching of that stable block (Module 7's cache-aware ordering).
- It applies the 9x, 40x, and 900x per year declines to A1's gemini-3.5-flash triage bill. Six months is the square root of a yearly factor.
- It converts METR doubling times (D3) into months, and shows how much a factor-of-2 measurement error shifts the answer (exactly one doubling time).
- It prices the A1 triage call on Gemini 3.8 Flash at its introductory and later prices (D5) and solves for the output-token multiple at which the new model stops being cheaper per task.
- Comes out:text
whole help center: 1205 tokens; per 100k tickets with the whole help center in every prompt: openai/gpt-oss-120b 41.32 USD, with the help center cached 41.32 USD gemini-3.5-flash 518.25 USD, with the help center cached 355.57 USD same-capability price decline applied to a 194.85 USD/month bill 9x per year: after 12 months 21.65 USD, after 6 months 64.95 USD 40x per year: after 12 months 4.87 USD, after 6 months 30.81 USD 900x per year: after 12 months 0.22 USD, after 6 months 6.50 USD METR 50% time horizon: months until a horizon grows from 14.5 h to a 40 h work week 196-day doubling (2019 to 2025 trend): 1.46 doublings = 9.4 months (a 2x measurement error moves this by 6.4 months) 89-day doubling (since 2024, TH1.1): 1.46 doublings = 4.3 months (a 2x measurement error moves this by 2.9 months) per-task cost of one triage call: 141 input tokens, 193 output tokens (incl. 150 assumed reasoning) gemini-3.5-flash 194.85 USD per 100k 3.8 Flash intro price 82.95 USD per 100k at equal tokens; breaks even at 2.55x the output tokens 3.8 Flash from Jan 2027 165.90 USD per 100k at equal tokens; breaks even at 1.20x the output tokensThe first lines are the practical unlock. Brightlane's entire help center is 1,205 tokens. Putting all of it in every prompt costs about 41 USD per 100,000 tickets on gpt-oss-120b, and it removes the retrieval misses you measured in A3 (and Module 7's 52 of 62 top-1 hit rate) at a stroke. For a 12-article help center, long context beats retrieval; for 12,000 articles it does not, and retrieval comes back. The Gemini line shows why prompt caching matters for this layout: the cached help center cuts the bill by about a third. The trend lines say the same triage capability that costs 195 USD a month today would cost somewhere between 22 USD and 22 cents a year from now if those trends hold for this task, a range so wide that the only safe plan is to re-measure prices quarterly, not to forecast them.
| Situation | Use this | Why |
|---|---|---|
| Knowledge base of a few thousand tokens | Put it all in the prompt, cached | No retrieval misses; cost is small and falling |
| Knowledge base far beyond the effective context | Retrieval (Module 7) | Cost and effective-context limits still bind |
| Medium size, frequent updates | Retrieval with a generous k, measured against the whole-corpus baseline | Keep whichever wins on your eval, not on the trend chart |
D2. Test-time compute scaling
Test-time compute is computation spent while answering, not while training: longer hidden reasoning, several samples with a vote, or a search guided by a verifier (Modules 3 and 5). OpenAI's o1 announcement ("Learning to reason with LLMs", 12 September 2024, openai.com) reported that performance "consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute)." Snell, Lee, Xu, and Kumar ("Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters", August 2024, arXiv:2408.03314) found that allocating test-time compute adaptively per prompt improved efficiency by more than 4x over a best-of-N baseline, and that on problems where a smaller model already had some success, extra test-time compute could outperform a 14x larger model in a compute-matched comparison.
For an application builder, three consequences follow:
- Price per task, not per token. Reasoning tokens bill as output (Module 2). A model with a lower price per token can cost more per ticket if it thinks longer. The A1 cost table only held because we wrote the reasoning-token assumption down.
- Effort is a dial you tune per route. Triage of a clear ticket needs little thought; a multi-step refund investigation may need more. Measure accuracy and cost at each
reasoning_effortlevel on your eval set and pick per route (Module 3). - Verifiers multiply the value of compute. Extra samples help most when something can check them, which is why the SQL pattern in A4 pairs so well with more attempts.
D3. Agentic capability trajectories
The best-known measurement of agent progress comes from METR, a nonprofit that evaluates frontier models. Their 50% time horizon is the length of task (measured by how long it takes skilled humans) that a model completes successfully half the time, on a suite of mostly software and research tasks.
- The original paper ("Measuring AI Ability to Complete Long Software Tasks", 19 March 2025, metr.org) found the horizon doubling about every 7 months over six years, with Claude 3.7 Sonnet at about one hour.
- Time Horizon 1.1 (29 January 2026, metr.org) expanded the suite from 170 to 228 tasks and reported a doubling time of 196 days over 2019 to 2025, 131 days since 2023, and 89 days since 2024. Its top estimates were Claude Opus 4.5 at 320 minutes (95% interval 170 to 729) and GPT-5 at 214 minutes.
- On 20 February 2026 METR estimated Claude Opus 4.6 at about 14.5 hours (95% interval 6 to 98 hours), adding that the measurement is "extremely noisy because our current task suite is nearly saturated" (METR on X). METR's live page (metr.org/time-horizons, last updated 8 May 2026 when this was written) has since added more models; check it for current figures rather than trusting any number printed here.
METR is also explicit about what the metric is not ("Clarifying limitations of time horizon", 22 January 2026, metr.org): it is not how long a model can work unattended, the error bars are roughly a factor of 2 in each direction, horizons differ between domains by orders of magnitude (visual computer-use tasks are 40x to 100x lower), and the task suite is mostly software.
So what should Brightlane do with this? The frontier math output above shows that a factor-of-2 measurement error moves any extrapolation by a full doubling time, 3 to 6 months. The durable lessons are about design, not dates. A 50% success rate is a research metric; Part C showed a customer-facing action needs 84% to 99.6% depending on the cost of an error, and tasks at the edge of a model's horizon are exactly where it fails most. Agents will keep taking on longer tasks, so design the approval gates from B1 around the consequence of each action, and they will stay correct however capable the model becomes.
D4. Small models and on-device inference
At the other end of the scale, small models run on phones and laptops with no network, no per-token price, and no data leaving the device. Apple's documentation for its on-device foundation model describes "a compact, approximately 3-billion-parameter model" compressed to 2 bits per weight, and states it "is not designed to be a chatbot for general world knowledge", while it does well at summarization, extraction, text understanding, and short dialog (Apple Machine Learning Research). That list is the pattern catalogue from Part A with the knowledge-heavy patterns removed.
TinyLM, at 1.07 million parameters, is the extreme small case: about 3,000 times smaller than that on-device model. Here is its real CPU latency.
# examples/m14_tinylm_latency.py
"""Frontier: small models and on-device inference, with TinyLM as the extreme small case.
Measures real CPU latency for a 1.07M-parameter model: prefill time, decode time
per token and tokens per second (wall clock and CPU time), plus memory footprint.
"""
from __future__ import annotations
import statistics
import time
import torch
from supportdesk import tinylm
PROMPT = "Customer: I was charged twice for the Team plan this month.\nAgent:"
def measure(model, tokenizer, threads: int, runs: int = 9) -> tuple[float, float, float]:
torch.set_num_threads(threads)
params = tinylm.SamplingParams(max_new_tokens=40, temperature=0)
walls, cpus, prefills = [], [], []
for i in range(runs):
cpu0, wall0 = time.process_time(), time.perf_counter()
gen = tinylm.generate(model, tokenizer, PROMPT, params)
cpu, wall = time.process_time() - cpu0, time.perf_counter() - wall0
if i >= 2: # the first two runs warm up; keep the rest
walls.append(wall * 1000 / len(gen.token_ids))
cpus.append(cpu * 1000 / len(gen.token_ids))
prefills.append(gen.prefill_ms)
return statistics.median(prefills), statistics.median(walls), statistics.median(cpus)
if __name__ == "__main__":
model, tokenizer = tinylm.load()
n = model.num_parameters()
print(f"parameters {n:,}; fp32 weights {n * 4 / 1e6:.1f} MB; at 2 bits per weight {n * 2 / 8 / 1e6:.2f} MB")
print(f"prompt: {len(tokenizer.encode(PROMPT).ids)} tokens, generating 40")
with open("/proc/loadavg") as f:
print("machine load average (1 min):", f.read().split()[0])
for threads in (1, 2):
prefill, wall, cpu = measure(model, tokenizer, threads)
print(f"threads={threads}: prefill {prefill:.1f} ms, whole generation {wall:.1f} ms/token wall "
f"({1000 / wall:.0f} tokens/s), {cpu:.1f} ms/token CPU")
torch.set_num_threads(1)
gen = tinylm.generate(model, tokenizer, PROMPT, tinylm.SamplingParams(max_new_tokens=40, temperature=0))
print("output:", repr(gen.text[:110]))
print("capital of France:", repr(tinylm.generate(model, tokenizer, "The capital of France is",
tinylm.SamplingParams(max_new_tokens=16, temperature=0)).text))
Code explained
- In simple words: time a very small model on this machine's CPU and look at what it produces.
- What happens:
- It computes the weight size at 32-bit floats and at 2 bits per weight, the compression level Apple reports.
measureruns greedy generation of 40 tokens nine times, discards two warm-up runs, and records the median prefill time plus wall-clock and CPU time per generated token. CPU time (time.process_time) counts only time this process actually computed, so it is less sensitive to other programs on the machine.- It measures once with 1 thread and once with 2, and prints the machine's load average so the numbers can be read in context.
- It prints a real continuation of a support prompt and of "The capital of France is".
- Comes out: this run shared a 2-core machine with other heavy jobs (load average about 17), so treat the wall-clock numbers as a snapshot.text
parameters 1,071,872; fp32 weights 4.3 MB; at 2 bits per weight 0.27 MB prompt: 16 tokens, generating 40 machine load average (1 min): 17.07 threads=1: prefill 1.6 ms, whole generation 4.5 ms/token wall (220 tokens/s), 1.0 ms/token CPU threads=2: prefill 896.1 ms, whole generation 131.0 ms/token wall (8 tokens/s), 28.6 ms/token CPU output: ' Team plan this month. Can you refund the duplicate?\nAgent (Eli): Sorry about the duplicate charge. Duplicate ' capital of France: ' not loading. Is there an outage?\nAgent (Nia): Please check status'With one thread, TinyLM generates at about 220 tokens a second of wall time and needs about 1 ms of CPU per token. The weights are 4.3 MB, or about a quarter of a megabyte at 2 bits per weight. With two threads on the busy machine it collapsed to 8 tokens a second, because the two threads kept waiting for each other and for other processes. That is a real on-device lesson: a phone or laptop runs your model alongside everything else, so measure under realistic load and prefer fewer threads when the CPU is contended. The outputs are the other lesson. The support continuation is fluent because TinyLM memorized its templated corpus; the France prompt produces confident support-desk text, which is Module 1's hallucination demo in one line. A small model is fast and private, and it knows only what its training data taught it.
| Situation | Use this | Why |
|---|---|---|
| Privacy-sensitive or offline, narrow task (classify, extract, short rewrite) | A small on-device model, evaluated on your task | No data leaves the device; latency is local |
| Tasks that need world knowledge or long reasoning | A hosted frontier or mid-tier model | Small models do not hold the knowledge |
| High volume, narrow task, a big model works but costs too much | Distill or fine-tune a small model (Module 9) | Keeps quality on the narrow task at a fraction of the cost |
D5. Reading releases critically: separating capability from marketing
Model announcements are written to sell. That does not make them false, but it means the questions you care about are often not the ones the post answers. Apply the same checklist to every release. Here it is applied to a real one: Google's "Introducing Gemini 3.8 Flash and 3.8 Flash Cyber" (Tulsee Doshi and Raluca Ada Popa, 2 September 2026, blog.google), as read on 21 September 2026.
| Question | What the announcement says | What it means for Brightlane |
|---|---|---|
| Which benchmarks, and are they like my task? | 54.9% on HLE-Verified; claims on DeepSWE v1.1 (long-horizon software engineering), Vals Finance Agent V2, and Harvey's Legal Agent Benchmark | None is support triage or grounded answering. Expect nothing until our own eval runs |
| Compared with what, under which settings? | "Outperforms most larger frontier models" on DeepSWE; "outperforms 3.7 Flash and other frontier models" on two agent benchmarks | "Most" and unnamed competitors; the post links no methodology page, effort level, or number of attempts |
| What does it cost per task, not per token? | 0.75 USD input and 3.75 USD output per million tokens, an introductory rate through 31 December 2026, then 1.50 and 7.50 | The frontier math shows the intro price breaks even with 3.5 Flash at 2.55x the output tokens, but the later price at only 1.20x |
| Does it change how many tokens it uses? | It "might use more tokens to maximize performance, especially at higher effort levels" | Per-token price is not per-task price. Measure tokens per ticket on our eval |
| Context, latency, availability? | No context window or latency figures in the post; available in the Gemini API via Google AI Studio, Antigravity, Android Studio, Gemini Enterprise, the Gemini app, and more | Look up the model card and measure latency ourselves |
| Who measured it? | The vendor | Wait for independent measurements (for example Artificial Analysis, METR, Epoch AI) and, above all, run our own |
Only the last question has a universal answer, and it is the one you control. Rerun the Module 10 eval suite with the candidate model (D6 automates the decision), measure tokens and latency per ticket, and let those numbers decide. A benchmark gain on legal agents says nothing about Maya's refund tickets.
D6. Keeping a system adaptable as models change under it
Models are deprecated, repriced, and replaced (Module 13's forced migrations). A system stays adaptable when three things are true, all of which this course has already built:
- Prompts are versioned data with a content hash, not strings scattered through code (Modules 4 and 5).
- Each feature names its provider, model, and prompt version in one place, behind the provider-neutral
supportdesk.llm.chat(Modules 1 and 13). - No change ships without the same eval gate: golden set, invalid-output count, cost and latency budgets (Module 10).
# examples/m14_adaptable.py
"""Frontier: keeping a system adaptable as models change underneath it.
Three habits from earlier modules, in one small file: prompts are versioned
data with a content hash (Module 4 and 5), each feature names its model in one
config (Module 13), and no model or prompt change ships without passing the
same golden-set gate (Module 10).
"""
from __future__ import annotations
import hashlib
import json
from dataclasses import dataclass
from pydantic import ValidationError
from examples.m14_common import fmt_rate
from supportdesk.data import load_tickets
from supportdesk.llm import ChatResult, Usage
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.schemas import Triage
from supportdesk.stand_in import ScriptedLLM
PROMPTS = {
("triage", "v3"): "Classify the ticket. Reply with only JSON: category, priority, language, summary, needs_human.",
("triage", "v4"): "Classify the Brightlane support ticket. Reply with only a JSON object with keys category, "
"priority, language, summary, needs_human. No prose, no code fences.",
}
def prompt_hash(feature: str, version: str) -> str:
return hashlib.sha256(PROMPTS[(feature, version)].encode()).hexdigest()[:8]
@dataclass(frozen=True)
class FeatureConfig:
feature: str
provider: str
model: str
prompt_version: str
max_cost_per_1k_usd: float
max_p95_latency_ms: float
PRODUCTION = FeatureConfig("triage", "groq", "openai/gpt-oss-120b", "v3", 0.50, 3000)
def run_gate(config: FeatureConfig, llm, tickets) -> dict:
correct = invalid = 0
priced_as = config.model if config.model in PRICES else "openai/gpt-oss-120b" # add new models to PRICES first
latencies, cost = [], 0.0
for t in tickets:
messages = [{"role": "system", "content": PROMPTS[(config.feature, config.prompt_version)]},
{"role": "user", "content": t.text}]
result = llm(messages, provider=config.provider, model=config.model, temperature=0.0)
latencies.append(result.latency_ms)
cost += cost_usd(result.usage, priced_as)
try:
correct += Triage.model_validate_json(result.text).category == t.gold["category"]
except ValidationError:
invalid += 1
p95 = sorted(latencies)[int(0.95 * (len(latencies) - 1))]
return {"correct": correct, "n": len(tickets), "invalid": invalid, "p95_ms": p95,
"cost_per_1k": cost / len(tickets) * 1000}
def decide(baseline: dict, candidate: dict, config: FeatureConfig, noise: int = 2) -> str:
"""Promote only if not worse beyond noise, no new invalid outputs, and within cost and latency budgets."""
if candidate["invalid"] > baseline["invalid"]:
return f"HOLD: {candidate['invalid']} invalid outputs vs {baseline['invalid']}"
if candidate["correct"] < baseline["correct"] - noise:
return f"HOLD: accuracy dropped beyond the noise allowance of {noise} tickets"
if candidate["cost_per_1k"] > config.max_cost_per_1k_usd or candidate["p95_ms"] > config.max_p95_latency_ms:
return "HOLD: over cost or latency budget"
return "PROMOTE"
def fake_model(fence_every: int, wrong_every: int, latency_ms: float):
"""STAND-IN 'models' with known behaviour (not real models): some wrap JSON in fences, some mislabel."""
counter = {"i": 0}
def responder(messages, kwargs):
counter["i"] += 1
t = next(t for t in TICKETS if t.text == messages[-1]["content"])
category = "how_to" if counter["i"] % wrong_every == 0 else t.gold["category"]
body = json.dumps({"category": category, "priority": t.gold["priority"], "language": t.language,
"summary": f"About {t.subject}.", "needs_human": False})
text = f"```json\n{body}\n```" if counter["i"] % fence_every == 0 else body
return ChatResult(text=text, usage=Usage(input_tokens=120, output_tokens=45), latency_ms=latency_ms)
return ScriptedLLM(responder=responder)
TICKETS = load_tickets("test")
if __name__ == "__main__":
for key in PROMPTS:
print(f"prompt {key[0]}/{key[1]} sha256:{prompt_hash(*key)}")
baseline = run_gate(PRODUCTION, fake_model(fence_every=10**6, wrong_every=8, latency_ms=900), TICKETS)
print("baseline ", PRODUCTION.model, PRODUCTION.prompt_version, fmt_rate(baseline["correct"], baseline["n"]))
candidates = {
"new model, old prompt": (FeatureConfig("triage", "groq", "new-model-x", "v3", 0.50, 3000), 6, 12, 700),
"new model, prompt v4": (FeatureConfig("triage", "groq", "new-model-x", "v4", 0.50, 3000), 10**6, 12, 700),
}
for name, (config, fence_every, wrong_every, latency) in candidates.items():
result = run_gate(config, fake_model(fence_every, wrong_every, latency), TICKETS)
print(f"{name:22} {fmt_rate(result['correct'], result['n'])}, invalid {result['invalid']}, "
f"p95 {result['p95_ms']:.0f} ms -> {decide(baseline, result, config)}")
Code explained
- In simple words: every model or prompt upgrade has to pass the same driving test as the one it replaces, and "drives about as well but crashes sometimes" is a fail.
- What happens:
PROMPTSstores prompt text by feature and version;prompt_hashgives a short SHA-256 that goes into logs and feedback events (B6), so any output can be traced to the exact prompt text.FeatureConfigpins provider, model, prompt version, and the cost and latency budgets for one feature.run_gatesends the 24 test tickets through a configuration, counting correct categories, invalid outputs, p95 latency, and cost per 1,000 tickets viacost_usd. A model missing fromPRICESis priced at gpt-oss-120b rates as a placeholder, with a comment telling you to add it first.decidepromotes only if there are no new invalid outputs, accuracy has not dropped by more than a stated noise allowance (2 tickets of 24), and budgets hold.fake_modelbuilds stand-in "models" with known behaviour, not real models: the new one wraps every sixth reply in a code fence (a common real change between model versions) and mislabels every twelfth.
- Comes out:text
prompt triage/v3 sha256:031fc023 prompt triage/v4 sha256:e7835c15 baseline openai/gpt-oss-120b v3 21/24 = 87.5% (95% CI 69.0% to 95.7%) new model, old prompt 20/24 = 83.3% (95% CI 64.1% to 93.3%), invalid 4, p95 700 ms -> HOLD: 4 invalid outputs vs 0 new model, prompt v4 22/24 = 91.7% (95% CI 74.2% to 97.7%), invalid 0, p95 700 ms -> PROMOTEThe new model with the old prompt is held because 4 replies came back fenced and failed validation, a regression an accuracy-only comparison would have hidden. With prompt v4, which says "no code fences" explicitly, the same model passes. Read PROMOTE carefully: 22 of 24 versus 21 of 24 is well within noise, so the gate's claim is "not worse and within budget", not "better". That is the right bar for a routine migration.
Module Lab
The lab joins the module's pieces into one intake pipeline for held-out tickets. Its rule is the main lesson of Part C: run the cheapest reliable step first, and call the model only for what the cheap steps cannot do.
# examples/m14_lab.py
"""Module 14 Lab: the Brightlane intake pipeline, cheapest reliable step first.
For each held-out ticket: regex extracts invoice ids; keyword rules classify
when they match and the LLM classifies only when they do not; the grounded
drafter answers or abstains; the autonomy policy decides who acts; every step
writes a feedback event; the run ends with measured counts and projected cost.
"""
from __future__ import annotations
import json
import re
from collections import Counter
from pathlib import Path
from pydantic import ValidationError
from examples.m14_autonomy import ACTIONS, autonomy_for
from examples.m14_classical_baselines import ROBUST, RULES, extract
from examples.m14_common import fmt_rate, pick_llm
from examples.m14_feedback_events import FeedbackEvent, log
from examples.m14_grounded_answer import answer, choose_threshold
from examples.m14_triage_at_scale import build_messages as triage_messages
from supportdesk.data import load_tickets
from supportdesk.llm import Usage
from supportdesk.pricing import cost_usd
from supportdesk.schemas import Triage
from supportdesk.stand_in import ScriptedLLM
MODEL = "openai/gpt-oss-120b"
POLICY = {a.name: autonomy_for(a) for a in ACTIONS}
def rule_match(text: str) -> str | None:
low = text.lower()
return next((category for category, pattern in RULES if re.search(pattern, low)), None)
def responder(messages, kwargs):
"""STAND-IN (not a model): triage always says how_to; drafts cite the first source they were shown."""
if "You triage support tickets" in messages[0]["content"]:
return json.dumps({"category": "how_to", "priority": "normal", "language": "en",
"summary": "Customer needs help with a feature.", "needs_human": False})
source = re.search(r'<source id="([^"]+)">', messages[-1]["content"]).group(1)
return json.dumps({"reply": f"Here is what our help center says ({source}).", "cited_articles": [source],
"confidence": "medium"})
class Metered:
"""Wraps any chat function and adds up calls and token usage, whatever the backend."""
def __init__(self, llm) -> None:
self.llm, self.calls, self.usage = llm, 0, Usage()
def __call__(self, messages, **kwargs):
result = self.llm(messages, **kwargs)
self.calls += 1
self.usage.input_tokens += result.usage.input_tokens
self.usage.output_tokens += result.usage.output_tokens
return result
if __name__ == "__main__":
llm = Metered(pick_llm(ScriptedLLM(responder=responder)))
threshold = choose_threshold(load_tickets("dev"))
events = Path("m14_out/lab_events.jsonl")
events.parent.mkdir(exist_ok=True)
events.unlink(missing_ok=True)
routes, rule_hits, invoice_count = Counter(), [], 0
for t in load_tickets("test"):
invoices = extract(ROBUST, t.text)
category = rule_match(t.text)
if category:
rule_hits.append(category == t.gold["category"])
routes["category by rule"] += 1
else:
result = llm(triage_messages(t), temperature=0.0)
try:
category = Triage.model_validate_json(result.text).category
routes["category by LLM"] += 1
except ValidationError:
category, routes["category fallback"] = "how_to", routes["category fallback"] + 1
draft, why = answer(t, llm, threshold)
action = "send_reply_with_kb_link" if why == "answered" else "route_to_queue"
routes[f"{action} ({POLICY[action]})"] += 1
log(FeedbackEvent(event="draft_shown" if why == "answered" else "escalated", ticket_id=t.id,
draft_id=f"lab-{t.id}", feature="draft_reply", prompt_version="lab-v1", model=MODEL,
agent="agent_lab", cited_articles=draft.cited_articles), events)
invoice_count += len(invoices)
print(json.dumps(dict(routes), indent=1))
print("rule accuracy where rules fired:", fmt_rate(sum(rule_hits), len(rule_hits)))
print(f"invoice ids extracted: {invoice_count}; events logged: {sum(1 for _ in events.open())}")
per_ticket = cost_usd(llm.usage, MODEL) / 24
print(f"LLM calls: {llm.calls} for 24 tickets ({llm.usage.input_tokens} in, {llm.usage.output_tokens} out tokens)")
print(f"projected {MODEL} cost: {per_ticket * 100_000:.2f} USD per 100k tickets "
f"(stand-in token counts; reasoning tokens not included)")
Code explained
- In simple words: one conveyor belt through every station built in this module, with a meter on the expensive machine.
- What happens:
- It imports the pieces it needs from this module's own examples: the autonomy policy (B1), the robust regex and rules (C2), the feedback schema (B6), the grounded answerer and threshold (A3), and the triage prompt (A1).
rule_matchreturns a category only when a rule actually fires, instead of falling back tohow_to, so the pipeline knows when to escalate to the model.Meteredwraps whichever chat function is in use and counts calls and tokens, so the cost report works the same for the stand-in and for a live model.- For each test ticket: extract invoice ids, classify by rule or by LLM (with a fallback on invalid output), draft or abstain, look up the autonomy level for the resulting action, and log a feedback event.
- The stand-in always answers
how_tofor triage and cites the first source it was shown for drafts. It exists to drive the plumbing; its category answers are not scored.
- Comes out:text
[stand-in] ScriptedLLM: tests the plumbing only, not model quality { "category by rule": 11, "send_reply_with_kb_link (suggest_and_approve)": 15, "category by LLM": 13, "route_to_queue (auto)": 9 } rule accuracy where rules fired: 9/11 = 81.8% (95% CI 52.3% to 94.9%) invoice ids extracted: 0; events logged: 24 LLM calls: 28 for 24 tickets (7584 in, 968 out tokens) projected openai/gpt-oss-120b cost: 7.16 USD per 100k tickets (stand-in token counts; reasoning tokens not included)Rules fire on 11 of 24 tickets and are right on 9 of them (82%, interval 52% to 95%). Compare that with 58% for the rules on all tickets in C2: letting rules abstain when they are unsure makes them much more accurate on the tickets they keep, which is the whole idea behind a cascade. The other 13 go to the model. Fifteen tickets get a draft for agent approval and nine are routed straight to the queue, matching A3. The test split contains no invoice ids, so the regex finds none, which is the right answer. The model was called 28 times for 24 tickets (13 triage calls plus 15 drafts), and the projected cost is about 7 USD per 100,000 tickets on gpt-oss-120b. That projection uses the stand-in's short replies and no reasoning tokens, so rerun with
M14_LIVE=1to replace it with a measured figure before quoting it.
Extend the lab in this order, measuring after each step: replace the stand-in with a real model and record its category accuracy on the 13 escalated tickets; train the C2 classifier on a larger labelled set and put it between the rules and the model; and apply the D6 gate before switching the model behind any stage.
Project Milestone
The Brightlane repository (work/m14) now contains:
| File | What it adds |
|---|---|
examples/m14_common.py | Live or stand-in switch, Wilson intervals |
examples/m14_triage_at_scale.py | Concurrent triage with dead-lettering and a 100k-per-month cost table |
examples/m14_weekly_digest.py | Map-reduce digest with a number checker |
examples/m14_grounded_answer.py | Retrieval gate chosen on dev, citation gate, abstention |
examples/m14_sql_codegen.py | Text-to-SQL with a read-only database and tests as the verifier |
examples/m14_review_queue.py | Review queue state machine for article drafts |
examples/m14_enrich_pipeline.py | Idempotent enrichment with a content-addressed cache |
examples/m14_ticket_search.py | Permission-filtered search over past tickets |
examples/m14_autonomy.py | Autonomy policy, undo log, fallbacks |
examples/m14_reviewer_fatigue.py | Review placement model, exact and simulated |
examples/m14_draft_card.py | Agent card with measured confidence bands, customer disclosure lines |
examples/m14_feedback_events.py | Feedback event schema and weekly summary |
examples/m14_classical_baselines.py | Regex, rules, and TF-IDF classifier measured against each other |
examples/m14_reliability_bar.py | Expected-cost comparison and break-even accuracy |
examples/m14_limitations_onepager.py | Stakeholder brief generated from measurements |
examples/m14_frontier_math.py | Trend and release arithmetic for Brightlane |
examples/m14_tinylm_latency.py | TinyLM CPU latency |
examples/m14_adaptable.py | Prompt registry and model-upgrade gate |
examples/m14_lab.py | The cascade pipeline |
tests/test_m14_patterns.py | 14 tests pinning the behaviour above |
Run the tests with PYTHONPATH=. python -m pytest -q tests/test_m14_patterns.py; in this build all 14 pass in about 9 seconds. The canonical supportdesk/ package is unchanged: everything here imports it.
The assistant can now do more than answer tickets. It can say which of its jobs it should not be doing, what each of its actions is allowed to do alone, what its limits are in numbers a support lead can read, and how it will be checked the next time its model changes.
Interview Questions
1. A product manager wants an LLM to classify 100,000 support tickets a month. What do you do first? Measure the cheap baselines on a labelled sample: majority class, keyword rules, and a classical classifier trained on historical labels, each with a confidence interval on a held-out split. Then price the LLM per ticket from real token counts, including reasoning tokens, and compare batch pricing and caching. In the Brightlane data, rules and TF-IDF are within noise of each other at 48 training tickets, and the right design is usually a cascade: rules or classifier when confident, LLM for the rest, all validated against a schema with a dead-letter path.
2. How do you decide what an AI system may do without human approval? Derive autonomy from the action's consequences, not from how good the model seems: reversibility, customer visibility, whether it moves money, data, or access, and blast radius. Reversible internal actions can be automatic with undo; customer-visible or access-changing actions need approval; irreversible money or data actions stay human-only. Then relax a level only with evidence from approval and edit rates.
3. Your reviewers approve 99% of AI drafts. Is that good news? Not necessarily. A very high approval rate is consistent with good drafts and with fatigued reviewers. Plant known-wrong drafts, measure the catch rate with an interval (you need on the order of 100 planted errors for a usable interval), shorten sessions if catch rates decay, and target review at items a measured signal flags, while auditing a random sample of the rest.
4. When is a regex better than an LLM for extraction? When the field has a fixed format: invoice ids, VAT numbers, order numbers. A tested regex is deterministic, explainable, and runs in hundredths of a millisecond. The work is in the hard cases: case, spacing, neighbours like SINV-, and Unicode word boundaries (a \b after Japanese text fails in Python). Use an LLM only when the format genuinely varies, and even then validate its output with a regex.
5. How do you know whether 90% accuracy is good enough to automate? Price the outcomes. Compute the expected cost per item of each option (manual, draft and review, automatic) from agent time, review time, fix time, catch rate, and the cost of an error that reaches the customer, then find the break-even accuracy. With Brightlane's assumptions, auto-send wins for how-to answers above 84%, but for errors costing 300 USD the bar is 99.6%. Then plan with the lower end of your measured accuracy interval, not the point estimate.
6. How should an assistant decide to say "I don't know"? Use evidence available before generation and after it. Before: a retrieval score threshold chosen on dev data for a target wrong-answer rate, so weak matches never reach the model. After: validation of the output and a citation check that every cited source was actually retrieved. Measure both the unsafe-answer rate and the rate of safe answers you gave up, on held-out data, and report both.
7. What is wrong with filtering search results by permission after ranking? Two things. Any bug in the post-filter leaks restricted content directly. And even a correct post-filter lets restricted documents influence collection-wide statistics such as IDF, so everyone's rankings depend on documents they cannot see; in the Brightlane example a shared ticket's score changed from 3.649 to 3.495 depending on the index. Filter before ranking: per-permission indexes or filtered queries.
8. How do you make an LLM enrichment pipeline safe to rerun? Extract deterministic fields with code, cache model replies under a hash of the exact request including the prompt version, validate every reply, write failures to a dead-letter file, and rewrite outputs rather than appending. Cache only validated replies, or failures replay forever. Changing the prompt version should invalidate the cache on purpose.
9. A vendor announces a model that "outperforms most frontier models" and is half the price. How do you evaluate it? Check whether the benchmarks resemble your task, what the comparisons were against and under which settings, whether the price is introductory, and whether the model uses more tokens per task. Then run your own eval suite through the same gate as any change: accuracy within noise, no new invalid outputs, cost per task and p95 latency within budget. Per-token price is not per-task price.
10. What does METR's time horizon tell you, and what does it not? It estimates the length of task, in skilled-human time, that a model completes with 50% success on a mostly software suite, and it has grown with a doubling time of roughly 3 to 7 months depending on the period. It does not say how long a model can work unattended, its error bars are about a factor of 2 either way, and it varies by orders of magnitude across domains. For product design it means agents will attempt longer tasks, while customer-facing actions still need reliability far above 50%.
11. What feedback should an LLM feature log from day one? Structured events with a fixed schema: draft shown, sent with a measured edit ratio, discarded, escalated, wrong source flagged, undo, each carrying the ticket, prompt version, model, cost, latency, and a pseudonymous user id. That makes "did the new prompt help?" a query with confidence intervals, and it produces labels for future evaluation and fine-tuning.
12. How do you keep an LLM system adaptable as models change? Keep prompts as versioned, hashed data; route every call through a provider-neutral interface with the model named in one configuration per feature; and gate every model or prompt change on the same golden set, invalid-output count, cost, and latency budgets. Expect format regressions (such as new code fences) as often as accuracy changes, and write the promotion rule as "not worse beyond noise", since small test sets cannot show improvement.
Other Tools and Providers
| Tool or provider | What it does | When to consider it instead |
|---|---|---|
| fastText, SetFit | Fast text classification; SetFit fine-tunes a small sentence-transformer from few labels | When the C2 classifier needs more accuracy but an LLM is too slow or costly |
| spaCy rule matchers, Duckling | Rule-based extraction with linguistic features; Duckling parses dates, amounts, and durations | When regexes grow unmaintainable but the fields are still well defined |
| Label Studio, Argilla | Human review and labelling queues | Instead of the hand-built A5 queue when many reviewers, guidelines, and agreement metrics are involved |
| Langfuse, Arize Phoenix, LangSmith, Braintrust | Tracing, feedback capture, and eval dashboards | Instead of the B6 JSONL file once several features and teams log events |
| Elasticsearch or OpenSearch, Vespa, pgvector, Qdrant | Search engines and vector stores with filtered queries | For A7 when per-group indexes stop scaling |
| Provider batch APIs (OpenAI, Anthropic, Google, Groq) | Asynchronous processing at a discount | For backlog classification and enrichment jobs that can wait |
| Apple Foundation Models framework, Gemini Nano, llama.cpp, Ollama, MLC LLM | On-device and local inference | For D4-style private, offline, narrow tasks |
| Epoch AI, METR, Artificial Analysis, LMArena | Independent trend data, agent evaluations, speed and price benchmarks, human preference rankings | For D5: context before you trust a launch post (still no substitute for your own eval) |
Coming Up: Capstone Project
The course ends with a capstone in module15-capstone.md. You pick one of four tracks and build it on the Brightlane assistant you now have:
- Product track: ship one LLM feature end to end, with prompt versioning, structured outputs, an evaluation suite with CI gates, cost and latency budgets, a safety review, and a documented failure-mode analysis.
- Agent track: build an agent with real tools, sandboxing, termination guards, trajectory evaluation, human approval gates, and full tracing, and show it behaving correctly under injected adversarial content.
- Adaptation track: take a task where prompting plateaus, build a fine-tuning dataset, adapt a model, and prove lift over the prompted baseline while measuring general-capability regression.
- Evaluation track: build a rigorous eval harness for an existing system (assertion tier, calibrated judge tier, noise-floor measurement, CI gates) and use it to find and fix three real quality regressions.
Every track shares four requirements that this module has been practicing: a baseline the final system must beat (Part C), error analysis on at least 50 real failures clustered by cause, cost and latency measured rather than estimated, and a written account of what did not work. The capstone file has the full brief, the milestones, and how to present your results.