Part G: The Model Landscape
Frontier, mid-tier, and small
Models cluster into rough tiers. The names change every few months, so learn the tiers and look up the current members when you choose.
| Tier | Typical size and access | Strengths | Costs | Examples (September 2026) |
|---|---|---|---|---|
| Frontier | Largest proprietary models via API; some very large open-weight MoEs | Hardest reasoning, coding, long agentic tasks | Highest price per token, slower, often reasoning by default | Flagship models from OpenAI, Anthropic, and Google; DeepSeek's V4 preview line was reported at 1.6T total parameters (49B active) with a 1M-token context (Digital Applied retrospective) |
| Mid-tier | Fast API models; open-weight models of roughly 20B to 120B total | Most application work: triage, extraction, grounded answers, drafting | Much cheaper and faster | gemini-3.5-flash (1M-token context, GA); openai/gpt-oss-120b on Groq |
| Small | Roughly 1B to 10B dense, or small MoEs | Classification, extraction, routing, on-device use | Cheapest; weaker at open-ended reasoning and rare facts | qwen3:8b on Ollama; the Qwen 3.5 small models (0.8B to 9B, Apache 2.0, reported March 2026) |
"Tier" is about capability per task, not a fixed size: a well-chosen small model can match a frontier model on a narrow task like Brightlane's six-way ticket classification, and that is exactly what you will measure in later modules.
Proprietary API models vs open-weight families
| Situation | Use this | Why |
|---|---|---|
| You need the strongest capability now, with low setup effort | Proprietary API model | Best models are often API-only; no hardware to run |
| Data must not leave your network, or regulation requires residency you control | Open-weight model, self-hosted | You control where prompts and outputs live (Module 13) |
| You need to fine-tune deeply, pin a version forever, or inspect probabilities | Open-weight model | Weights never change under you; full access to internals |
| You want open-weight economics without running GPUs | An open-weight model on a hosted provider (as the course does with gpt-oss-120b on Groq) | Pay per token, keep the option to self-host later |
| Low volume, spiky traffic | API (either kind) | Self-hosted GPUs cost money while idle |
A practical point about version pinning: API model names can be retired or silently improved; open weights you download are frozen. Module 13 covers deprecations and forced migrations.
Reasoning models vs standard models: cost and latency
A reasoning model spends extra output tokens thinking before it answers. Output tokens are what drive latency (they are generated one at a time) and usually cost (output is priced higher than input). Two measured and published facts make this concrete for the course's default:
- Groq serves gpt-oss-120b at about 500 output tokens per second, with a 131,072-token context and prices of 0.15 USD per million input tokens and 0.60 USD per million output tokens (Groq docs, checked 21 September 2026).
- At 500 tokens per second, every 200 reasoning tokens adds about 0.4 seconds before the first visible word of the answer, and bills as 200 output tokens.
For a one-word triage label, reasoning can multiply the output tokens many times over. For a tricky refund-policy question that combines three rules, it may be what makes the answer correct. The lab below prices both profiles; Module 3 measures when extra reasoning helps and when it only wastes money.
Small language models and on-device use
Small language models (SLMs) run on a laptop, phone, or a single modest GPU. Apple's on-device foundation model is about 3 billion parameters, compressed to 2 bits per weight with quantization-aware training, and exposed to app developers through the Foundation Models framework with guided (constrained) generation and tool calling (Apple Machine Learning Research). OpenAI describes gpt-oss-20b as running within 16 GB of memory. Using the arithmetic from Part D, a 3B model at 2 bits is about 0.75 GB of weights: that is why it fits on a phone.
On-device wins on privacy (nothing leaves the device), offline use, and zero per-token cost. It loses on capability, especially rare facts and long reasoning. For Brightlane, a small local model is a reasonable candidate for the first step (triage) and a poor one for drafting policy answers without retrieval.
Model cards, benchmarks, and why leaderboard rank rarely transfers
A model card is the document published with a model: intended use, training data summary, evaluation results, limitations, licence. Read it for facts (context, cutoff, licence, languages), and read its benchmark table with suspicion. A benchmark is a fixed public test set; a leaderboard ranks models on one or more of them. Rank transfers poorly to your task for three reasons:
- Different task: public benchmarks test general skills; your task has its own inputs, labels, and failure costs.
- Contamination: benchmark questions leak into training data, so scores can measure memorization.
- Leaderboard dynamics: a 2025 study of a popular crowd-voted leaderboard found that some providers privately tested many variants and published only the best, and that access to leaderboard data favored those who had more of it (Singh et al., "The Leaderboard Illusion").
You can reproduce the first two effects in miniature. Task: pick the help-center article that answers a question. System A is TinyLM (generate the agent's reply, then match it to an article). System B is not a model at all: match the customer's own words against each article and pick the one with the most words in common. The "benchmark" is the 12 question templates TinyLM's corpus was built from, which is to say a contaminated benchmark. "Our tickets" are the 62 answerable Brightlane tickets.
"""Module 1: why a benchmark rank may not transfer to your task.
Task: pick the help-center article that answers a customer question.
Two systems:
A. TinyLM: generate the agent's reply, then match the reply to an article.
B. Keyword overlap: match the customer's own words to an article (no model at all).
Two test sets:
"benchmark": the 12 question templates TinyLM's corpus was generated from (contaminated).
"our tickets": the 62 answerable Brightlane tickets, written independently.
"""
import math
import re
import torch
from supportdesk.data import load_articles, load_tickets
from supportdesk.tinylm import SamplingParams, generate, load
torch.set_num_threads(1) # one thread is plenty for a 1M-parameter model
model, tokenizer = load()
articles = load_articles()
STOP = set("a an and are as at be by can do does for from get how i in is it my of on or our the to "
"we what when why will with you your this that there".split())
def words(text: str) -> set[str]:
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOP}
ARTICLE_WORDS = {a.id: words(a.title + " " + a.body) for a in articles}
def best_article(text: str) -> str:
"""The article sharing the most distinct words with `text` (ties: first by id)."""
return max(ARTICLE_WORDS, key=lambda aid: len(words(text) & ARTICLE_WORDS[aid]))
def tinylm_route(question: str) -> str:
prompt = f"Customer (Maya): Hi, {question}\nAgent (Dara):"
reply = generate(model, tokenizer, prompt, SamplingParams(max_new_tokens=40, temperature=0,
stop=["\nCustomer"])).text
return best_article(reply)
def keyword_route(question: str) -> str:
return best_article(question)
BENCHMARK = [ # the corpus templates, with {plan}/{area} filled in
("I was charged twice for the Team plan this month. Can you refund the duplicate?", "billing-refunds"),
("how do I cancel my Team subscription?", "billing-refunds"),
("I forgot my password and the reset link expired.", "account-login"),
("my account is locked after too many attempts.", "account-login"),
("does the Team plan include SSO?", "account-sso"),
("our automations stopped and there is a banner about a limit.", "boards-automations"),
("how do I export my board data?", "exports-data"),
("slack notifications stopped posting.", "integrations-slack"),
("can I get a refund on my annual Team plan?", "billing-refunds"),
("where can I find my invoices?", "billing-invoices"),
("how much does the Business plan cost?", "billing-plans"),
("the board is not loading. Is there an outage?", "status-incidents"),
]
TICKETS = [(t.body, t.gold["kb_article"]) for t in load_tickets() if t.gold["answerable"]]
def wilson(k: int, n: int) -> str:
z = 1.96
c = (k / n + z * z / (2 * n)) / (1 + z * z / n)
h = z * math.sqrt(k / n * (1 - k / n) / n + z * z / (4 * n * n)) / (1 + z * z / n)
return f"{max(c - h, 0):.2f} to {min(c + h, 1):.2f}"
print(f"{'system':18} {'benchmark (n=12)':>24} {'our tickets (n=62)':>26}")
for name, route in [("A. TinyLM", tinylm_route), ("B. keyword overlap", keyword_route)]:
cells = []
for dataset in (BENCHMARK, TICKETS):
k = sum(route(q) == gold for q, gold in dataset)
cells.append(f"{k}/{len(dataset)} ({wilson(k, len(dataset))})")
print(f"{name:18} {cells[0]:>24} {cells[1]:>26}")Code explained
- In simple words: run two systems on a public-looking benchmark and on our real tickets, and see whether the ranking holds.
- What happens:
wordslowercases text, keeps letters and digits, and drops common stop words.best_articlepicks the article sharing the most distinct words with a piece of text.tinylm_routegenerates a greedy reply (stopping at the next customer line) and routes the reply;keyword_routeroutes the question directly.wilsonprints a 95 percent interval for each score. - Comes out: real output. It takes about 5 seconds of compute; on a busy machine allow up to a minute.
textCopy
system benchmark (n=12) our tickets (n=62)
A. TinyLM 12/12 (0.76 to 1.00) 10/62 (0.09 to 0.27)
B. keyword overlap 8/12 (0.39 to 0.86) 43/62 (0.57 to 0.79)On the benchmark, TinyLM wins 12 to 8 and would top the leaderboard. On our tickets it loses 10 to 43, and the intervals do not overlap, so this is not noise. The benchmark rewarded memorization of the exact templates (contamination) and said nothing about real customers' wording. The keyword baseline's 43 of 62 (0.69) is also a useful number in its own right: it is the floor any model-based router for Brightlane must beat, and you will meet it again.
Large models are nothing like TinyLM, but the lesson is the same one practitioners keep relearning: choose models with an evaluation built from your own data (Module 10), and use public benchmarks only to build a shortlist.
A selection framework for Brightlane
Five questions decide most model choices. Answer them in order, because the first one is a filter, not a trade-off:
- Capability floor: does the model pass your own eval at the required quality? If not, nothing else matters.
- Latency: does it meet your time budget (for example, time to first token for a chat widget, total time for background triage)?
- Cost: what is the cost per task times your volume?
- Privacy: where do prompts and outputs go, who can see them, and what do your customers and regulators require?
- Control: can you pin the version, fine-tune it, inspect probabilities, and keep running if the provider changes terms?
Applied to the Brightlane assistant's first two jobs (numbers marked as assumptions must be replaced with your measurements):
| Criterion | Ticket triage (label + priority) | Drafting answers from the help center |
|---|---|---|
| Capability floor | Must beat the 0.69 keyword baseline on article routing and reach Maya's target on category accuracy (set in Module 10), on the dev split | Must be grounded in retrieved articles with no invented policy; measured with a judge and human review (Module 10) |
| Latency | Background job; seconds are fine | Agent is waiting; aim for first words within about a second (Module 3) |
| Cost | High volume, tiny outputs: small or mid-tier standard model, reasoning off or low | Lower volume, longer outputs: a mid-tier model; reasoning only if evals show it helps |
| Privacy | Tickets contain names, emails, invoice ids: redact before sending or self-host (Module 11) | Same, plus drafts must not leak one customer's data to another |
| Control | Pin the model version; keep a fallback provider (Module 13) | Pin the version and re-run the eval on every model change |
| Starting candidates | qwen3:8b locally, openai/gpt-oss-120b at low effort on Groq | openai/gpt-oss-120b or gemini-3.5-flash, with retrieval |
| Situation | Use this | Why |
|---|---|---|
| You do not yet have an eval for the task | Build the eval first, with the cheapest plausible model | Without it you cannot tell whether a pricier model is better |
| Two models pass the floor | Pick the cheaper and faster one | Capability above the floor is not worth paying for twice |
| Only a frontier model passes | Use it, and log inputs and outputs to build a fine-tuning or routing dataset later | Modules 9 and 13 show how to move traffic to cheaper models safely |
| Privacy rules forbid sending ticket text out | Self-hosted open weights, or redaction before any API call | Privacy is a hard constraint, not a weight in a score |
Module Lab
The lab combines the module into one script that writes a measured fact sheet for a model: spec and memory (Part D), what training did (Part C), behavior probes for phrasing, hallucination, and calibration (Part F), a check on the real Brightlane task against a no-model baseline (Part G), and the monthly cost of hosted candidates for the same job. Its rule is the module's rule: every line is either measured, taken from a cited source, or labelled as an assumption.
"""Module 1 Lab: write a one-page fact sheet for a model, measured rather than assumed.
It inspects TinyLM (spec, memory, training), probes its behaviour (memorized vs reworded
questions, hallucination, calibration), checks it on the real Brightlane task, and prices
the hosted candidates for the same job. Output: reports/m01_model_card.json plus a printout.
"""
import json
import math
import re
from pathlib import Path
import torch
from supportdesk.data import load_articles, load_tickets
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.tinylm import MODEL_DIR, SamplingParams, generate, load
torch.set_num_threads(1) # one thread is plenty for a 1M-parameter model
model, tokenizer = load()
card: dict = {"model": "tinylm-base"}
# 1. Spec and memory -------------------------------------------------------------
params = model.num_parameters()
card["spec"] = {"parameters": params, "layers": model.cfg.n_layers, "heads": model.cfg.n_heads,
"d_model": model.cfg.d_model, "context_tokens": model.cfg.context,
"embedding_rows": model.cfg.vocab_size, "tokenizer_vocab": tokenizer.get_vocab_size()}
card["weights_mb"] = {dtype: round(params * b / 1e6, 2) for dtype, b in
{"fp32": 4, "bf16": 2, "int8": 1, "4-bit": 0.5}.items()}
# 2. What pretraining did ----------------------------------------------------------
log = json.loads((MODEL_DIR / "training_log.json").read_text())
best = min(log, key=lambda r: r["val_loss"])
card["training"] = {"best_step": best["step"], "best_val_loss": best["val_loss"],
"final_train_loss": log[-1]["train_loss"], "final_val_loss": log[-1]["val_loss"],
"epochs_over_train_tokens": round(log[-1]["tokens_seen"] / 156_620, 1),
"overfitting": log[-1]["val_loss"] > best["val_loss"]}
@torch.no_grad()
def next_probs(prompt: str) -> torch.Tensor:
return torch.softmax(model(torch.tensor([tokenizer.encode(prompt).ids]))[0, -1], dim=-1)
# 3. Behaviour probes ----------------------------------------------------------------
# (question as trained, reworded question, correct first token of the reply)
PAIRS = [
("my account is locked after too many attempts.", "it says my account is locked. Why?", " After"),
("I was charged twice for the Team plan this month. Can you refund the duplicate?",
"you billed my card two times this month for Team, please refund one.", " Sorry"),
("how do I export my board data?", "how can I download my board as a spreadsheet?", " Export"),
("I forgot my password and the reset link expired.", "my password link does not work anymore.", " Reset"),
("does the Team plan include SSO?", "can we log in with Okta on Team?", " SSO"),
]
probe = {"as_trained": [], "reworded": []}
for trained, reworded, correct in PAIRS:
target = tokenizer.encode(correct).ids[0]
for key, q in (("as_trained", trained), ("reworded", reworded)):
probs = next_probs(f"Customer (Maya): Hi, {q}\nAgent (Dara):")
probe[key].append((int(probs.argmax()) == target, probs.max().item()))
card["phrasing"] = {k: f"{sum(ok for ok, _ in v)}/{len(v)} top-1 correct" for k, v in probe.items()}
france = next_probs("The capital of France is")
card["hallucination"] = {"prompt": "The capital of France is",
"top_token": tokenizer.decode([int(france.argmax())]),
"top_prob": round(france.max().item(), 3)}
pairs = [c for v in probe.values() for c in v]
card["calibration_on_probes"] = {
"n": len(pairs), "accuracy": round(sum(ok for ok, _ in pairs) / len(pairs), 2),
"mean_confidence": round(sum(p for _, p in pairs) / len(pairs), 2),
"confident_wrong(p>=0.8)": sum(1 for ok, p in pairs if p >= 0.8 and not ok)}
# 4. The real task: pick the right help-center article for each answerable ticket ----
STOP = set("a an and are as at be by can do does for from get how i in is it my of on or our the to "
"we what when why will with you your this that there".split())
words = lambda t: {w for w in re.findall(r"[a-z0-9]+", t.lower()) if w not in STOP} # noqa: E731
ARTICLE_WORDS = {a.id: words(a.title + " " + a.body) for a in load_articles()}
best_article = lambda t: max(ARTICLE_WORDS, key=lambda a: len(words(t) & ARTICLE_WORDS[a])) # noqa: E731
tickets = [t for t in load_tickets() if t.gold["answerable"]]
greedy = SamplingParams(max_new_tokens=40, temperature=0, stop=["\nCustomer"])
tiny_hits = sum(best_article(generate(model, tokenizer, f"Customer (Maya): Hi, {t.body}\nAgent (Dara):",
greedy).text) == t.gold["kb_article"] for t in tickets)
rule_hits = sum(best_article(t.body) == t.gold["kb_article"] for t in tickets)
card["task_article_routing"] = {"n": len(tickets), "tinylm": tiny_hits, "keyword_rule": rule_hits}
# 5. Price the hosted candidates for the same job ------------------------------------
TICKETS_PER_MONTH = 3000 # assumption for Brightlane; use your own volume
STANDARD = Usage(input_tokens=90, output_tokens=13) # measured with the stand-in prompt
REASONING = Usage(input_tokens=90, output_tokens=13 + 200, reasoning_tokens=200) # assumption: 200 thinking tokens
card["monthly_cost_usd"] = {
f"{m} ({'reasoning' if u is REASONING else 'standard'})": round(cost_usd(u, m) * TICKETS_PER_MONTH, 2)
for m, u in [("openai/gpt-oss-120b", REASONING), ("gemini-3.5-flash", REASONING),
("llama-3.1-8b-instant", STANDARD), ("qwen3:8b", STANDARD)]
if m in PRICES}
out = Path("reports/m01_model_card.json")
out.parent.mkdir(exist_ok=True)
out.write_text(json.dumps(card, indent=2) + "\n")
for section, value in card.items():
print(f"{section}: {value}")
k, n = rule_hits, len(tickets)
print(f"\nkeyword rule accuracy {k / n:.2f} +/- {1.96 * math.sqrt(k / n * (1 - k / n) / n):.2f} (95%, normal approx.)")
print("saved", out)Code explained
- In simple words: a model report card generator. Point it at a model, and it fills in each section with a number you can defend.
- What happens:
- Section 1 reads the spec from the loaded model and tokenizer (not from memory or marketing) and computes weight memory at four precisions.
- Section 2 summarizes
training_log.json: best validation step, final losses, how many times the training set was repeated (tokens_seen / 156,620), and whether validation loss rose after its best point (overfitting). - Section 3 runs five intents twice, once in training wording and once reworded, and records top-1 correctness and confidence; it checks the France prompt for a hallucination; and it summarizes calibration on those ten probes, counting confident-wrong answers (probability at least 0.8 and wrong).
- Section 4 runs the article-routing task on all 62 answerable tickets for TinyLM and for the keyword rule.
- Section 5 prices hosted candidates with
pricing.cost_usd(the price table is explained in Module 2). The standard profile uses the 90 input and 13 output tokens counted for the Part A triage prompt. Two numbers are assumptions, marked in the code: 3,000 tickets per month, and 200 reasoning tokens per call for reasoning models. Replace both with your own measurements fromresult.usage. The localqwen3:8bshows 0.0 because there is no per-token price; your hardware and electricity are not free. - Everything is saved to
reports/m01_model_card.json. - Comes out: real output from this build (runtime is under a minute; the routing step dominates).
model: tinylm-base
spec: {'parameters': 1071872, 'layers': 4, 'heads': 4, 'd_model': 128, 'context_tokens': 128, 'embedding_rows': 2048, 'tokenizer_vocab': 1503}
weights_mb: {'fp32': 4.29, 'bf16': 2.14, 'int8': 1.07, '4-bit': 0.54}
training: {'best_step': 400, 'best_val_loss': 0.6553, 'final_train_loss': 0.2636, 'final_val_loss': 0.6799, 'epochs_over_train_tokens': 26.2, 'overfitting': True}
phrasing: {'as_trained': '5/5 top-1 correct', 'reworded': '1/5 top-1 correct'}
hallucination: {'prompt': 'The capital of France is', 'top_token': ' not', 'top_prob': 0.196}
calibration_on_probes: {'n': 10, 'accuracy': 0.6, 'mean_confidence': 0.65, 'confident_wrong(p>=0.8)': 0}
task_article_routing: {'n': 62, 'tinylm': 10, 'keyword_rule': 43}
monthly_cost_usd: {'openai/gpt-oss-120b (reasoning)': 0.42, 'gemini-3.5-flash (reasoning)': 6.16, 'llama-3.1-8b-instant (standard)': 0.02, 'qwen3:8b (standard)': 0.0}
keyword rule accuracy 0.69 +/- 0.11 (95%, normal approx.)
saved reports/m01_model_card.jsonReading the fact sheet as Maya would:
- TinyLM overfit (best step 400) after reading its training set 26 times.
- It answers 5 of 5 trained phrasings and 1 of 5 rewordings correctly: it matches strings, it does not understand requests.
- On the real routing task it gets 10 of 62, against 43 of 62 for a keyword rule that needs no model at all. With n = 62, the keyword rule's accuracy is 0.69 plus or minus 0.11.
- For the hosted candidates, cost is not the bottleneck at Brightlane's assumed volume: even the most expensive option in the table is a few dollars a month under these assumptions. The capability floor, latency, and privacy will decide, which is why the next modules build measurement before anything else.
To run the lab against a different model directory (for example, one you fine-tune in Module 9), change the load() call to load("models/<your-model>").
Project Milestone
After this module, your supportdesk repository contains:
- The dataset:
data/tickets.jsonl(72 labelled tickets, dev 48 and test 24) anddata/kb/(12 articles), loaded throughsupportdesk/data.py. - The provider helper
supportdesk/llm.py, configured byLLM_PROVIDER,LLM_MODEL, and a key in your environment, and the stand-insupportdesk/stand_in.pyfor key-free tests. - The pretrained TinyLM in
models/tinylm-base/, which you have probed but not changed. - Module 1 scripts in
examples/:m01_explore_data.py,m01_first_call.py,m01_next_token.py,m01_training_log.py,m01_memorize_hallucinate.py,m01_base_model.py,m01_model_spec.py,m01_token_view.py,m01_phrasing.py,m01_calibration.py,m01_leaderboard.py, andm01_lab.py. tests/test_m01_foundations.py(8 tests) andreports/m01_model_card.json.- Two numbers you will reuse: the keyword baseline for article routing (43 of 62 answerable tickets) and the triage prompt's token counts (90 input, 13 output with the stand-in's estimate).
If you have a key, run examples/m01_first_call.py against your provider now and note the real usage and latency. Module 2 turns them into cost per ticket.
Interview Questions
1. What does a large language model actually compute?
A probability distribution over its vocabulary for the next token, given the tokens so far. Generation calls that repeatedly, appending each chosen token. Everything else (chat roles, tools, reasoning) is built on top by formatting the input and by training the model to produce useful continuations. In practice this means every output is a plausible continuation, not a retrieved fact, so anything that must be true needs grounding or verification.
2. Why do models hallucinate, and why can't a vendor just fix it?
The training objective rewards assigning high probability to the text that actually came next in training data. It does not reward truth, and it has no built-in "I don't know": a distribution must put its probability mass somewhere, and generation always emits a token. For rare or unseen facts, the most plausible continuation is often wrong. Post-training reduces the rate and teaches some abstention, but the cause is the objective itself. The practical defenses are retrieval of trusted sources, constrained outputs, verification by code, and evaluation.
3. What is the difference between a base model and an instruction-tuned model? When would you use a base model?
A base model continues documents: it ignores instructions, does not stop at the end of a turn, and may write the other speaker's lines. Instruction tuning (SFT, then usually preference optimization) trains it on conversations ending with ideal answers, so it follows instructions, formats, and stops. Use a base model to start your own fine-tune, to score text likelihood, or for pure continuation in a fixed style; use an instruction-tuned model for almost all application work.
4. Walk through the stages of an LLM's lifecycle and name one behavior each stage causes.
Pretraining on trillions of tokens: knowledge, skills, and hallucination on rare facts. Mid-training or continued pretraining on curated or domain data: better math and code, longer context. Supervised fine-tuning: instruction following and output formats. Preference optimization (RLHF, DPO): the assistant's style, refusals, sycophancy, and typically worse calibration. Reasoning training with verifiable rewards: long thinking before answers, better math and code, more tokens and latency.
5. How much GPU memory does an 8B model need?
Weights alone: parameters times bytes per parameter. At bf16 that is 8e9 x 2 = 16 GB; at 8-bit about 8 GB; at 4-bit about 4 GB. Then add the KV cache, which grows with context length and batch size (for Llama 3.1 8B at bf16, about 131 KB per token, so about 4.3 GB at 32k tokens), plus runtime working memory. So an 8B model at 4-bit with a moderate context fits a 16 GB GPU; at bf16 with long contexts it needs more than 24 GB.
6. What do "total" and "active" parameters mean for a mixture-of-experts model, and why does it matter?
An MoE layer has many expert MLPs but routes each token to only a few. Total parameters (116.83B for gpt-oss-120b) must all be in memory. Active parameters (5.13B) are used per token and set the compute per token, which largely drives speed. So an MoE can be as fast as a much smaller dense model while needing the memory of a large one.
7. A model is described as "open". What do you check before using it commercially?
The licence: Apache 2.0 or MIT is permissive; custom licences (for example Llama 3.1's) can add user caps, naming and attribution requirements, and acceptable-use rules. Whether it is open weights only or open source in the OSI sense (training data information and code published too). And for distillation or fine-tuning on another model's outputs, the terms of that provider.
8. Why do LLMs struggle to count letters or do long arithmetic?
They see tokens, not characters or digits. "strawberry" arrives as three tokens like st, raw, berry, so the letters are not directly visible, and a long number is split into chunks that do not align with place value. The fix is to route exact work to code (tools) and let the model handle language; reasoning models help by writing steps out, but results still need verification.
9. What is calibration, and how would you check it for a triage classifier?
A model is calibrated when its stated confidence matches its accuracy. Collect confidence (token probability, or agreement across several samples) and correctness on a labelled set, bin by confidence, and compare mean confidence with accuracy per bin (a reliability table), with confidence intervals because bins are small. Use it to choose an escalation threshold, and recalibrate whenever the model or prompt changes. Never let confidence alone gate an irreversible action.
10. A new model tops a public leaderboard. Should you switch your production system to it?
Not on that evidence. Leaderboards measure different tasks, can be inflated by contamination and selective reporting, and rank transfer to a specific task is weak. Run your own eval on your own data, check latency, cost, privacy, and licence, and switch only if it clears the capability floor at acceptable cost. In this module a model that won a contaminated benchmark 12 to 8 lost on real tickets 10 to 43.
11. What is a knowledge cutoff, and how do you design around it?
The date after which the model saw no training data; it has no knowledge of later events and cannot tell that it is out of date. Keep changing facts (prices, known issues, roadmaps) in documents retrieved at request time, date them, and let the model phrase answers from them. Fine-tune for behavior, not for facts that will change.
12. When would you choose a reasoning model over a standard one?
When the task needs several dependent steps (math, code, policy questions that combine rules) and your eval shows reasoning improves accuracy enough to justify extra output tokens and latency. For short, high-volume tasks like classification, use a standard model or the lowest reasoning effort, because thinking tokens are billed as output and delay the first visible token.
Other Tools and Providers
| What this module used | Alternatives | When to prefer the alternative |
|---|---|---|
Groq (openai/gpt-oss-120b) via llm.py | Together AI, Fireworks, OpenRouter, Cerebras, cloud catalogues (AWS Bedrock, Google Vertex AI, Azure AI Foundry) | You need a specific model, region, or enterprise contract; most expose an OpenAI-compatible endpoint that llm.py can target by adding a PROVIDERS entry |
Gemini (gemini-3.5-flash) | OpenAI and Anthropic APIs | You need a particular frontier model's quality; add a provider entry or use their SDKs |
Ollama (qwen3:8b) locally | llama.cpp, LM Studio, vLLM, SGLang | llama.cpp and LM Studio for laptops; vLLM or SGLang for serving many users on GPUs (Module 13) |
| TinyLM for mechanisms | nanoGPT, Hugging Face transformers with a small open model (for example a sub-1B Qwen) | You have network access to download weights and want a model that generalizes beyond a toy corpus |
ScriptedLLM for plumbing tests | unittest.mock, VCR-style recorded responses (for example vcrpy), provider sandbox modes | You want to replay real recorded responses in tests instead of scripted ones |
| Model cards and our own probes | Hugging Face model pages, Artificial Analysis, LMArena, provider docs | Building a shortlist; always confirm with your own eval |
tokens.py offline tokenizers | tiktoken directly, Hugging Face tokenizers with each model's own tokenizer file, provider token-count endpoints | You need exact counts for a specific model family (Module 2) |
Coming Up in Module 2
You now know the model reads tokens, not words, and that it can only see what fits in its context window. Module 2 makes both concrete and puts a price on them. You will read supportdesk/tokens.py and supportdesk/pricing.py in full, measure how many tokens each Brightlane ticket costs under four real tokenizers (and why the Japanese and Hindi tickets cost more), build a context budget that reserves room for the answer, compare truncation and summarization strategies, and turn the usage numbers from your first call into a cost per ticket that Maya can put in a budget.