CourseLarge Language Models · Module 1: What a Large Language Model Actually Is · part 7 of 80
Part 7 · Module 1: What a Large Language Model Actually Is

Part G: The Model Landscape

22 min read·22 Sept 2026

Frontier, mid-tier, and small

Models cluster into rough tiers. The names change every few months, so learn the tiers and look up the current members when you choose.

TierTypical size and accessStrengthsCostsExamples (September 2026)
FrontierLargest proprietary models via API; some very large open-weight MoEsHardest reasoning, coding, long agentic tasksHighest price per token, slower, often reasoning by defaultFlagship models from OpenAI, Anthropic, and Google; DeepSeek's V4 preview line was reported at 1.6T total parameters (49B active) with a 1M-token context (Digital Applied retrospective)
Mid-tierFast API models; open-weight models of roughly 20B to 120B totalMost application work: triage, extraction, grounded answers, draftingMuch cheaper and fastergemini-3.5-flash (1M-token context, GA); openai/gpt-oss-120b on Groq
SmallRoughly 1B to 10B dense, or small MoEsClassification, extraction, routing, on-device useCheapest; weaker at open-ended reasoning and rare factsqwen3:8b on Ollama; the Qwen 3.5 small models (0.8B to 9B, Apache 2.0, reported March 2026)

"Tier" is about capability per task, not a fixed size: a well-chosen small model can match a frontier model on a narrow task like Brightlane's six-way ticket classification, and that is exactly what you will measure in later modules.

Proprietary API models vs open-weight families

SituationUse thisWhy
You need the strongest capability now, with low setup effortProprietary API modelBest models are often API-only; no hardware to run
Data must not leave your network, or regulation requires residency you controlOpen-weight model, self-hostedYou control where prompts and outputs live (Module 13)
You need to fine-tune deeply, pin a version forever, or inspect probabilitiesOpen-weight modelWeights never change under you; full access to internals
You want open-weight economics without running GPUsAn open-weight model on a hosted provider (as the course does with gpt-oss-120b on Groq)Pay per token, keep the option to self-host later
Low volume, spiky trafficAPI (either kind)Self-hosted GPUs cost money while idle

A practical point about version pinning: API model names can be retired or silently improved; open weights you download are frozen. Module 13 covers deprecations and forced migrations.

Reasoning models vs standard models: cost and latency

A reasoning model spends extra output tokens thinking before it answers. Output tokens are what drive latency (they are generated one at a time) and usually cost (output is priced higher than input). Two measured and published facts make this concrete for the course's default:

  • Groq serves gpt-oss-120b at about 500 output tokens per second, with a 131,072-token context and prices of 0.15 USD per million input tokens and 0.60 USD per million output tokens (Groq docs, checked 21 September 2026).
  • At 500 tokens per second, every 200 reasoning tokens adds about 0.4 seconds before the first visible word of the answer, and bills as 200 output tokens.

For a one-word triage label, reasoning can multiply the output tokens many times over. For a tricky refund-policy question that combines three rules, it may be what makes the answer correct. The lab below prices both profiles; Module 3 measures when extra reasoning helps and when it only wastes money.

Small language models and on-device use

Small language models (SLMs) run on a laptop, phone, or a single modest GPU. Apple's on-device foundation model is about 3 billion parameters, compressed to 2 bits per weight with quantization-aware training, and exposed to app developers through the Foundation Models framework with guided (constrained) generation and tool calling (Apple Machine Learning Research). OpenAI describes gpt-oss-20b as running within 16 GB of memory. Using the arithmetic from Part D, a 3B model at 2 bits is about 0.75 GB of weights: that is why it fits on a phone.

On-device wins on privacy (nothing leaves the device), offline use, and zero per-token cost. It loses on capability, especially rare facts and long reasoning. For Brightlane, a small local model is a reasonable candidate for the first step (triage) and a poor one for drafting policy answers without retrieval.

Model cards, benchmarks, and why leaderboard rank rarely transfers

A model card is the document published with a model: intended use, training data summary, evaluation results, limitations, licence. Read it for facts (context, cutoff, licence, languages), and read its benchmark table with suspicion. A benchmark is a fixed public test set; a leaderboard ranks models on one or more of them. Rank transfers poorly to your task for three reasons:

  1. Different task: public benchmarks test general skills; your task has its own inputs, labels, and failure costs.
  2. Contamination: benchmark questions leak into training data, so scores can measure memorization.
  3. Leaderboard dynamics: a 2025 study of a popular crowd-voted leaderboard found that some providers privately tested many variants and published only the best, and that access to leaderboard data favored those who had more of it (Singh et al., "The Leaderboard Illusion").

You can reproduce the first two effects in miniature. Task: pick the help-center article that answers a question. System A is TinyLM (generate the agent's reply, then match it to an article). System B is not a model at all: match the customer's own words against each article and pick the one with the most words in common. The "benchmark" is the 12 question templates TinyLM's corpus was built from, which is to say a contaminated benchmark. "Our tickets" are the 62 answerable Brightlane tickets.

python
"""Module 1: why a benchmark rank may not transfer to your task.

Task: pick the help-center article that answers a customer question.
Two systems:
  A. TinyLM: generate the agent's reply, then match the reply to an article.
  B. Keyword overlap: match the customer's own words to an article (no model at all).
Two test sets:
  "benchmark": the 12 question templates TinyLM's corpus was generated from (contaminated).
  "our tickets": the 62 answerable Brightlane tickets, written independently.
"""
import math
import re

import torch

from supportdesk.data import load_articles, load_tickets
from supportdesk.tinylm import SamplingParams, generate, load

torch.set_num_threads(1)  # one thread is plenty for a 1M-parameter model
model, tokenizer = load()
articles = load_articles()
STOP = set("a an and are as at be by can do does for from get how i in is it my of on or our the to "
           "we what when why will with you your this that there".split())


def words(text: str) -> set[str]:
    return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOP}


ARTICLE_WORDS = {a.id: words(a.title + " " + a.body) for a in articles}


def best_article(text: str) -> str:
    """The article sharing the most distinct words with `text` (ties: first by id)."""
    return max(ARTICLE_WORDS, key=lambda aid: len(words(text) & ARTICLE_WORDS[aid]))


def tinylm_route(question: str) -> str:
    prompt = f"Customer (Maya): Hi, {question}\nAgent (Dara):"
    reply = generate(model, tokenizer, prompt, SamplingParams(max_new_tokens=40, temperature=0,
                                                              stop=["\nCustomer"])).text
    return best_article(reply)


def keyword_route(question: str) -> str:
    return best_article(question)


BENCHMARK = [  # the corpus templates, with {plan}/{area} filled in
    ("I was charged twice for the Team plan this month. Can you refund the duplicate?", "billing-refunds"),
    ("how do I cancel my Team subscription?", "billing-refunds"),
    ("I forgot my password and the reset link expired.", "account-login"),
    ("my account is locked after too many attempts.", "account-login"),
    ("does the Team plan include SSO?", "account-sso"),
    ("our automations stopped and there is a banner about a limit.", "boards-automations"),
    ("how do I export my board data?", "exports-data"),
    ("slack notifications stopped posting.", "integrations-slack"),
    ("can I get a refund on my annual Team plan?", "billing-refunds"),
    ("where can I find my invoices?", "billing-invoices"),
    ("how much does the Business plan cost?", "billing-plans"),
    ("the board is not loading. Is there an outage?", "status-incidents"),
]
TICKETS = [(t.body, t.gold["kb_article"]) for t in load_tickets() if t.gold["answerable"]]


def wilson(k: int, n: int) -> str:
    z = 1.96
    c = (k / n + z * z / (2 * n)) / (1 + z * z / n)
    h = z * math.sqrt(k / n * (1 - k / n) / n + z * z / (4 * n * n)) / (1 + z * z / n)
    return f"{max(c - h, 0):.2f} to {min(c + h, 1):.2f}"


print(f"{'system':18} {'benchmark (n=12)':>24} {'our tickets (n=62)':>26}")
for name, route in [("A. TinyLM", tinylm_route), ("B. keyword overlap", keyword_route)]:
    cells = []
    for dataset in (BENCHMARK, TICKETS):
        k = sum(route(q) == gold for q, gold in dataset)
        cells.append(f"{k}/{len(dataset)} ({wilson(k, len(dataset))})")
    print(f"{name:18} {cells[0]:>24} {cells[1]:>26}")

Code explained

  • In simple words: run two systems on a public-looking benchmark and on our real tickets, and see whether the ranking holds.
  • What happens: words lowercases text, keeps letters and digits, and drops common stop words. best_article picks the article sharing the most distinct words with a piece of text. tinylm_route generates a greedy reply (stopping at the next customer line) and routes the reply; keyword_route routes the question directly. wilson prints a 95 percent interval for each score.
  • Comes out: real output. It takes about 5 seconds of compute; on a busy machine allow up to a minute.

textCopy

text
  system                     benchmark (n=12)         our tickets (n=62)
  A. TinyLM              12/12 (0.76 to 1.00)       10/62 (0.09 to 0.27)
  B. keyword overlap      8/12 (0.39 to 0.86)       43/62 (0.57 to 0.79)

On the benchmark, TinyLM wins 12 to 8 and would top the leaderboard. On our tickets it loses 10 to 43, and the intervals do not overlap, so this is not noise. The benchmark rewarded memorization of the exact templates (contamination) and said nothing about real customers' wording. The keyword baseline's 43 of 62 (0.69) is also a useful number in its own right: it is the floor any model-based router for Brightlane must beat, and you will meet it again.

Large models are nothing like TinyLM, but the lesson is the same one practitioners keep relearning: choose models with an evaluation built from your own data (Module 10), and use public benchmarks only to build a shortlist.

A selection framework for Brightlane

Five questions decide most model choices. Answer them in order, because the first one is a filter, not a trade-off:

  1. Capability floor: does the model pass your own eval at the required quality? If not, nothing else matters.
  2. Latency: does it meet your time budget (for example, time to first token for a chat widget, total time for background triage)?
  3. Cost: what is the cost per task times your volume?
  4. Privacy: where do prompts and outputs go, who can see them, and what do your customers and regulators require?
  5. Control: can you pin the version, fine-tune it, inspect probabilities, and keep running if the provider changes terms?

Applied to the Brightlane assistant's first two jobs (numbers marked as assumptions must be replaced with your measurements):

CriterionTicket triage (label + priority)Drafting answers from the help center
Capability floorMust beat the 0.69 keyword baseline on article routing and reach Maya's target on category accuracy (set in Module 10), on the dev splitMust be grounded in retrieved articles with no invented policy; measured with a judge and human review (Module 10)
LatencyBackground job; seconds are fineAgent is waiting; aim for first words within about a second (Module 3)
CostHigh volume, tiny outputs: small or mid-tier standard model, reasoning off or lowLower volume, longer outputs: a mid-tier model; reasoning only if evals show it helps
PrivacyTickets contain names, emails, invoice ids: redact before sending or self-host (Module 11)Same, plus drafts must not leak one customer's data to another
ControlPin the model version; keep a fallback provider (Module 13)Pin the version and re-run the eval on every model change
Starting candidatesqwen3:8b locally, openai/gpt-oss-120b at low effort on Groqopenai/gpt-oss-120b or gemini-3.5-flash, with retrieval
SituationUse thisWhy
You do not yet have an eval for the taskBuild the eval first, with the cheapest plausible modelWithout it you cannot tell whether a pricier model is better
Two models pass the floorPick the cheaper and faster oneCapability above the floor is not worth paying for twice
Only a frontier model passesUse it, and log inputs and outputs to build a fine-tuning or routing dataset laterModules 9 and 13 show how to move traffic to cheaper models safely
Privacy rules forbid sending ticket text outSelf-hosted open weights, or redaction before any API callPrivacy is a hard constraint, not a weight in a score

Module Lab

The lab combines the module into one script that writes a measured fact sheet for a model: spec and memory (Part D), what training did (Part C), behavior probes for phrasing, hallucination, and calibration (Part F), a check on the real Brightlane task against a no-model baseline (Part G), and the monthly cost of hosted candidates for the same job. Its rule is the module's rule: every line is either measured, taken from a cited source, or labelled as an assumption.

python
"""Module 1 Lab: write a one-page fact sheet for a model, measured rather than assumed.

It inspects TinyLM (spec, memory, training), probes its behaviour (memorized vs reworded
questions, hallucination, calibration), checks it on the real Brightlane task, and prices
the hosted candidates for the same job. Output: reports/m01_model_card.json plus a printout.
"""
import json
import math
import re
from pathlib import Path

import torch

from supportdesk.data import load_articles, load_tickets
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.tinylm import MODEL_DIR, SamplingParams, generate, load

torch.set_num_threads(1)  # one thread is plenty for a 1M-parameter model
model, tokenizer = load()
card: dict = {"model": "tinylm-base"}

# 1. Spec and memory -------------------------------------------------------------
params = model.num_parameters()
card["spec"] = {"parameters": params, "layers": model.cfg.n_layers, "heads": model.cfg.n_heads,
                "d_model": model.cfg.d_model, "context_tokens": model.cfg.context,
                "embedding_rows": model.cfg.vocab_size, "tokenizer_vocab": tokenizer.get_vocab_size()}
card["weights_mb"] = {dtype: round(params * b / 1e6, 2) for dtype, b in
                      {"fp32": 4, "bf16": 2, "int8": 1, "4-bit": 0.5}.items()}

# 2. What pretraining did ----------------------------------------------------------
log = json.loads((MODEL_DIR / "training_log.json").read_text())
best = min(log, key=lambda r: r["val_loss"])
card["training"] = {"best_step": best["step"], "best_val_loss": best["val_loss"],
                    "final_train_loss": log[-1]["train_loss"], "final_val_loss": log[-1]["val_loss"],
                    "epochs_over_train_tokens": round(log[-1]["tokens_seen"] / 156_620, 1),
                    "overfitting": log[-1]["val_loss"] > best["val_loss"]}


@torch.no_grad()
def next_probs(prompt: str) -> torch.Tensor:
    return torch.softmax(model(torch.tensor([tokenizer.encode(prompt).ids]))[0, -1], dim=-1)


# 3. Behaviour probes ----------------------------------------------------------------
# (question as trained, reworded question, correct first token of the reply)
PAIRS = [
    ("my account is locked after too many attempts.", "it says my account is locked. Why?", " After"),
    ("I was charged twice for the Team plan this month. Can you refund the duplicate?",
     "you billed my card two times this month for Team, please refund one.", " Sorry"),
    ("how do I export my board data?", "how can I download my board as a spreadsheet?", " Export"),
    ("I forgot my password and the reset link expired.", "my password link does not work anymore.", " Reset"),
    ("does the Team plan include SSO?", "can we log in with Okta on Team?", " SSO"),
]
probe = {"as_trained": [], "reworded": []}
for trained, reworded, correct in PAIRS:
    target = tokenizer.encode(correct).ids[0]
    for key, q in (("as_trained", trained), ("reworded", reworded)):
        probs = next_probs(f"Customer (Maya): Hi, {q}\nAgent (Dara):")
        probe[key].append((int(probs.argmax()) == target, probs.max().item()))
card["phrasing"] = {k: f"{sum(ok for ok, _ in v)}/{len(v)} top-1 correct" for k, v in probe.items()}

france = next_probs("The capital of France is")
card["hallucination"] = {"prompt": "The capital of France is",
                         "top_token": tokenizer.decode([int(france.argmax())]),
                         "top_prob": round(france.max().item(), 3)}

pairs = [c for v in probe.values() for c in v]
card["calibration_on_probes"] = {
    "n": len(pairs), "accuracy": round(sum(ok for ok, _ in pairs) / len(pairs), 2),
    "mean_confidence": round(sum(p for _, p in pairs) / len(pairs), 2),
    "confident_wrong(p>=0.8)": sum(1 for ok, p in pairs if p >= 0.8 and not ok)}

# 4. The real task: pick the right help-center article for each answerable ticket ----
STOP = set("a an and are as at be by can do does for from get how i in is it my of on or our the to "
           "we what when why will with you your this that there".split())
words = lambda t: {w for w in re.findall(r"[a-z0-9]+", t.lower()) if w not in STOP}  # noqa: E731
ARTICLE_WORDS = {a.id: words(a.title + " " + a.body) for a in load_articles()}
best_article = lambda t: max(ARTICLE_WORDS, key=lambda a: len(words(t) & ARTICLE_WORDS[a]))  # noqa: E731
tickets = [t for t in load_tickets() if t.gold["answerable"]]
greedy = SamplingParams(max_new_tokens=40, temperature=0, stop=["\nCustomer"])
tiny_hits = sum(best_article(generate(model, tokenizer, f"Customer (Maya): Hi, {t.body}\nAgent (Dara):",
                                      greedy).text) == t.gold["kb_article"] for t in tickets)
rule_hits = sum(best_article(t.body) == t.gold["kb_article"] for t in tickets)
card["task_article_routing"] = {"n": len(tickets), "tinylm": tiny_hits, "keyword_rule": rule_hits}

# 5. Price the hosted candidates for the same job ------------------------------------
TICKETS_PER_MONTH = 3000                        # assumption for Brightlane; use your own volume
STANDARD = Usage(input_tokens=90, output_tokens=13)                            # measured with the stand-in prompt
REASONING = Usage(input_tokens=90, output_tokens=13 + 200, reasoning_tokens=200)  # assumption: 200 thinking tokens
card["monthly_cost_usd"] = {
    f"{m} ({'reasoning' if u is REASONING else 'standard'})": round(cost_usd(u, m) * TICKETS_PER_MONTH, 2)
    for m, u in [("openai/gpt-oss-120b", REASONING), ("gemini-3.5-flash", REASONING),
                 ("llama-3.1-8b-instant", STANDARD), ("qwen3:8b", STANDARD)]
    if m in PRICES}

out = Path("reports/m01_model_card.json")
out.parent.mkdir(exist_ok=True)
out.write_text(json.dumps(card, indent=2) + "\n")
for section, value in card.items():
    print(f"{section}: {value}")
k, n = rule_hits, len(tickets)
print(f"\nkeyword rule accuracy {k / n:.2f} +/- {1.96 * math.sqrt(k / n * (1 - k / n) / n):.2f} (95%, normal approx.)")
print("saved", out)

Code explained

  • In simple words: a model report card generator. Point it at a model, and it fills in each section with a number you can defend.
  • What happens:
  • Section 1 reads the spec from the loaded model and tokenizer (not from memory or marketing) and computes weight memory at four precisions.
  • Section 2 summarizes training_log.json: best validation step, final losses, how many times the training set was repeated (tokens_seen / 156,620), and whether validation loss rose after its best point (overfitting).
  • Section 3 runs five intents twice, once in training wording and once reworded, and records top-1 correctness and confidence; it checks the France prompt for a hallucination; and it summarizes calibration on those ten probes, counting confident-wrong answers (probability at least 0.8 and wrong).
  • Section 4 runs the article-routing task on all 62 answerable tickets for TinyLM and for the keyword rule.
  • Section 5 prices hosted candidates with pricing.cost_usd (the price table is explained in Module 2). The standard profile uses the 90 input and 13 output tokens counted for the Part A triage prompt. Two numbers are assumptions, marked in the code: 3,000 tickets per month, and 200 reasoning tokens per call for reasoning models. Replace both with your own measurements from result.usage. The local qwen3:8b shows 0.0 because there is no per-token price; your hardware and electricity are not free.
  • Everything is saved to reports/m01_model_card.json.
  • Comes out: real output from this build (runtime is under a minute; the routing step dominates).
text
  model: tinylm-base
  spec: {'parameters': 1071872, 'layers': 4, 'heads': 4, 'd_model': 128, 'context_tokens': 128, 'embedding_rows': 2048, 'tokenizer_vocab': 1503}
  weights_mb: {'fp32': 4.29, 'bf16': 2.14, 'int8': 1.07, '4-bit': 0.54}
  training: {'best_step': 400, 'best_val_loss': 0.6553, 'final_train_loss': 0.2636, 'final_val_loss': 0.6799, 'epochs_over_train_tokens': 26.2, 'overfitting': True}
  phrasing: {'as_trained': '5/5 top-1 correct', 'reworded': '1/5 top-1 correct'}
  hallucination: {'prompt': 'The capital of France is', 'top_token': ' not', 'top_prob': 0.196}
  calibration_on_probes: {'n': 10, 'accuracy': 0.6, 'mean_confidence': 0.65, 'confident_wrong(p>=0.8)': 0}
  task_article_routing: {'n': 62, 'tinylm': 10, 'keyword_rule': 43}
  monthly_cost_usd: {'openai/gpt-oss-120b (reasoning)': 0.42, 'gemini-3.5-flash (reasoning)': 6.16, 'llama-3.1-8b-instant (standard)': 0.02, 'qwen3:8b (standard)': 0.0}

  keyword rule accuracy 0.69 +/- 0.11 (95%, normal approx.)
  saved reports/m01_model_card.json

Reading the fact sheet as Maya would:

  • TinyLM overfit (best step 400) after reading its training set 26 times.
  • It answers 5 of 5 trained phrasings and 1 of 5 rewordings correctly: it matches strings, it does not understand requests.
  • On the real routing task it gets 10 of 62, against 43 of 62 for a keyword rule that needs no model at all. With n = 62, the keyword rule's accuracy is 0.69 plus or minus 0.11.
  • For the hosted candidates, cost is not the bottleneck at Brightlane's assumed volume: even the most expensive option in the table is a few dollars a month under these assumptions. The capability floor, latency, and privacy will decide, which is why the next modules build measurement before anything else.

To run the lab against a different model directory (for example, one you fine-tune in Module 9), change the load() call to load("models/<your-model>").

Project Milestone

After this module, your supportdesk repository contains:

  • The dataset: data/tickets.jsonl (72 labelled tickets, dev 48 and test 24) and data/kb/ (12 articles), loaded through supportdesk/data.py.
  • The provider helper supportdesk/llm.py, configured by LLM_PROVIDER, LLM_MODEL, and a key in your environment, and the stand-in supportdesk/stand_in.py for key-free tests.
  • The pretrained TinyLM in models/tinylm-base/, which you have probed but not changed.
  • Module 1 scripts in examples/: m01_explore_data.py, m01_first_call.py, m01_next_token.py, m01_training_log.py, m01_memorize_hallucinate.py, m01_base_model.py, m01_model_spec.py, m01_token_view.py, m01_phrasing.py, m01_calibration.py, m01_leaderboard.py, and m01_lab.py.
  • tests/test_m01_foundations.py (8 tests) and reports/m01_model_card.json.
  • Two numbers you will reuse: the keyword baseline for article routing (43 of 62 answerable tickets) and the triage prompt's token counts (90 input, 13 output with the stand-in's estimate).

If you have a key, run examples/m01_first_call.py against your provider now and note the real usage and latency. Module 2 turns them into cost per ticket.

Interview Questions

1. What does a large language model actually compute?
A probability distribution over its vocabulary for the next token, given the tokens so far. Generation calls that repeatedly, appending each chosen token. Everything else (chat roles, tools, reasoning) is built on top by formatting the input and by training the model to produce useful continuations. In practice this means every output is a plausible continuation, not a retrieved fact, so anything that must be true needs grounding or verification.

2. Why do models hallucinate, and why can't a vendor just fix it?
The training objective rewards assigning high probability to the text that actually came next in training data. It does not reward truth, and it has no built-in "I don't know": a distribution must put its probability mass somewhere, and generation always emits a token. For rare or unseen facts, the most plausible continuation is often wrong. Post-training reduces the rate and teaches some abstention, but the cause is the objective itself. The practical defenses are retrieval of trusted sources, constrained outputs, verification by code, and evaluation.

3. What is the difference between a base model and an instruction-tuned model? When would you use a base model?
A base model continues documents: it ignores instructions, does not stop at the end of a turn, and may write the other speaker's lines. Instruction tuning (SFT, then usually preference optimization) trains it on conversations ending with ideal answers, so it follows instructions, formats, and stops. Use a base model to start your own fine-tune, to score text likelihood, or for pure continuation in a fixed style; use an instruction-tuned model for almost all application work.

4. Walk through the stages of an LLM's lifecycle and name one behavior each stage causes.
Pretraining on trillions of tokens: knowledge, skills, and hallucination on rare facts. Mid-training or continued pretraining on curated or domain data: better math and code, longer context. Supervised fine-tuning: instruction following and output formats. Preference optimization (RLHF, DPO): the assistant's style, refusals, sycophancy, and typically worse calibration. Reasoning training with verifiable rewards: long thinking before answers, better math and code, more tokens and latency.

5. How much GPU memory does an 8B model need?
Weights alone: parameters times bytes per parameter. At bf16 that is 8e9 x 2 = 16 GB; at 8-bit about 8 GB; at 4-bit about 4 GB. Then add the KV cache, which grows with context length and batch size (for Llama 3.1 8B at bf16, about 131 KB per token, so about 4.3 GB at 32k tokens), plus runtime working memory. So an 8B model at 4-bit with a moderate context fits a 16 GB GPU; at bf16 with long contexts it needs more than 24 GB.

6. What do "total" and "active" parameters mean for a mixture-of-experts model, and why does it matter?
An MoE layer has many expert MLPs but routes each token to only a few. Total parameters (116.83B for gpt-oss-120b) must all be in memory. Active parameters (5.13B) are used per token and set the compute per token, which largely drives speed. So an MoE can be as fast as a much smaller dense model while needing the memory of a large one.

7. A model is described as "open". What do you check before using it commercially?
The licence: Apache 2.0 or MIT is permissive; custom licences (for example Llama 3.1's) can add user caps, naming and attribution requirements, and acceptable-use rules. Whether it is open weights only or open source in the OSI sense (training data information and code published too). And for distillation or fine-tuning on another model's outputs, the terms of that provider.

8. Why do LLMs struggle to count letters or do long arithmetic?
They see tokens, not characters or digits. "strawberry" arrives as three tokens like st, raw, berry, so the letters are not directly visible, and a long number is split into chunks that do not align with place value. The fix is to route exact work to code (tools) and let the model handle language; reasoning models help by writing steps out, but results still need verification.

9. What is calibration, and how would you check it for a triage classifier?
A model is calibrated when its stated confidence matches its accuracy. Collect confidence (token probability, or agreement across several samples) and correctness on a labelled set, bin by confidence, and compare mean confidence with accuracy per bin (a reliability table), with confidence intervals because bins are small. Use it to choose an escalation threshold, and recalibrate whenever the model or prompt changes. Never let confidence alone gate an irreversible action.

10. A new model tops a public leaderboard. Should you switch your production system to it?
Not on that evidence. Leaderboards measure different tasks, can be inflated by contamination and selective reporting, and rank transfer to a specific task is weak. Run your own eval on your own data, check latency, cost, privacy, and licence, and switch only if it clears the capability floor at acceptable cost. In this module a model that won a contaminated benchmark 12 to 8 lost on real tickets 10 to 43.

11. What is a knowledge cutoff, and how do you design around it?
The date after which the model saw no training data; it has no knowledge of later events and cannot tell that it is out of date. Keep changing facts (prices, known issues, roadmaps) in documents retrieved at request time, date them, and let the model phrase answers from them. Fine-tune for behavior, not for facts that will change.

12. When would you choose a reasoning model over a standard one?
When the task needs several dependent steps (math, code, policy questions that combine rules) and your eval shows reasoning improves accuracy enough to justify extra output tokens and latency. For short, high-volume tasks like classification, use a standard model or the lowest reasoning effort, because thinking tokens are billed as output and delay the first visible token.

Other Tools and Providers

What this module usedAlternativesWhen to prefer the alternative
Groq (openai/gpt-oss-120b) via llm.pyTogether AI, Fireworks, OpenRouter, Cerebras, cloud catalogues (AWS Bedrock, Google Vertex AI, Azure AI Foundry)You need a specific model, region, or enterprise contract; most expose an OpenAI-compatible endpoint that llm.py can target by adding a PROVIDERS entry
Gemini (gemini-3.5-flash)OpenAI and Anthropic APIsYou need a particular frontier model's quality; add a provider entry or use their SDKs
Ollama (qwen3:8b) locallyllama.cpp, LM Studio, vLLM, SGLangllama.cpp and LM Studio for laptops; vLLM or SGLang for serving many users on GPUs (Module 13)
TinyLM for mechanismsnanoGPT, Hugging Face transformers with a small open model (for example a sub-1B Qwen)You have network access to download weights and want a model that generalizes beyond a toy corpus
ScriptedLLM for plumbing testsunittest.mock, VCR-style recorded responses (for example vcrpy), provider sandbox modesYou want to replay real recorded responses in tests instead of scripted ones
Model cards and our own probesHugging Face model pages, Artificial Analysis, LMArena, provider docsBuilding a shortlist; always confirm with your own eval
tokens.py offline tokenizerstiktoken directly, Hugging Face tokenizers with each model's own tokenizer file, provider token-count endpointsYou need exact counts for a specific model family (Module 2)

Coming Up in Module 2

You now know the model reads tokens, not words, and that it can only see what fits in its context window. Module 2 makes both concrete and puts a price on them. You will read supportdesk/tokens.py and supportdesk/pricing.py in full, measure how many tokens each Brightlane ticket costs under four real tokenizers (and why the Japanese and Hindi tickets cost more), build a context budget that reserves room for the answer, compare truncation and summarization strategies, and turn the usage numbers from your first call into a cost per ticket that Maya can put in a budget.