CourseLarge Language Models · Module 1: What a Large Language Model Actually Is · part 1 of 80
Part 1 · Module 1: What a Large Language Model Actually Is

Part A: The Project and the Toolkit

19 min read·22 Sept 2026

By the end of this module, you'll have:

  • Run a real language model (TinyLM, 1.07 million parameters) on your own CPU, read its next-token probabilities, and generated a support reply one token at a time.
  • Read a real training log, pointed at the step where the model started to overfit, and measured how much of its output is memorized text.
  • Read a model spec (parameters, layers, context window, vocabulary) for TinyLM and for published open-weight models, and turned parameter counts into gigabytes of memory.
  • Measured four limits yourself: tokenization hiding letters and digits, sensitivity to phrasing, calibration, and hallucination outside the training data.
  • Built a selection framework (capability floor, latency, cost, privacy, control) and applied it to the Brightlane support assistant, plus a one-page measured fact sheet for a model.
  • Set up the supportdesk repository with the dataset and three helper files (data.py, llm.py, stand_in.py) that every later module imports.

Prerequisites: Working Python (functions, classes, dataclasses, running scripts from a terminal). No machine learning background. A free API key (Groq or Gemini) or a local Ollama install is optional for this module: every example runs without one.

Where we are: This is the first module. Before we prompt, retrieve, fine-tune, or evaluate anything, we need an honest picture of the object itself: what a large language model computes, how it was made, what it is reliably good and bad at, and how to choose one.

How this module is organized

PartWhat it covers
Part A: The Project and the ToolkitThe Brightlane support desk, the ticket dataset, and the helpers data.py, llm.py, and stand_in.py; a first triage call
Part B: Next-Token PredictionTinyLM, next-token probabilities, generating a reply token by token, why one objective yields many skills
Part C: What Pretraining LearnedThe training log, loss, overfitting, memorization, and hallucination as a property of the objective
Part D: Reading a Model SpecParameters, layers, context window, vocabulary; checking published configs; memory arithmetic
Part E: The LifecyclePretraining, mid-training, supervised fine-tuning, preference optimization, reasoning training; base vs tuned models; open weights vs open source
Part F: Capabilities and Limits, MeasuredTokens vs letters and digits, phrasing sensitivity, calibration, knowledge cutoffs, self-knowledge
Part G: The Model LandscapeModel tiers, proprietary vs open weights, reasoning vs standard, small on-device models, benchmarks, a selection framework

Setting up

The course has one code repository, supportdesk. Every module adds example scripts to examples/ and tests to tests/, and imports the shared helpers under supportdesk/. All commands run from the repository root with PYTHONPATH=. so that import supportdesk works without installing anything.

python
cd supportdesk
python3.11 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
export PYTHONPATH=.
python -m pytest -q tests/test_m01_foundations.py

Code explained

  • In simple words: make an isolated Python environment, install the pinned libraries, and run this module's tests to prove the setup works.
  • What happens: venv creates a private Python 3.11 environment so course libraries do not collide with anything else on your machine. requirements.txt pins exact versions (openai 3.16.2, pydantic 2.13.5, tiktoken 0.14.0, tokenizers 0.23.2, torch 2.14.0, numpy 2.4.6, pytest 9.1.1). PYTHONPATH=. lets Python find the supportdesk package in the current folder. The tests load the dataset, load TinyLM, and check a few behaviors you will see in this module.
  • Comes out: pytest prints one dot per passing test. Timing depends on your machine.
text
  ........                                                                 [100%]
  8 passed in 13.28s

The examples in this module are complete scripts in examples/. Run them in the order they appear; later ones reuse ideas (not variables) from earlier ones. PyTorch examples call torch.set_num_threads(1): a 1-million-parameter model does not benefit from more threads, and one thread keeps timing steady on a busy laptop.

Part A: The Project and the Toolkit

The Brightlane support desk

Brightlane is a fictional project-management SaaS with four plans: Free, Team (12 USD per user per month), Business (24 USD), and Enterprise (custom pricing). Its support desk receives tickets in several languages. Maya, the support lead, wants an assistant that triages incoming tickets, answers from the help center, and drafts replies that her team reviews before sending. Over the course you will build that assistant, then evaluate it, secure it, adapt it, and operate it.

Two data sources ship with the repository:

  • data/tickets.jsonl: 72 hand-written tickets. Each has a gold label, meaning the correct answer a human decided in advance: a category, a priority, which help-center article answers it, and whether it is answerable from the help center at all. The tickets are split into a dev set (48 tickets you may look at while building) and a test set (24 tickets you keep for the final check, so you do not tune to them by accident).
  • data/kb/*.md: 12 help-center articles. The file name without .md is the article id, for example billing-refunds.

Here is one line of tickets.jsonl, pretty-printed:

json
{
  "id": "T-1001",
  "subject": "Charged twice this month",
  "body": "Hi, my card was charged 288 USD twice on 3 September for the Team plan (invoice INV-2026-004512). Please refund the duplicate.",
  "customer_tier": "team",
  "language": "en",
  "split": "dev",
  "gold": {"category": "billing", "priority": "high", "kb_article": "billing-refunds", "answerable": true}
}

Code explained

  • In simple words: one ticket is one JSON object on one line, like a row in a spreadsheet with a nested "answer key".
  • What happens: subject and body are what the customer wrote. customer_tier is their plan. language is an ISO code (en, es, de, ja, hi). split says whether the ticket belongs to dev or test. gold holds the labels: category is one of six values, priority one of four, kb_article names the article that answers it (or null when no article does), and answerable says whether the help center covers it.
  • Comes out: nothing runs here; this is the data format. In the file each ticket sits on a single line (the JSON Lines format), which makes it easy to append and stream.

Now load everything and look at the shape of the data.

python
"""Module 1: a first look at the Brightlane dataset the whole course uses."""
from collections import Counter

from supportdesk.data import CATEGORIES, PRIORITIES, get_article, load_articles, load_tickets

tickets = load_tickets()
dev, test = load_tickets("dev"), load_tickets("test")
articles = load_articles()

print(f"tickets: {len(tickets)} (dev {len(dev)}, test {len(test)})")
print("languages:", dict(Counter(t.language for t in tickets).most_common()))
print("categories:", {c: sum(t.gold["category"] == c for t in tickets) for c in CATEGORIES})
print("priorities:", {p: sum(t.gold["priority"] == p for t in tickets) for p in PRIORITIES})
print("answerable from the help center:", sum(t.gold["answerable"] for t in tickets))
print(f"articles: {len(articles)}:", ", ".join(a.id for a in articles))

first = tickets[0]
print("\n--- one ticket, as an agent reads it ---")
print(first.id, first.customer_tier, first.language)
print(first.text)
print("gold:", first.gold)

article = get_article(first.gold["kb_article"])
print("\n--- the article that answers it ---")
print(article.title, article.tags)
print(article.body[:160] + "...")

Code explained

  • In simple words: count what is in the dataset and print one ticket next to the article that answers it, the way a new support agent would study the queue on day one.
  • What happens: load_tickets() returns all 72 tickets as Ticket objects; load_tickets("dev") and load_tickets("test") filter by split. Counter tallies languages. The category and priority loops count gold labels using the canonical CATEGORIES and PRIORITIES tuples, so a typo in a label would show up as a missing count. ticket.text joins subject and body. get_article fetches the gold article by id.
  • Comes out: the dataset is small and uneven: 64 of 72 tickets are English, only 4 are urgent, and 10 cannot be answered from the help center. Keep those numbers in mind. With 72 tickets, one ticket is 1.4 percentage points of accuracy, so small differences you measure later will often be noise.
python
  tickets: 72 (dev 48, test 24)
  languages: {'en': 64, 'de': 3, 'es': 2, 'ja': 2, 'hi': 1}
  categories: {'billing': 16, 'cancellation': 6, 'account_access': 17, 'bug': 9, 'how_to': 18, 'feature_request': 6}
  priorities: {'low': 24, 'normal': 31, 'high': 13, 'urgent': 4}
  answerable from the help center: 62
  articles: 12: account-login, account-sso, billing-invoices, billing-plans, billing-refunds, boards-automations, data-privacy, exports-data, feature-requests, integrations-slack, mobile-app, status-incidents

  --- one ticket, as an agent reads it ---
  T-1001 team en
  Subject: Charged twice this month

  Hi, my card was charged 288 USD twice on 3 September for the Team plan (invoice INV-2026-004512). Please refund the duplicate.
  gold: {'category': 'billing', 'priority': 'high', 'kb_article': 'billing-refunds', 'answerable': True}

  --- the article that answers it ---
  Refunds and cancellations ('billing', 'cancellation')
  You can cancel at any time from Settings > Billing > Cancel plan. Monthly plans stay active until the end of the current billing period and are not refunded pro...

The shared helpers

Three files are introduced in this module and imported by every later one. You do not need to memorize them. Skim each, read the explanation, and come back when a later module uses a function.

supportdesk/data.py

python
"""Load the Brightlane support dataset: tickets with gold labels and help-center articles."""
from __future__ import annotations

import json
from dataclasses import dataclass
from pathlib import Path

DATA_DIR = Path(__file__).resolve().parents[1] / "data"
CATEGORIES = ("billing", "cancellation", "account_access", "bug", "how_to", "feature_request")
PRIORITIES = ("low", "normal", "high", "urgent")


@dataclass(frozen=True)
class Ticket:
    id: str
    subject: str
    body: str
    customer_tier: str
    language: str
    split: str
    gold: dict

    @property
    def text(self) -> str:
        """Subject and body as one string, the way an agent reads a ticket."""
        return f"Subject: {self.subject}\n\n{self.body}"


@dataclass(frozen=True)
class Article:
    id: str
    title: str
    tags: tuple[str, ...]
    body: str


def load_tickets(split: str | None = None) -> list[Ticket]:
    """All tickets, or only the 'dev' or 'test' split. Order is stable."""
    tickets = []
    with (DATA_DIR / "tickets.jsonl").open(encoding="utf-8") as f:
        for line in f:
            t = Ticket(**json.loads(line))
            if split is None or t.split == split:
                tickets.append(t)
    return tickets


def load_articles() -> list[Article]:
    """Help-center articles from data/kb, sorted by id."""
    articles = []
    for path in sorted((DATA_DIR / "kb").glob("*.md")):
        text = path.read_text(encoding="utf-8")
        _, header, body = text.split("---\n", 2)
        meta = dict(line.split(": ", 1) for line in header.strip().splitlines())
        tags = tuple(t.strip() for t in meta.get("tags", "").strip("[]").split(",") if t.strip())
        articles.append(Article(path.stem, meta["title"], tags, body.strip()))
    return articles


def get_article(article_id: str) -> Article:
    for article in load_articles():
        if article.id == article_id:
            return article
    raise KeyError(f"No article {article_id!r}")

Code explained

  • In simple words: data.py turns the files in data/ into Python objects so no example ever parses JSON or Markdown by hand.
  • What happens:
  • DATA_DIR, CATEGORIES, PRIORITIES: the data folder (found relative to this file, so it works from any working directory) and the fixed label sets. Later modules build schemas and metrics from these tuples, so there is one source of truth.
  • Ticket: a frozen dataclass (its fields cannot be changed after creation, which prevents a script from accidentally editing gold labels). The text property formats subject and body the way an agent reads them.
  • Article: the same idea for a help-center article: id, title, tags, and body.
  • load_tickets(split): reads the JSON Lines file line by line, builds a Ticket from each, and keeps only the requested split. Order is stable, so results are reproducible.
  • load_articles(): reads every kb/*.md file. Each file starts with a small header between --- lines (called front matter); the function splits it off, parses title and tags, and keeps the rest as the body.
  • get_article(article_id): finds one article by id and raises a clear KeyError if it does not exist.
  • Comes out: nothing when imported; the functions return the objects you saw in the previous example.

supportdesk/llm.py

python
"""One helper for three LLM providers: Groq, Gemini, and Ollama.

All three expose an OpenAI-compatible Chat Completions endpoint, so a single
client library (`openai`) talks to each of them. Choose a provider with the
LLM_PROVIDER environment variable (groq, gemini, ollama) and override the
model with LLM_MODEL. Keys come from GROQ_API_KEY and GEMINI_API_KEY; Ollama
runs locally and needs no key.

Every call returns a ChatResult with the text, any tool calls, token usage,
and latency, because the course measures cost and speed on every example.
"""
from __future__ import annotations

import json
import os
import time
from collections.abc import Iterator
from dataclasses import dataclass, field
from typing import Any

from openai import OpenAI

PROVIDERS: dict[str, dict[str, str]] = {
    "groq": {
        "base_url": "https://api.groq.com/openai/v1",
        "key_env": "GROQ_API_KEY",
        "default_model": "openai/gpt-oss-120b",
    },
    "gemini": {
        "base_url": "https://generativelanguage.googleapis.com/v1beta/openai/",
        "key_env": "GEMINI_API_KEY",
        "default_model": "gemini-3.5-flash",
    },
    "ollama": {
        "base_url": "http://localhost:11434/v1",
        "key_env": "",
        "default_model": "qwen3:8b",
    },
}


@dataclass
class ToolCall:
    id: str
    name: str
    arguments: dict[str, Any]
    raw_arguments: str = ""


@dataclass
class Usage:
    input_tokens: int = 0
    output_tokens: int = 0
    cached_tokens: int = 0
    reasoning_tokens: int = 0


@dataclass
class ChatResult:
    text: str
    tool_calls: list[ToolCall] = field(default_factory=list)
    usage: Usage = field(default_factory=Usage)
    latency_ms: float = 0.0
    finish_reason: str = ""
    model: str = ""
    provider: str = ""

    def as_message(self) -> dict[str, Any]:
        """The assistant message to append to the conversation history."""
        message: dict[str, Any] = {"role": "assistant", "content": self.text or ""}
        if self.tool_calls:
            message["tool_calls"] = [
                {
                    "id": call.id,
                    "type": "function",
                    "function": {"name": call.name, "arguments": call.raw_arguments or json.dumps(call.arguments)},
                }
                for call in self.tool_calls
            ]
        return message


def resolve(provider: str | None = None, model: str | None = None) -> tuple[str, str]:
    """Pick the provider and model from arguments first, then the environment, then defaults."""
    name = (provider or os.environ.get("LLM_PROVIDER", "groq")).lower()
    if name not in PROVIDERS:
        raise ValueError(f"Unknown LLM_PROVIDER {name!r}. Choose one of {sorted(PROVIDERS)}.")
    return name, model or os.environ.get("LLM_MODEL") or PROVIDERS[name]["default_model"]


def make_client(provider: str) -> OpenAI:
    """Build an OpenAI-compatible client for one provider, reading its key from the environment."""
    settings = PROVIDERS[provider]
    if settings["key_env"]:
        api_key = os.environ.get(settings["key_env"], "")
        if not api_key:
            raise RuntimeError(f"Set {settings['key_env']} in your environment to use {provider}.")
        base_url = settings["base_url"]
    else:
        api_key = "ollama"
        base_url = os.environ.get("OLLAMA_BASE_URL", settings["base_url"])
    return OpenAI(api_key=api_key, base_url=base_url, max_retries=2, timeout=120.0)


def _usage(raw: Any) -> Usage:
    if raw is None:
        return Usage()
    prompt_details = getattr(raw, "prompt_tokens_details", None)
    completion_details = getattr(raw, "completion_tokens_details", None)
    return Usage(
        input_tokens=raw.prompt_tokens or 0,
        output_tokens=raw.completion_tokens or 0,
        cached_tokens=(getattr(prompt_details, "cached_tokens", 0) or 0) if prompt_details else 0,
        reasoning_tokens=(getattr(completion_details, "reasoning_tokens", 0) or 0) if completion_details else 0,
    )


def _parse_arguments(raw: str) -> dict[str, Any]:
    try:
        value = json.loads(raw or "{}")
    except json.JSONDecodeError:
        return {"_unparseable": raw}
    return value if isinstance(value, dict) else {"_value": value}


def chat(
    messages: list[dict[str, Any]],
    *,
    provider: str | None = None,
    model: str | None = None,
    temperature: float | None = 0.0,
    top_p: float | None = None,
    max_tokens: int | None = None,
    stop: list[str] | None = None,
    seed: int | None = None,
    tools: list[dict[str, Any]] | None = None,
    tool_choice: str | dict[str, Any] | None = None,
    response_format: dict[str, Any] | None = None,
    reasoning_effort: str | None = None,
    extra: dict[str, Any] | None = None,
) -> ChatResult:
    """Send one Chat Completions request and return text, tool calls, usage, and latency.

    Parameters left as None are not sent, so each provider uses its own default.
    `extra` passes provider-specific fields through unchanged.
    """
    name, model_id = resolve(provider, model)
    kwargs: dict[str, Any] = {"model": model_id, "messages": messages}
    optional = {
        "temperature": temperature, "top_p": top_p, "max_tokens": max_tokens, "stop": stop,
        "seed": seed, "tools": tools, "tool_choice": tool_choice,
        "response_format": response_format, "reasoning_effort": reasoning_effort,
    }
    kwargs.update({k: v for k, v in optional.items() if v is not None})
    if extra:
        kwargs["extra_body"] = extra
    started = time.perf_counter()
    response = make_client(name).chat.completions.create(**kwargs)
    latency_ms = (time.perf_counter() - started) * 1000
    choice = response.choices[0]
    calls = [
        ToolCall(c.id, c.function.name, _parse_arguments(c.function.arguments), c.function.arguments or "")
        for c in (choice.message.tool_calls or [])
        if c.type == "function"
    ]
    return ChatResult(
        text=choice.message.content or "",
        tool_calls=calls,
        usage=_usage(response.usage),
        latency_ms=round(latency_ms, 1),
        finish_reason=choice.finish_reason or "",
        model=model_id,
        provider=name,
    )


@dataclass
class StreamStats:
    ttft_ms: float = 0.0
    total_ms: float = 0.0
    chunks: int = 0


def stream_chat(
    messages: list[dict[str, Any]],
    stats: StreamStats | None = None,
    *,
    provider: str | None = None,
    model: str | None = None,
    temperature: float | None = 0.0,
    max_tokens: int | None = None,
) -> Iterator[str]:
    """Yield text pieces as they arrive. Fills `stats` with time to first token and total time."""
    name, model_id = resolve(provider, model)
    kwargs: dict[str, Any] = {"model": model_id, "messages": messages, "stream": True}
    if temperature is not None:
        kwargs["temperature"] = temperature
    if max_tokens is not None:
        kwargs["max_tokens"] = max_tokens
    stats = stats if stats is not None else StreamStats()
    started = time.perf_counter()
    for chunk in make_client(name).chat.completions.create(**kwargs):
        if not chunk.choices:
            continue
        piece = chunk.choices[0].delta.content or ""
        if piece:
            if stats.chunks == 0:
                stats.ttft_ms = round((time.perf_counter() - started) * 1000, 1)
            stats.chunks += 1
            yield piece
    stats.total_ms = round((time.perf_counter() - started) * 1000, 1)

Code explained

  • In simple words: llm.py is one telephone that can dial three different companies. You always speak the same language (a list of messages), and it always hands back the same kind of answer (a ChatResult).
  • What happens:
  • PROVIDERS: the three supported providers. Groq and Gemini are hosted services with free tiers; Ollama runs models on your own machine. All three speak the OpenAI-compatible Chat Completions format (a de facto standard request shape: a model name plus a list of {"role", "content"} messages), so one client library covers them. The default models are openai/gpt-oss-120b on Groq, gemini-3.5-flash on Gemini, and qwen3:8b on Ollama.
  • ToolCall, Usage, ChatResult: plain dataclasses for what comes back. Usage records tokens (the units models read and write, roughly word pieces; Module 2 covers them in depth) split into input, output, cached, and reasoning tokens, because cost is billed per token. ChatResult.as_message() converts a reply back into a message you can append to the conversation.
  • resolve(provider, model): picks the provider and model from function arguments, then the LLM_PROVIDER and LLM_MODEL environment variables, then the defaults. Unknown providers fail loudly.
  • make_client(provider): builds an openai.OpenAI client pointed at the provider's URL, with the key read from the environment (never from code). It raises a clear error if the key is missing. max_retries=2 retries transient network errors.
  • _usage and _parse_arguments: small private helpers. _usage copes with providers that omit usage details. _parse_arguments never raises on malformed tool-call JSON; it wraps the raw text so the caller can decide what to do (Module 6 builds on this).
  • chat(messages, ...): sends one request. Parameters left as None are not sent, so each provider keeps its own default. It times the call, extracts text, tool calls, usage, and finish reason, and returns a ChatResult. temperature defaults to 0.0 (the least random setting; Module 3 explains why that still is not fully deterministic).
  • StreamStats and stream_chat(...): the streaming version, which yields text as it arrives and records time to first token. Module 3 uses it.
  • Comes out: nothing when imported. Calling chat needs a key or a running Ollama server.

supportdesk/stand_in.py

python
"""A scripted stand-in for `llm.chat`, for running examples and tests without an API key.

It has the same call signature as `supportdesk.llm.chat` and returns the same
ChatResult type. It is NOT a language model: it replays scripted replies or
applies a rule you give it. Use it to test plumbing (parsing, loops, retries,
budgets), never to measure model quality.
"""
from __future__ import annotations

import itertools
from collections.abc import Callable
from typing import Any

from supportdesk.llm import ChatResult, Usage
from supportdesk.tokens import count_messages, count_tokens

Responder = Callable[[list[dict[str, Any]], dict[str, Any]], ChatResult | str]


class ScriptedLLM:
    """Replays a list of replies in order, or calls `responder(messages, kwargs)`."""

    _ids = itertools.count(1)

    def __init__(self, replies: list[ChatResult | str] | None = None, responder: Responder | None = None,
                 model: str = "scripted-stand-in") -> None:
        if (replies is None) == (responder is None):
            raise ValueError("Pass exactly one of `replies` or `responder`.")
        self.replies = list(replies or [])
        self.responder = responder
        self.model = model
        self.calls: list[dict[str, Any]] = []

    def __call__(self, messages: list[dict[str, Any]], **kwargs: Any) -> ChatResult:
        self.calls.append({"messages": [dict(m) for m in messages], **kwargs})
        if self.responder is not None:
            reply = self.responder(messages, kwargs)
        elif self.replies:
            reply = self.replies.pop(0)
        else:
            raise RuntimeError("ScriptedLLM ran out of scripted replies.")
        result = ChatResult(text=reply) if isinstance(reply, str) else reply
        if result.usage.input_tokens == 0:
            result.usage = Usage(input_tokens=count_messages(messages), output_tokens=count_tokens(result.text or ""))
        result.model = result.model or self.model
        result.provider = result.provider or "stand-in"
        result.finish_reason = result.finish_reason or ("tool_calls" if result.tool_calls else "stop")
        return result

Code explained

  • In simple words: ScriptedLLM is a crash-test dummy. It has the same shape as a model call, so you can test everything around the model without paying for, or waiting on, a real one. It has no intelligence.
  • What happens:
  • ScriptedLLM(replies=...) replays a list of replies in order. ScriptedLLM(responder=...) instead calls a function you write, which receives the messages and keyword arguments and returns text or a ChatResult. Passing both, or neither, is an error.
  • __call__ makes the object callable with the same signature as llm.chat, so code written as chat(messages, temperature=0) works with either. It records every call in self.calls (handy for tests), fills in estimated token usage with the offline tokenizer from tokens.py when you did not script it, and sets provider="stand-in" so logs never confuse it with a real model.
  • When the scripted replies run out, it raises instead of inventing text.
  • Comes out: a ChatResult, exactly like llm.chat.

Your first call: a triage question

Now use the helpers together. The script below asks a model to triage one real ticket. If a key is configured, it calls the real provider through llm.chat. If not, it falls back to ScriptedLLM so the rest of the code still runs.

python
"""Module 1: the first call to a hosted model, and the same call through the stand-in.

With a key set (for example GROQ_API_KEY, or LLM_PROVIDER=ollama with Ollama running),
this sends a real request through supportdesk.llm.chat. Without one, it falls back to
ScriptedLLM, which is NOT a model: it replays a scripted reply so the plumbing runs.
"""
import os

from supportdesk import llm
from supportdesk.data import CATEGORIES, load_tickets
from supportdesk.stand_in import ScriptedLLM

ticket = next(t for t in load_tickets() if t.id == "T-1008")
messages = [
    {"role": "system", "content": "You triage support tickets for Brightlane, a project-management SaaS. "
                                  f"Reply with one category from: {', '.join(CATEGORIES)}. "
                                  "Then one short sentence explaining why."},
    {"role": "user", "content": ticket.text},
]

provider, model = llm.resolve()
key_env = llm.PROVIDERS[provider]["key_env"]
use_real = provider == "ollama" or bool(os.environ.get(key_env))

if use_real:
    chat = llm.chat
    print(f"calling {provider} / {model}")
else:
    chat = ScriptedLLM(replies=["account_access\nThe customer is locked out after failed password attempts."])
    print(f"no {key_env} set: using ScriptedLLM (scripted reply, not model output)")

result = chat(messages, temperature=0.0, max_tokens=200)
print("text:         ", result.text.replace("\n", " | "))
print("provider/model:", result.provider, "/", result.model)
print("finish_reason: ", result.finish_reason)
print("usage:         ", result.usage)
print("latency_ms:    ", result.latency_ms)
print("gold category: ", ticket.gold["category"])

Code explained

  • In simple words: build a two-message conversation (instructions plus the ticket), send it to whichever "model" is available, and print everything the result object carries.
  • What happens: the system message sets the job and the allowed categories (taken from CATEGORIES, not typed by hand). The user message is the ticket text for T-1008, a customer who is locked out. llm.resolve() reports which provider and model would be used. The script checks the provider's key variable; Ollama needs no key. Both paths call chat(messages, temperature=0.0, max_tokens=200) with identical arguments, which is the point of the shared signature.
  • Comes out: this build has no API key, so the output below is real plumbing output from ScriptedLLM. The text is the scripted reply, not a model's judgment. The usage numbers are real counts from the offline tokenizer (90 input tokens, 13 output tokens) and latency is 0 because nothing went over the network.
text
  no GROQ_API_KEY set: using ScriptedLLM (scripted reply, not model output)
  text:          account_access | The customer is locked out after failed password attempts.
  provider/model: stand-in / scripted-stand-in
  finish_reason:  stop
  usage:          Usage(input_tokens=90, output_tokens=13, cached_tokens=0, reasoning_tokens=0)
  latency_ms:     0.0
  gold category:  account_access

With a Groq key set (export GROQ_API_KEY=...), the same script calls openai/gpt-oss-120b. The following is an illustrative sample run (not captured in this build; produced for teaching). Your output will differ, including the wording, the token counts, and the latency:

text
calling groq / openai/gpt-oss-120b
text:          account_access | The customer is locked out after too many failed password attempts and needs access restored.
provider/model: groq / openai/gpt-oss-120b
finish_reason:  stop
usage:          Usage(input_tokens=131, output_tokens=74, cached_tokens=0, reasoning_tokens=48)
latency_ms:     640.2
gold category:  account_access

Code explained

  • In simple words: the shape a real reply takes, so you know what to look for when you run it yourself.
  • What happens: a real provider counts input tokens with its own tokenizer and chat formatting, so the input count differs from the stand-in's estimate. gpt-oss-120b is a reasoning model: it writes hidden "thinking" tokens before the answer, and they show up in reasoning_tokens and are billed as output tokens.
  • Comes out: labelled illustrative. Run the script with your key to get real numbers; you will use them in Module 2 to compute cost per ticket.

Two failures you will almost certainly hit, and how to read them:

What you seeWhat it meansFix
RuntimeError: Set GROQ_API_KEY in your environment to use groq.make_client found no key for the selected providerexport GROQ_API_KEY=..., or export LLM_PROVIDER=gemini with GEMINI_API_KEY, or export LLM_PROVIDER=ollama
openai.APIConnectionError: Connection error. with LLM_PROVIDER=ollamaNothing is listening on localhost:11434Start Ollama (ollama serve) and pull the model (ollama pull qwen3:8b), or set OLLAMA_BASE_URL

Both are real errors captured in this build. Read the last line of the traceback first: llm.py is written so that line names the missing piece.

Across the course you will choose between three ways of running an example:

SituationUse thisWhy
You want to know how well a model does a taskllm.chat with a real providerOnly a real model produces evidence about quality
You are testing loops, parsers, retries, budgets, or routingScriptedLLMDeterministic, free, instant; failures you see are in your code, not the model
You want to see a mechanism (probabilities, sampling, training) with real numbersTinyLMYou can open the model up; no hosted API exposes its internals this way