Part A: How a response is produced
Module 3: Inference Behaviour and Decoding Control
By the end of this module, you'll have:
- A measured picture of how a reply is produced: prefill and decode timings from a real model, with the KV cache on and off, and a clear answer to why long outputs, not long prompts, make replies slow.
- A token-by-token streaming loop with time to first token and inter-token latency, plus a fix for the classic "stop sequence leaks into the stream" bug.
- Working intuition for temperature, top-k, top-p, min-p, penalties, stop sequences, max tokens, and seeds, backed by entropy numbers, survivor counts, and a 20-samples-per-setting diversity experiment.
- A constrained decoder that forces TinyLM to emit one of the six Brightlane ticket categories or a small triage JSON object, and a "constraint pressure" measurement that tells you when the format is doing all the work.
- A cost and latency model for reasoning effort levels built on
pricing.py, a harness to measure them on real tickets, and a grounded view of how far to trust visible reasoning traces. - Decoding presets for the Brightlane assistant (triage, reply drafts, brainstorming) that translate into
llm.chatarguments for Groq, Gemini, or Ollama.
Prerequisites: Module 1 (next-token prediction, TinyLM at a high level, llm.chat) and Module 2 (tokens, counting, pricing.py, why output tokens cost more). Working Python.
Where we are: Module 2 treated a request as a bill: tokens in, tokens out, dollars. This module opens the box between "request sent" and "reply received": how the model turns one probability distribution per step into text, which knobs change that, and what each knob costs in time, money, and quality.
How this module is organized
| Part | What it covers |
|---|---|
| Part A: How a response is produced | The decoding loop, prefill vs decode, the KV cache, time to first token, inter-token latency, streaming, and why output length drives latency |
| Part B: Sampling parameters | Temperature, top-k, top-p, min-p, frequency and presence penalties, stop sequences, max tokens, seeds, why hosted models vary, and choosing settings by task |
| Part C: Constrained generation | Forcing a format token by token, logit bias, what constraints cost in probability and quality, and the production tools that do this |
| Part D: Reasoning-model behaviour | Thinking tokens, effort levels and budgets, cost and latency math, when reasoning pays, and reading traces without over-trusting them |
Setup: TinyLM as the lab instrument
Hosted models hide their probabilities and run on hardware you cannot see. TinyLM does not. It is the 1,071,872-parameter GPT from Module 1, trained on this project's support corpus, and it exposes every step: logits, probabilities, the sampling filters, prefill and decode timings. Its output quality says nothing about real LLMs. Its mechanisms are the same ones real LLMs use, so we measure those.
One property of TinyLM shapes every experiment here. It memorized its templated corpus (validation loss bottoms out near step 400 while training loss keeps falling, as Module 1 showed). So on most support prompts its next-token distribution is extremely peaked: after Customer (Ana): Hi, my it puts 99.9 percent on account, because every conversation in the corpus that starts that way continues that way. Sampling settings do almost nothing to a distribution like that. To see them work we need prompts where TinyLM is genuinely unsure, and there are a few: right after Customer (Ana): Hi, the corpus has about nine different questions a customer might ask, and after Agent ( there are sixteen agent names, all roughly equally likely. We will lean on those.
The examples run in order from the repository root with PYTHONPATH=.. The first file loads TinyLM once and defines the small helpers every later example imports.
examples/m03_setup.py (click to expand)
pythonCopy
"""Module 3 shared setup: load TinyLM once and define small measurement helpers.
Every other m03 example imports from here, so run them from the repo root with
PYTHONPATH=. (for example: PYTHONPATH=. python examples/m03_temperature.py).
"""
import math
import statistics
import torch
from supportdesk.tinylm import SamplingParams, apply_sampling, load
torch.set_num_threads(1) # one thread: steadier timings for a model this small (see Part A)
torch.manual_seed(0)
MODEL, TOK = load() # the pretrained TinyLM from models/tinylm-base
# Prompts we reuse. TinyLM memorized its templated corpus, so most support
# prompts have one obvious continuation. These two are the exceptions we want.
UNCERTAIN = "Customer (Ana): Hi," # which question comes next is open
CERTAIN = "Customer (Ana): Hi, my" # the corpus always continues with " account"
@torch.no_grad()
def logits_for(text: str) -> torch.Tensor:
"""Raw next-token scores (one per vocabulary entry) after `text`."""
ids = TOK.encode(text).ids[-MODEL.cfg.context:]
return MODEL(torch.tensor([ids]))[0, -1]
def distribution(text: str, params: SamplingParams) -> torch.Tensor:
"""The probabilities the sampler would draw from, after all sampling settings."""
return apply_sampling(logits_for(text), params, generated=[])
def entropy_bits(probs: torch.Tensor) -> float:
"""How spread out a distribution is, in bits. 0 = certain; log2(2048) = 11 = uniform."""
p = probs[probs > 0]
return float(-(p * p.log2()).sum()) + 0.0 # "+ 0.0" turns -0.0 into 0.0
def top_tokens(probs: torch.Tensor, n: int = 5) -> list[tuple[str, float]]:
"""The n most likely tokens as (text, probability) pairs."""
values, ids = probs.topk(n)
return [(TOK.decode([i]), round(v, 3)) for v, i in zip(values.tolist(), ids.tolist())]
def survivors(probs: torch.Tensor) -> int:
"""How many tokens still have nonzero probability after filtering."""
return int((probs > 0).sum())
def median_and_spread(values: list[float]) -> str:
"""'median (min to max)' so every timing shows its noise."""
return f"{statistics.median(values):.2f} ({min(values):.2f} to {max(values):.2f})"
@torch.no_grad()
def token_probs(prompt: str, continuation_ids: list[int]) -> list[float]:
"""The raw model's probability (T=1, no filters) for each token of a continuation."""
ids = TOK.encode(prompt).ids
probs = torch.softmax(MODEL(torch.tensor([ids + continuation_ids]))[0], dim=-1)
return [probs[len(ids) - 1 + i, t].item() for i, t in enumerate(continuation_ids)]
if __name__ == "__main__":
print(f"TinyLM loaded: {MODEL.num_parameters():,} parameters, context {MODEL.cfg.context} tokens")
for prompt in (UNCERTAIN, CERTAIN):
probs = distribution(prompt, SamplingParams(temperature=1.0))
print(f"{prompt!r:28} entropy {entropy_bits(probs):.2f} bits, top: {top_tokens(probs, 3)}")
print("max possible entropy:", round(math.log2(MODEL.cfg.vocab_size), 2), "bits")
Code explained
- In simple words: this is the lab bench: it loads the model, picks two reference prompts (one where TinyLM is sure, one where it is not), and defines the measuring tools we use all module.
- What happens:
load()readsmodels/tinylm-base.logits_forruns the model on a prompt and keeps the scores for the next position only.distributionpasses those scores through the realapply_samplingfromsupportdesk/tinylm.py, so every number you see is what the sampler would actually draw from.entropy_bitsmeasures spread: 0 bits means one certain token, 11 bits means all 2,048 tokens equally likely.survivorscounts tokens a filter left alive.median_and_spreadprints every timing with its range, andtoken_probsscores a continuation under the unmodified model, which we use later as a quality check. Thetorch.set_num_threads(1)line is explained in Part A: on a busy machine, two threads made this tiny model up to 100 times slower. - Comes out: run
PYTHONPATH=. python examples/m03_setup.py:
TinyLM loaded: 1,071,872 parameters, context 128 tokens
'Customer (Ana): Hi,' entropy 3.06 bits, top: [(' how', 0.27), (' I', 0.153), (' can', 0.108)]
'Customer (Ana): Hi, my' entropy 0.02 bits, top: [(' account', 0.999), (' password', 0.0), (' annual', 0.0)]
max possible entropy: 11.0 bits
The uncertain prompt has 3.06 bits of entropy (roughly "eight or nine live options"), the certain one 0.02 bits. Keep both numbers in mind: they explain why the same temperature changes one prompt's output a lot and the other's not at all.
Part A: How a response is produced
The decoding loop
A language model produces one token at a time. For each step it outputs logits: one raw score per vocabulary entry (2,048 for TinyLM, around 200,000 for recent hosted models). A sampler turns those scores into probabilities, applies your settings, and picks one token. That token is appended to the input, and the model runs again. The whole loop is called decoding. Here is generate from supportdesk/tinylm.py, shown as an excerpt (Module 1 showed the full file):
Excerpt from supportdesk/tinylm.py:
@torch.no_grad()
def generate(model: TinyGPT, tokenizer: Tokenizer, prompt: str, params: SamplingParams | None = None,
use_cache: bool = True, allowed=None) -> Generation:
"""Generate a continuation of `prompt`.
`allowed`, if given, is a function (generated_ids) -> set of token ids that
may come next; everything else is masked out (constrained decoding).
"""
params = params or SamplingParams()
generator = torch.Generator().manual_seed(params.seed) if params.seed is not None else None
ids = tokenizer.encode(prompt).ids[-(model.cfg.context - params.max_new_tokens):]
caches = [dict() for _ in model.blocks] if use_cache else None
generated: list[int] = []
started = time.perf_counter()
logits = model(torch.tensor([ids]), caches)[0, -1]
prefill_ms = (time.perf_counter() - started) * 1000
decode_started = time.perf_counter()
stop_reason = "length"
for _ in range(params.max_new_tokens):
probs = apply_sampling(logits, params, generated)
if allowed is not None:
mask = torch.zeros_like(probs)
permitted = list(allowed(generated))
if not permitted:
stop_reason = "constraint"
break
mask[permitted] = 1.0
probs = probs * mask
probs = probs / probs.sum() if probs.sum() > 0 else mask / mask.sum()
next_id = int(torch.multinomial(probs, 1, generator=generator)) if params.temperature > 0 else int(probs.argmax())
generated.append(next_id)
text = tokenizer.decode(generated)
if any(s in text for s in params.stop):
stop_reason = "stop"
break
position = len(ids) + len(generated) - 1
if position >= model.cfg.context:
break
if use_cache:
logits = model(torch.tensor([[next_id]]), caches, start_pos=position)[0, -1]
else:
logits = model(torch.tensor([ids + generated]))[0, -1]
decode_ms = (time.perf_counter() - decode_started) * 1000
text = tokenizer.decode(generated)
for s in params.stop:
if s in text:
text = text[: text.index(s)]
return Generation(text, generated, round(prefill_ms, 2), round(decode_ms / max(len(generated), 1), 3), stop_reason)
Code explained
- In simple words: read the whole prompt once, then loop: turn scores into probabilities, pick a token, check whether to stop, and run the model on just the new token.
- What happens: the prompt is tokenized and trimmed so prompt plus output fits the 128-token context. The first
model(...)call processes every prompt token at once; that is prefill, timed asprefill_ms. Each loop iteration is one decode step:apply_samplingreshapes the probabilities (Part B), the optionalallowedmask removes forbidden tokens (Part C), then eithertorch.multinomialdraws a random token or, at temperature 0,argmaxtakes the top one. After each token it decodes the text so far and checks stop sequences. Withuse_cache=Truethe next model call sees only the new token and reuses stored attention keys and values (the KV cache, below); withuse_cache=Falseit reprocesses the whole sequence every step. At the end it trims the stop sequence from the returned text and reports why it stopped:"stop","length", or"constraint". - Comes out: nothing yet; this is the function the rest of the module measures. Two details matter later. First,
token_idsstill contains the stop-sequence tokens even thoughtextdoes not, so you paid for them. Second, whentemperature <= 0and anallowedmask is given, the code falls back in a way that surprises people; Part C diagnoses it.
Let's run one step by hand, then the whole loop.
examples/m03_one_step.py:
"""One decoding step by hand, then the same thing through generate()."""
import torch
from examples.m03_setup import MODEL, TOK, UNCERTAIN, logits_for, top_tokens
from supportdesk.tinylm import SamplingParams, apply_sampling, generate
logits = logits_for(UNCERTAIN) # 2048 raw scores, one per token
print("logits shape:", tuple(logits.shape), " highest score:", round(logits.max().item(), 2))
probs = apply_sampling(logits, SamplingParams(temperature=1.0), generated=[])
print("probabilities sum to", round(probs.sum().item(), 4), " top 5:", top_tokens(probs))
greedy_id = int(probs.argmax()) # temperature 0 picks this every time
g = torch.Generator().manual_seed(1)
sampled = [TOK.decode([int(torch.multinomial(probs, 1, generator=g))]) for _ in range(8)]
print("greedy pick:", repr(TOK.decode([greedy_id])), " 8 sampled picks:", sampled)
# generate() repeats that step: sample, append, run the model on the new token, repeat.
out = generate(MODEL, TOK, UNCERTAIN, SamplingParams(max_new_tokens=30, temperature=0, stop=["\n"]))
print("greedy continuation:", repr(out.text), "| stop_reason:", out.stop_reason, "| tokens:", len(out.token_ids))
Code explained
- In simple words: we look at the model's scores for one position, turn them into probabilities, and pick a token two ways: always the top one (greedy) or by rolling weighted dice (sampling).
- What happens:
logits_forgives 2,048 scores.apply_samplingat temperature 1.0 applies a softmax, which turns scores into positive numbers that sum to 1. Greedy decoding takes theargmax. Sampling draws withtorch.multinomial, where a token with probability 0.27 wins about 27 percent of the time. Finallygenerateruns the full loop greedily until the newline stop sequence. - Comes out:
logits shape: (2048,) highest score: 10.23
probabilities sum to 1.0 top 5: [(' how', 0.27), (' I', 0.153), (' can', 0.108), (' where', 0.091), (' slack', 0.084)]
greedy pick: ' how' 8 sampled picks: [' does', ' does', ' does', ' the', ' how', ' my', ' can', ' where']
greedy continuation: ' how do I export my invoice data?' | stop_reason: stop | tokens: 9
Greedy always says how. Eight samples produced six different first words, including does three times even though it is not in the top five: that is what "the tail" means, and with only eight draws, three hits on a lower-probability token is ordinary luck. The greedy continuation is a memorized customer question.
Prefill and decode are separate phases
The two phases have very different costs. Prefill processes all prompt tokens in one parallel pass: the hardware multiplies big matrices once. Decode produces output tokens one after another, and each step needs a full pass through the model for a single token. Parallel work is cheap per token; sequential work is not. On GPUs the effect is larger still: decode is limited by how fast weights can be read from memory, not by arithmetic, so each step costs roughly the same no matter how little work it does.
Prompt: N tokens → →→ Prefill: one parallel pass over all N tokens → → → (KV cache: keys and values for every prompt token) → → → First output token → → → Decode step: one token in, one token out → → →
Next output token
Let's measure both phases on TinyLM, varying prompt length and output length independently.
examples/m03_prefill_decode.py:
"""Measure prefill time and per-token decode time for different prompt and output lengths."""
import statistics
from pathlib import Path
from examples.m03_setup import MODEL, TOK, median_and_spread
from supportdesk.tinylm import SamplingParams, generate
corpus_ids = TOK.encode(Path("data/corpus.txt").read_text(encoding="utf-8")[:4000]).ids
REPEATS = 15
def prompt_of(n_tokens: int) -> str:
"""A slice of real corpus text that encodes to exactly n_tokens tokens."""
text = TOK.decode(corpus_ids[:n_tokens])
assert len(TOK.encode(text).ids) == n_tokens, "re-encoding changed the length"
return text
def measure(prompt: str, new_tokens: int) -> dict:
params = SamplingParams(max_new_tokens=new_tokens, temperature=0) # greedy, no stop: always runs to length
runs = [generate(MODEL, TOK, prompt, params) for _ in range(REPEATS)]
assert all(len(r.token_ids) == new_tokens for r in runs)
total = [r.prefill_ms + r.decode_ms_per_token * new_tokens for r in runs]
return {
"prefill": [r.prefill_ms for r in runs],
"per_token": [r.decode_ms_per_token for r in runs],
"total": total,
}
generate(MODEL, TOK, prompt_of(16), SamplingParams(max_new_tokens=8, temperature=0)) # warm-up, not timed
print(f"Each cell: median (min to max) over {REPEATS} runs, in milliseconds\n")
print(f"{'prompt':>6} {'output':>6} | {'prefill ms':>20} | {'decode ms/token':>20} | {'total ms':>22}")
rows = {}
for p_len in (8, 32, 64, 96):
for n_out in (8, 32):
m = measure(prompt_of(p_len), n_out)
rows[(p_len, n_out)] = m
print(f"{p_len:>6} {n_out:>6} | {median_and_spread(m['prefill']):>20} | "
f"{median_and_spread(m['per_token']):>20} | {median_and_spread(m['total']):>22}")
# Minimums are the cleanest estimate of the true cost: noise from other processes only ever adds time.
best = {k: {name: min(vals) for name, vals in v.items()} for k, v in rows.items()}
prefill_best = {p: min(best[(p, 8)]["prefill"], best[(p, 32)]["prefill"]) for p in (8, 32, 64, 96)}
prompt_cost = (prefill_best[96] - prefill_best[8]) / 88
output_cost = statistics.median(statistics.median(v["per_token"]) for v in rows.values()) # already an average of many tokens
print("\nbest prefill by prompt length:", {p: round(v, 2) for p, v in prefill_best.items()})
print(f"prefill, 8 to 96 prompt tokens: {prefill_best[8]:.2f} to {prefill_best[96]:.2f} ms,"
f" so about {prompt_cost:.3f} ms per extra prompt token")
print(f"decode (typical): about {output_cost:.2f} ms per output token, "
f"{output_cost / prompt_cost:.0f} times the cost of a prompt token")
Code explained
- In simple words: time a stopwatch around "read the prompt" and "write the answer" separately, for short and long prompts and short and long answers, and repeat enough times to see the noise.
- What happens:
prompt_ofcuts real corpus text to exactly 8, 32, 64, or 96 tokens (it asserts that re-encoding gives the same length, so the lengths in the table are true). Each combination runs 15 times greedily with no stop sequence, so every run produces exactly the requested number of tokens. Each cell prints the median and the full range. The summary uses minimums for prefill, because background noise only ever adds time, so the fastest run is the cleanest estimate of the true cost, and the median of the per-token decode times, which are already averages over many tokens. - Comes out: on the 2-core course machine, when it was otherwise idle:
Each cell: median (min to max) over 15 runs, in milliseconds
prompt output | prefill ms | decode ms/token | total ms
8 8 | 1.19 (1.11 to 1.54) | 0.88 (0.85 to 0.93) | 8.18 (8.00 to 9.00)
8 32 | 1.24 (1.12 to 2.22) | 0.91 (0.85 to 1.06) | 30.44 (28.47 to 35.57)
32 8 | 1.73 (1.65 to 2.37) | 0.89 (0.84 to 1.26) | 9.01 (8.41 to 12.15)
32 32 | 1.75 (1.70 to 1.88) | 0.87 (0.84 to 0.97) | 29.65 (28.58 to 32.67)
64 8 | 2.43 (2.37 to 2.87) | 0.88 (0.85 to 0.91) | 9.47 (9.23 to 9.85)
64 32 | 2.55 (2.42 to 2.91) | 0.93 (0.86 to 1.01) | 32.21 (30.07 to 34.91)
96 8 | 3.44 (3.23 to 4.43) | 1.02 (0.90 to 1.08) | 11.61 (10.40 to 12.74)
96 32 | 3.50 (3.32 to 3.86) | 0.97 (0.92 to 1.09) | 34.54 (32.79 to 38.73)
best prefill by prompt length: {8: 1.11, 32: 1.65, 64: 2.37, 96: 3.23}
prefill, 8 to 96 prompt tokens: 1.11 to 3.23 ms, so about 0.024 ms per extra prompt token
decode (typical): about 0.90 ms per output token, 37 times the cost of a prompt token
Read the table two ways. Down a column, prompt length quadruples-plus from 8 to 96 tokens and prefill grows from about 1.1 to 3.2 ms, while decode stays near 0.9 ms per token. Across a row, going from 8 to 32 output tokens adds about 21 ms regardless of prompt length. One more output token costs roughly 37 times as much wall-clock time as one more prompt token here. Your absolute numbers will differ; the ratio is the lesson.
Noise, honestly. The first time this ran, other heavy jobs were sharing the machine (load average above 12 on 2 cores). The same script then reported median decode times of 2.4 to 3.2 ms per token and single prefill outliers above 20 ms, three times slower overall and with ranges wider than some of the differences we care about. The conclusion (output length dominates) held in both runs, but you could not have read a 10 percent difference out of the noisy one. Two habits follow. Always print the spread, not just one number. And when two numbers differ by less than their ranges overlap, call it noise.
A side finding from that noisy run: with torch.set_num_threads(2) on the overloaded machine, a 20-token generation took about 15 seconds instead of about 0.13 seconds with one thread, because the two worker threads kept waiting on each other while other processes held the cores. On the idle machine one and two threads were the same (18.3 vs 17.7 ms per generation). That is why the setup file uses one thread. It is a TinyLM-on-a-shared-CPU quirk, not a property of LLMs, but "check what else is running before trusting a benchmark" is universal.
The KV cache: same tokens, less work
Attention lets each new token look back at every earlier token through two vectors per token per layer, called keys and values. Those vectors for old tokens never change. The KV cache stores them so each decode step only computes keys and values for the one new token. Without it, step 50 recomputes all 49 earlier positions again.
examples/m03_kv_cache.py:
"""KV cache on vs off: identical tokens, very different decode cost."""
import statistics
from examples.m03_setup import MODEL, TOK
from supportdesk.tinylm import SamplingParams, generate
PROMPT = "Customer (Ana): Hi, I was charged twice for the Team plan this month. Can you refund the duplicate?\n"
REPEATS = 7
print(f"prompt tokens: {len(TOK.encode(PROMPT).ids)}; each cell is the median (min) of {REPEATS} runs\n")
print(f"{'new tokens':>10} | {'cache ms/token':>15} | {'no-cache ms/token':>17} | {'slowdown':>8} | same tokens?")
for n_new in (8, 32, 64, 96):
params = SamplingParams(max_new_tokens=n_new, temperature=0)
cached, uncached = [], []
for _ in range(REPEATS): # interleave, so background load hits both equally
cached.append(generate(MODEL, TOK, PROMPT, params, use_cache=True))
uncached.append(generate(MODEL, TOK, PROMPT, params, use_cache=False))
c = [g.decode_ms_per_token for g in cached]
u = [g.decode_ms_per_token for g in uncached]
same = all(a.token_ids == b.token_ids for a, b in zip(cached, uncached))
print(f"{n_new:>10} | {statistics.median(c):>7.2f} ({min(c):5.2f}) | {statistics.median(u):>9.2f} ({min(u):5.2f}) |"
f" {statistics.median(u) / statistics.median(c):>7.1f}x | {same}")
print("\ncontinuation (first 90 characters):", repr(cached[0].text[:90]))
Code explained
- In simple words: generate the same greedy reply with and without the cache and compare speed and output.
- What happens: for 8, 32, 64, and 96 new tokens, the script alternates cached and uncached runs seven times each (interleaving means any background slowdown hits both equally), records per-token decode time, and checks every pair produced identical token ids.
- Comes out:
Time to first token, inter-token latency, and streaming
Two latency numbers describe what a user experiences:
- Time to first token (TTFT): from sending the request to seeing the first character. It covers network, queueing on the provider's side, and prefill. For reasoning models it also covers any hidden thinking.
- Inter-token latency (ITL): the gap between later tokens, set by decode speed. Its inverse is the familiar "tokens per second".
Total time is roughly TTFT plus ITL times the number of output tokens. Streaming sends each piece of text as soon as it is decoded instead of waiting for the whole reply. It does not make the model faster at all. It makes the product feel faster, because a reader starts reading after TTFT instead of after the total. For a support agent reviewing a 300-token draft at 50 tokens per second, that is the difference between staring at a spinner for 6 seconds and seeing words after half a second.
Here is streaming from TinyLM, with a timestamp on every piece:
examples/m03_stream_sim.py:
"""Stream TinyLM token by token with timestamps: time to first token vs inter-token latency."""
import statistics
import time
import torch
from examples.m03_setup import MODEL, TOK
from supportdesk.tinylm import SamplingParams, apply_sampling
@torch.no_grad()
def stream(prompt: str, params: SamplingParams):
"""Yield (new_text, ms_since_request) as each token is decoded, like a streaming API."""
started = time.perf_counter()
ids = TOK.encode(prompt).ids[-(MODEL.cfg.context - params.max_new_tokens):]
caches = [dict() for _ in MODEL.blocks]
logits = MODEL(torch.tensor([ids]), caches)[0, -1] # prefill: the whole prompt at once
generated, shown = [], ""
for _ in range(params.max_new_tokens):
next_id = int(apply_sampling(logits, params, generated).argmax()) # greedy for a repeatable demo
generated.append(next_id)
text = TOK.decode(generated)
piece, shown = text[len(shown):], text # send only the new characters
yield piece, (time.perf_counter() - started) * 1000
if any(s in text for s in params.stop):
return
logits = MODEL(torch.tensor([[next_id]]), caches, start_pos=len(ids) + len(generated) - 1)[0, -1]
PROMPT = "Customer (Ana): Hi, how do I export my board data?\nAgent (Lena):"
events = list(stream(PROMPT, SamplingParams(max_new_tokens=40, temperature=0, stop=["\nCustomer"])))
for piece, t in events[:6]:
print(f"{t:7.2f} ms {piece!r}")
print(" ...")
for piece, t in events[-3:]:
print(f"{t:7.2f} ms {piece!r}")
times = [t for _, t in events]
gaps = [b - a for a, b in zip(times, times[1:])]
print(f"\ntokens: {len(events)} time to first token: {times[0]:.2f} ms total: {times[-1]:.2f} ms")
print(f"inter-token latency: median {statistics.median(gaps):.2f} ms, "
f"p90 {sorted(gaps)[int(0.9 * len(gaps))]:.2f} ms, max {max(gaps):.2f} ms")
print(f"without streaming the reader sees nothing for {times[-1]:.1f} ms; "
f"with streaming, text appears after {times[0]:.1f} ms")
Code explained
- In simple words: the same loop as
generate, but instead of returning at the end it hands each new piece of text to the caller the moment it exists, stamped with the time. - What happens:
streamruns prefill, then each decode step picks a token, decodes the whole generated sequence, and yields only the characters that are new since last time. Decoding the whole sequence rather than one token at a time matters with byte-level tokenizers: a single token can be half of a multi-byte character, and decoding it alone would print garbage. The script then computes TTFT, the gaps between pieces (median, 90th percentile, max), and the total. - Comes out:
4.84 ms ' Export'
6.39 ms ' a'
7.54 ms ' board'
8.44 ms ' to'
9.45 ms ' CSV'
10.43 ms ' or'
...
26.76 ms '.'
27.67 ms '\n'
28.64 ms 'Customer'
tokens: 25 time to first token: 4.84 ms total: 28.64 ms
inter-token latency: median 0.95 ms, p90 1.08 ms, max 1.55 ms
without streaming the reader sees nothing for 28.6 ms; with streaming, text appears after 4.8 ms
TTFT is about 5 ms (prefill plus one decode step), the median inter-token gap about 1 ms, total about 29 ms. On a hosted model multiply everything by a few hundred, but the shape is the same.
Now look at the last two pieces: '\n' and 'Customer'. We asked to stop at "\nCustomer", and the loop did stop, but only after it had already streamed the stop sequence to the reader. generate trims the stop text from its return value because it has the whole text at the end; a streaming loop cannot un-send characters. This is a real bug class in streaming UIs. The fix, used in the Module Lab, is to hold back any trailing text that could be the start of a stop sequence until the next token proves otherwise. Hosted APIs do this for you on the server; your own streaming layers (for example, filtering or post-processing streams) must do it themselves.
On a hosted model, supportdesk.llm.stream_chat does the same job and fills a StreamStats object with TTFT and total time:
examples/m03_stream_hosted.py:
"""Stream a real hosted model through llm.stream_chat and measure TTFT and inter-token pace."""
import os
import sys
import time
from supportdesk.llm import PROVIDERS, StreamStats, resolve, stream_chat
from supportdesk.tokens import count_tokens
provider, model = resolve()
key_env = PROVIDERS[provider]["key_env"]
if key_env and not os.environ.get(key_env):
sys.exit(f"Set {key_env} (or LLM_PROVIDER=ollama) to run this example against {provider}.")
messages = [
{"role": "system", "content": "You are a support agent for Brightlane, a project-management app. Be concise."},
{"role": "user", "content": "How do I export my board data, and does the CSV include comments?"},
]
stats = StreamStats()
arrivals = []
started = time.perf_counter()
text = ""
for piece in stream_chat(messages, stats, temperature=None, max_tokens=400): # None = provider default
arrivals.append((time.perf_counter() - started) * 1000)
text += piece
print(piece, end="", flush=True)
out_tokens = count_tokens(text) # an o200k estimate; the provider's own count may differ a little
gaps = [b - a for a, b in zip(arrivals, arrivals[1:])]
print(f"\n\n{provider}/{model}: TTFT {stats.ttft_ms:.0f} ms, total {stats.total_ms:.0f} ms, "
f"{stats.chunks} chunks, about {out_tokens} tokens")
if gaps:
print(f"median gap between chunks {sorted(gaps)[len(gaps) // 2]:.1f} ms; "
f"decode pace about {out_tokens / max((stats.total_ms - stats.ttft_ms) / 1000, 1e-9):.0f} tokens/s")
Code explained
- In simple words: ask a real model a Brightlane question with streaming on, print text as it arrives, and report how long the first piece took and how fast the rest came.
- What happens:
resolve()picks the provider fromLLM_PROVIDER(Groq by default). The script exits with a clear message if the key is missing.stream_chatyields text pieces; we record an arrival time for each.temperature=Nonemeans "send no temperature, use the provider's default", which matters for Gemini 3 (Part B). Output tokens are estimated with the o200k tokenizer from Module 2, so the tokens-per-second figure is approximate. - Comes out: without a key in this build it prints
Set GROQ_API_KEY (or LLM_PROVIDER=ollama) to run this example against groq.With a key, the shape of the result looks like this.Illustrative sample run (not captured in this build; produced for teaching). Your output will differ.
Go to the board menu and choose Export, then pick CSV or JSON. CSV exports do not include comments or attachments; ...
groq/openai/gpt-oss-120b: TTFT 620 ms, total 1150 ms, 58 chunks, about 190 tokens
median gap between chunks 7.9 ms; decode pace about 360 tokens/s
Two things to check in your own run. First, openai/gpt-oss-120b is a reasoning model, and stream_chat only yields visible content, so the TTFT you see is "time to first visible token": it includes all the hidden reasoning. Raise reasoning_effort and watch TTFT grow while the inter-token gap stays the same. Second, chunks are not tokens: providers often batch several tokens per chunk, so count tokens from text, not chunks.
Why latency scales with output length, not input length
Put the measurements together. Prefill is parallel and cheap per token; decode is sequential and pays a full model pass per token. So for typical support traffic (a prompt of a few hundred to a few thousand tokens, a reply of a hundred or more), the reply length sets the wall-clock time. Very long prompts (tens of thousands of tokens) do push TTFT up noticeably, and providers with prompt caching can shave that, but the everyday lever is output length.
| Situation | Use this | Why |
|---|---|---|
| Replies feel slow in the agent UI | Stream the draft | Perceived wait drops from total time to TTFT; the model is no faster |
| Triage call returns a label only | Cap max_tokens low and ask for the label alone | Every output token is a sequential decode step and billed at the output rate |
| Long system prompt reused on every call | Prompt caching (Module 2) plus a stable prefix | Cached prefill is cheaper and faster; decode is unchanged |
| TTFT high on a reasoning model | Lower reasoning effort or use a non-reasoning model for simple tasks | Hidden thinking tokens are decoded before the first visible token |
| Need many labels at once | Batch API or concurrent requests | Throughput rises; per-request latency does not improve |