CourseLarge Language Models · Module 2 : Tokens -context-cost · part 11 of 80
Part 11 · Module 2 : Tokens -context-cost

Part D: Cost mechanics

25 min read·22 Sept 2026

The canonical price table: supportdesk/pricing.py

Every dollar figure in the course comes from one table, so a price change is a one-line edit.<strong>supportdesk/pricing.py</strong>

python
# supportdesk/pricing.py
"""Token prices per million tokens, used to turn usage into dollars.

Prices change often. These were checked on 21 September 2026 from the providers'
pricing pages (Groq figures via a third-party summary of Groq's page, so treat
them as directional). Verify current prices before relying on them, and edit
this table rather than hard-coding prices anywhere else.
"""
from __future__ import annotations

from dataclasses import dataclass

from supportdesk.llm import Usage


@dataclass(frozen=True)
class Price:
    input: float          # USD per 1M input tokens
    output: float         # USD per 1M output tokens (reasoning tokens bill as output)
    cached_input: float   # USD per 1M cached input tokens
    batch_discount: float = 0.5


PRICES: dict[str, Price] = {
    "openai/gpt-oss-120b": Price(input=0.15, output=0.60, cached_input=0.15),
    "openai/gpt-oss-20b": Price(input=0.075, output=0.30, cached_input=0.075),
    "llama-3.1-8b-instant": Price(input=0.05, output=0.08, cached_input=0.05),
    "llama-3.3-70b-versatile": Price(input=0.59, output=0.79, cached_input=0.59),
    "gemini-3.5-flash": Price(input=1.50, output=9.00, cached_input=0.15),
    "gemini-3.5-flash-lite": Price(input=0.30, output=2.50, cached_input=0.03),
    "qwen3:8b": Price(input=0.0, output=0.0, cached_input=0.0, batch_discount=0.0),
}


def cost_usd(usage: Usage, model: str, batch: bool = False) -> float:
    """Dollar cost of one call. Cached input tokens are billed at the cached rate."""
    price = PRICES[model]
    uncached = max(usage.input_tokens - usage.cached_tokens, 0)
    total = (
        uncached * price.input
        + usage.cached_tokens * price.cached_input
        + usage.output_tokens * price.output
    ) / 1_000_000
    return total * (1 - price.batch_discount) if batch else total

Code explained

  • In simple words: a price list per model and a function that turns a call's token usage into dollars.
  • What happens (per part):
  • Module docstring: prices were checked on 21 September 2026 and need verifying before you rely on them.
  • Price: a frozen dataclass with USD per 1 million tokens for input, output, and cached_input, plus batch_discount (0.5 means half price). The comment records that reasoning tokens bill as output.
  • PRICES: one entry per model id, keyed by the exact string llm.chat reports in ChatResult.model. Groq models, two Gemini models, and qwen3:8b at zero because Ollama runs on your own hardware (your electricity and hardware are not in the table).
  • cost_usd(usage, model, batch): splits input into uncached and cached tokens, multiplies each count by its price, adds output, divides by one million, and applies the batch discount to the whole bill if batch=True. Usage.output_tokens already includes reasoning tokens, because providers report reasoning inside the completion count, so they are priced as output without extra code.
  • Comes out: nothing by itself. Two notes from checking the table against provider pages on 21 Sep 2026. First, Groq's model pages list cached input for openai/gpt-oss-120b at 0.075 USD and for openai/gpt-oss-20b at 0.037 USD per million (a 50% discount), while the table has the full input price; Part D corrects this locally where it matters. Second, Groq states that its batch discount does not stack with caching (batch tokens bill at 50% regardless of cache status), so cost_usd(..., batch=True) slightly underestimates a Groq batch with cached tokens. Gemini 3.5 Flash's figures match Google's pricing page: 1.50 USD input, 9.00 USD output including thinking tokens, 0.15 USD cached input, batch at half price.

Input, output, and why output is the expensive half

python
# examples/m02_cost.py
"""Module 2: turn usage into dollars, and see where the money goes."""
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, cost_usd

# Realistic shapes for three Brightlane calls (token counts measured earlier in this module
# for triage; the reply and agent shapes are typical sizes, not measurements).
CALLS = {
    "triage (JSON label)": Usage(input_tokens=664, output_tokens=45),
    "draft reply": Usage(input_tokens=1_200, output_tokens=250),
    "draft reply, reasoning": Usage(input_tokens=1_200, output_tokens=1_050, reasoning_tokens=800),
}

print(f"{'model':26}" + "".join(f"{name:>26}" for name in CALLS))
for model in PRICES:
    print(f"{model:26}" + "".join(f"{cost_usd(u, model) * 1000:>21.4f} USD" for u in CALLS.values()))
print("(cost per 1,000 calls)\n")

model = "gemini-3.5-flash"
p = PRICES[model]
print(f"{model}: output costs {p.output / p.input:.0f}x input per token")
for name, u in CALLS.items():
    out_share = u.output_tokens * p.output / (u.input_tokens * p.input + u.output_tokens * p.output)
    print(f"  {name:24} {u.output_tokens / (u.input_tokens + u.output_tokens):4.0%} of tokens are output, "
          f"{out_share:4.0%} of the cost")

u = CALLS["draft reply, reasoning"]
visible = Usage(input_tokens=u.input_tokens, output_tokens=u.output_tokens - u.reasoning_tokens)
print(f"\nReasoning tokens bill as output: {cost_usd(u, model) / cost_usd(visible, model):.1f}x the cost of the "
      f"same reply without {u.reasoning_tokens} hidden reasoning tokens")
print(f"Batch tier for 10,000 draft replies on {model}: "
      f"{cost_usd(CALLS['draft reply'], model) * 10_000:.2f} USD live, "
      f"{cost_usd(CALLS['draft reply'], model, batch=True) * 10_000:.2f} USD batch")

Code explained

  • In simple words: price three realistic Brightlane calls on every model, then see which half of the bill is input and which is output.
  • What happens: the triage shape (664 in, 45 out) comes from Part A's measurement; the draft-reply shapes are typical sizes, not measurements. The reasoning variant has 800 hidden reasoning tokens inside its 1,050 output tokens, the same way providers report them. We print cost per 1,000 calls per model, then for gemini-3.5-flash the output share of tokens versus cost, the reasoning multiplier, and live versus batch pricing.
  • Comes out:
text
  model                            triage (JSON label)               draft reply    draft reply, reasoning
  openai/gpt-oss-120b                      0.1266 USD               0.3300 USD               0.8100 USD
  openai/gpt-oss-20b                       0.0633 USD               0.1650 USD               0.4050 USD
  llama-3.1-8b-instant                     0.0368 USD               0.0800 USD               0.1440 USD
  llama-3.3-70b-versatile                  0.4273 USD               0.9055 USD               1.5375 USD
  gemini-3.5-flash                         1.4010 USD               4.0500 USD              11.2500 USD
  gemini-3.5-flash-lite                    0.3117 USD               0.9850 USD               2.9850 USD
  qwen3:8b                                 0.0000 USD               0.0000 USD               0.0000 USD
  (cost per 1,000 calls)

  gemini-3.5-flash: output costs 6x input per token
    triage (JSON label)        6% of tokens are output,  29% of the cost
    draft reply               17% of tokens are output,  56% of the cost
    draft reply, reasoning    47% of tokens are output,  84% of the cost

  Reasoning tokens bill as output: 2.8x the cost of the same reply without 800 hidden reasoning tokens
  Batch tier for 10,000 draft replies on gemini-3.5-flash: 40.50 USD live, 20.25 USD batch

On gemini-3.5-flash an output token costs 6 times an input token. So a triage call whose tokens are only 6% output spends 29% of its money on output, and a draft reply that is 17% output spends 56% on it. Output is priced higher because it is produced one token at a time in the decode phase, each step a full pass through the model, while input is processed in parallel in the prefill phase (Module 3 measures both on TinyLM). On Groq the ratio is 4 for gpt-oss and 1.3 to 1.6 for the Llama models, so the effect is smaller but has the same direction.

Reasoning tokens are the hidden thinking a reasoning model generates before its visible answer. You never see them in ChatResult.text, but you pay for them as output, and they count against max_tokens. In this example 800 reasoning tokens make the reply 2.8 times as expensive as the same visible answer without them. llm.py exposes them as usage.reasoning_tokens when the provider reports them. Controlling reasoning effort is Module 3's topic; the cost lesson is to measure reasoning tokens per task before you choose a reasoning model for a high-volume job.

SituationUse thisWhy
Classification or extractionShort structured output, low or no reasoningOutput tokens are the expensive ones
Long draftsAsk for the length you need; set max_tokensEvery extra sentence is billed at the output rate
Reasoning model on a simple taskLow reasoning effort, or a non-reasoning modelHidden tokens can multiply the bill
Comparing modelsCompare cost per task with real usagePer-token prices hide output length and reasoning

Prompt caching and cache-aware layout

Prompt caching lets a provider reuse the work it did on a prompt's beginning when the next request starts with exactly the same tokens. Cached input tokens are billed at a discount and processed faster. Caching works on prefixes: a request hits the cache only for the tokens that are identical from the very first token up to the first difference. One changed character early in the prompt breaks the cache for everything after it. Provider details differ (all checked 21 Sep 2026):

ProviderHow it cachesDiscount on cached inputMinimum prefix
Groq (gpt-oss models)Automatic50%128 to 1,024 tokens depending on model; expires after 2 hours unused
Gemini (2.5 and newer)Implicit, automatic; explicit caches optional90% on 3.5 Flash (0.15 vs 1.50 USD), explicit caches also pay storage per hour4,096 tokens for 3.5 Flash
AnthropicYou mark breakpoints with cache_controlReads at 0.1x input; writes cost 1.25x (5-minute cache) or 2x (1-hour)512 to 4,096 tokens by model
OpenAIAutomaticVaries by model1,024 tokens, reused in 128-token blocks

So the layout rule is: stable content first, variable content last. Instructions, tool definitions, examples, and reference articles go at the top in a fixed order; the ticket, timestamps, and user-specific data go at the end. Let's measure how much of each triage request is shareable across the 24 test tickets under four layouts.

python
# examples/m02_cache_layout.py
"""Module 2: measure how much of each triage request a prefix cache could reuse, for three layouts."""
from dataclasses import replace

from m02_prompts import EXAMPLES, TRIAGE_INSTRUCTIONS, triage_messages

from supportdesk.data import load_articles, load_tickets
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.tokens import encode

tickets = load_tickets("test")  # 24 tickets, sent one after another
KB = "\n\n".join(f"[{a.id}] {a.title}\n{a.body}" for a in load_articles())


def render(messages: list[dict]) -> list[int]:
    """Approximate the provider's chat template: role header, content, separator."""
    return encode("".join(f"<|{m['role']}|>\n{m['content']}\n<|end|>\n" for m in messages))


def good(t):
    return triage_messages(t)


def good_with_kb(t):
    messages = triage_messages(t)
    messages[0] = {"role": "system", "content": TRIAGE_INSTRUCTIONS + "\n\nHelp-center articles:\n" + KB}
    return messages


def bad_metadata_first(t):
    header = f"Ticket {t.id}, received 2026-09-21, customer tier {t.customer_tier}.\n"
    messages = triage_messages(t)
    messages[0] = {"role": "system", "content": header + TRIAGE_INSTRUCTIONS}
    return messages


def bad_ticket_first(t):
    shots = "\n\n".join(f"Example ticket:\n{u}\nExample answer:\n{a}" for u, a in EXAMPLES)
    return [{"role": "user", "content": f"{t.text}\n\n---\n{TRIAGE_INSTRUCTIONS}\n\n{shots}"}]


def common_prefix(a: list[int], b: list[int]) -> int:
    n = 0
    for x, y in zip(a, b):
        if x != y:
            break
        n += 1
    return n


def openai_style(prefix: int, minimum: int = 1024, block: int = 128) -> int:
    """Cached tokens under a rule like OpenAI's: nothing below the minimum, then whole 128-token blocks."""
    return 0 if prefix < minimum else prefix // block * block


if __name__ == "__main__":
    print(f"{'layout':22}{'input tokens':>13}{'shared prefix':>15}{'share':>7}{'cacheable (1024 min, 128 blocks)':>34}")
    for layout in (good, good_with_kb, bad_metadata_first, bad_ticket_first):
        requests = [render(layout(t)) for t in tickets]
        total = sum(len(r) for r in requests)
        shared = [common_prefix(requests[i - 1], r) if i else 0 for i, r in enumerate(requests)]
        cacheable = sum(openai_style(s) for s in shared)
        print(f"{layout.__name__:22}{total:>13,}{sum(shared):>15,}{sum(shared) / total:>7.0%}{cacheable:>34,}")
        if layout is good:
            print(f"{'':22}per request: {len(requests[0])} tokens, of which {shared[1]} are identical to the previous request")

    # What caching is worth for the good_with_kb layout over the 24 tickets.
    # Groq's model page (checked 21 Sep 2026) lists cached input for gpt-oss-120b at 0.075 USD per 1M,
    # a 50% discount; the canonical table has 0.15. Correct it locally for this estimate.
    PRICES["openai/gpt-oss-120b"] = replace(PRICES["openai/gpt-oss-120b"], cached_input=0.075)

    requests = [render(good_with_kb(t)) for t in tickets]
    shared = [common_prefix(requests[i - 1], r) if i else 0 for i, r in enumerate(requests)]
    total_in, total_out = sum(len(r) for r in requests), 24 * 45  # about 45 output tokens per triage JSON
    print(f"\ngood_with_kb over 24 tickets: {total_in:,} input tokens, {total_out:,} output tokens")
    # Minimum cacheable prefix: Groq documents 128 to 1024 by model (1024 assumed); Gemini 3.5 Flash: 4,096.
    for model, minimum in [("openai/gpt-oss-120b", 1024), ("gemini-3.5-flash", 4096), ("gemini-3.5-flash-lite", 4096)]:
        cached = sum(s if s >= minimum else 0 for s in shared)
        plain = cost_usd(Usage(input_tokens=total_in, output_tokens=total_out), model)
        warm = cost_usd(Usage(input_tokens=total_in, output_tokens=total_out, cached_tokens=cached), model)
        print(f"  {model:22} minimum {minimum:>5}  cached {cached:>6,}  no cache {plain:.5f} USD  with cache {warm:.5f} USD  "
              f"({1 - warm / plain:.0%} saved)")

Code explained

  • In simple words: we line up 24 consecutive requests and count how many tokens at the start of each are identical to the one before, which is the most any prefix cache could reuse.
  • What happens: render() approximates the provider's chat template by joining role headers and contents, then tokenizes with o200k_base. The four layouts are: good (instructions and examples first, ticket last), good_with_kb (the same with all 12 help-center articles in the system prompt), bad_metadata_first (a ticket id and date line at the top of the system prompt, a very common habit), and bad_ticket_first (the ticket before the instructions). common_prefix() counts identical leading tokens between consecutive requests; openai_style() applies a 1,024-token minimum and 128-token blocks. The second half prices the good_with_kb run with and without caching, after correcting gpt-oss-120b's cached price to Groq's published 0.075 USD in a local copy of the table.
  • Comes out:
text
  layout                 input tokens  shared prefix  share  cacheable (1024 min, 128 blocks)
  good                         16,721         15,361    92%                                 0
                        per request: 697 tokens, of which 667 are identical to the previous request
  good_with_kb                 47,153         44,525    94%                            44,160
  bad_metadata_first           17,177            200     1%                                 0
  bad_ticket_first             15,713            161     1%                                 0

  good_with_kb over 24 tickets: 47,153 input tokens, 1,080 output tokens
    openai/gpt-oss-120b    minimum  1024  cached 44,525  no cache 0.00772 USD  with cache 0.00438 USD  (43% saved)
    gemini-3.5-flash       minimum  4096  cached      0  no cache 0.08045 USD  with cache 0.08045 USD  (0% saved)
    gemini-3.5-flash-lite  minimum  4096  cached      0  no cache 0.01685 USD  with cache 0.01685 USD  (0% saved)

The good layout makes 92% of all input tokens shareable (667 of each request's 697 tokens). Putting one line of ticket metadata at the top drops that to 1%, and so does putting the ticket first: the same content, reordered, turns a cacheable workload into an uncacheable one. But note the last column: the good layout's 667-token prefix is below a 1,024-token minimum, so under that rule nothing is cached at all. Adding the help center to the prefix makes each request larger (1,965 tokens) and cacheable, and on gpt-oss-120b that saves 43% of the bill even though the prompt nearly tripled. On Gemini 3.5 Flash the shared prefix (about 1,936 tokens per request) is still under the 4,096-token minimum, so it saves nothing. Caching is a property of your layout and your provider's rules together; measure both.

These are upper bounds from our own tokenizer. Real hit rates also depend on timing (caches expire), on requests reaching the same server, and on the provider's block size. Read the provider's reported usage.cached_tokens (exposed by llm.py) to see what you actually got; the lab prints it.

SituationUse thisWhy
Stable instructions, variable ticketInstructions first, ticket last92% shareable versus 1% with the order reversed
Per-request metadata (ids, dates, tier)Put it in the final user messageAny change near the top breaks the whole prefix
Prefix below the provider's minimumAccept no caching, or add useful stable context667 tokens caches nothing under a 1,024 minimum
Tool definitions and examplesFixed order, byte-identical every callKey reordering or whitespace changes break the prefix

Batch and async pricing tiers

A batch tier accepts a file of requests and returns results later, in exchange for a discount. Groq charges 50% of the on-demand rate with a completion window you choose from 24 hours to 7 days, and the discount does not stack with caching; Gemini's batch prices for 3.5 Flash are half its standard rates (0.75 USD input, 4.50 USD output per million). cost_usd(usage, model, batch=True) applies each model's batch_discount.

SituationUse thisWhy
A customer is waiting for the replyLive (synchronous) callsBatch can take hours
Nightly re-triage, backfills, evaluation runs (Module 11)Batch tierHalf price and no rate-limit juggling
Large cached prefix on GroqCompare bothBatch does not stack with cache discounts there

Cost per task: triaging all 72 tickets

Price per token is an input to the metric that matters: cost per task, what it costs to produce one usable result. For triage, a task is one ticket that ends with a valid label, including retries.

python
# examples/m02_cost_per_task.py
"""Module 2: what triaging all 72 Brightlane tickets costs under each model in PRICES."""
import json

from m02_prompts import triage_messages

from supportdesk.data import load_tickets
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.tokens import count_messages, count_tokens

tickets = load_tickets()
input_tokens = [count_messages(triage_messages(t)) for t in tickets]


def expected_reply(t) -> str:
    """A reply shaped like a correct answer, to size the output (the summary is a stand-in sentence)."""
    g = t.gold
    return json.dumps({"category": g["category"], "priority": g["priority"], "language": t.language,
                       "summary": f"Customer writes about: {t.subject}.", "needs_human": not g["answerable"]},
                      separators=(",", ":"))


output_tokens = [count_tokens(expected_reply(t)) for t in tickets]
print(f"{len(tickets)} tickets: input {sum(input_tokens):,} tokens (mean {sum(input_tokens) / len(tickets):.0f}), "
      f"output {sum(output_tokens):,} tokens (mean {sum(output_tokens) / len(tickets):.0f})")

REASONING = {"openai/gpt-oss-120b": 300, "openai/gpt-oss-20b": 300}  # assumed hidden tokens per call; measure yours
SUCCESS = 0.95  # assumed share of replies that parse and validate the first time; failures are retried once

attempts = 1 + (1 - SUCCESS)             # expected calls per ticket with one retry
succeeded = 1 - (1 - SUCCESS) ** 2        # share of tickets that end with a valid label
print(f"\n{'model':26}{'all 72, live':>14}{'all 72, batch':>15}{'per 1,000 tickets':>19}{'per 1,000 labels':>18}")
for model in PRICES:
    extra = REASONING.get(model, 0)
    live = sum(cost_usd(Usage(input_tokens=i, output_tokens=o + extra, reasoning_tokens=extra), model)
               for i, o in zip(input_tokens, output_tokens))
    batch = sum(cost_usd(Usage(input_tokens=i, output_tokens=o + extra, reasoning_tokens=extra), model, batch=True)
                for i, o in zip(input_tokens, output_tokens))
    per_ticket = live / len(tickets)
    per_label = per_ticket * attempts / succeeded  # retries cost money, and a few tickets still fail
    note = "  +300 reasoning/call (assumed)" if extra else ""
    print(f"{model:26}{live:>10.4f} USD{batch:>11.4f} USD{per_ticket * 1000:>15.4f} USD{per_label * 1000:>14.4f} USD{note}")

Code explained

  • In simple words: price the whole triage job on every model in the table, then convert it into cost per usable label.
  • What happens: input tokens are measured with count_messages() for all 72 real triage requests. Output is sized by counting a compact JSON reply built from each ticket's gold labels, with a stand-in summary sentence. Two numbers are assumptions, marked in the code: 300 hidden reasoning tokens per call for the gpt-oss models (measure yours with the calibration script) and a 95% first-try validity rate with one retry. per 1,000 labels divides the expected spend (1.05 calls per ticket) by the share of tickets that end with a valid label (99.75%).
  • Comes out:
text
  72 tickets: input 47,877 tokens (mean 665), output 2,279 tokens (mean 32)

  model                       all 72, live  all 72, batch  per 1,000 tickets  per 1,000 labels
  openai/gpt-oss-120b           0.0215 USD     0.0108 USD         0.2987 USD        0.3145 USD  +300 reasoning/call (assumed)
  openai/gpt-oss-20b            0.0108 USD     0.0054 USD         0.1494 USD        0.1572 USD  +300 reasoning/call (assumed)
  llama-3.1-8b-instant          0.0026 USD     0.0013 USD         0.0358 USD        0.0377 USD
  llama-3.3-70b-versatile       0.0300 USD     0.0150 USD         0.4173 USD        0.4393 USD
  gemini-3.5-flash              0.0923 USD     0.0462 USD         1.2823 USD        1.3498 USD
  gemini-3.5-flash-lite         0.0201 USD     0.0100 USD         0.2786 USD        0.2933 USD
  qwen3:8b                      0.0000 USD     0.0000 USD         0.0000 USD        0.0000 USD

Triaging all 72 tickets costs between a quarter of a cent and 9 cents depending on the model, so the dataset itself is cheap; the per-1,000 column is the one to plan with. Input dominates volume (665 tokens in, 32 out per ticket), which is why caching the system prompt matters more for triage than trimming the output. The reasoning assumption decides the gpt-oss ranking: without its 300 hidden tokens per call, gpt-oss-120b would cost about 0.12 USD per 1,000 tickets instead of 0.30. Gemini 3.5 Flash is the most expensive here mainly because of its input price; it also bills any thinking tokens as output, which this table sets to zero. The local qwen3:8b shows zero, but your hardware and electricity are not free and are not in the table.

Cost is only half of the decision. A cheaper model that labels 20% of tickets wrongly costs you agent time, which is far more expensive than tokens. Module 11 builds the evaluation harness that measures accuracy per model; combine its accuracy with this table to get cost per correct label.

Module Lab

The lab joins the module's pieces into one run: count every triage request before sending, check it against the model's window, send it, accumulate reported usage, compare estimate with reported usage, measure the shareable prefix, price the run live and in batch, and compact a long chat with a fact-survival guard that falls back to importance ranking.

python
# examples/m02_lab.py
"""Module 2 lab: a token-aware, budget-checked, cost-reported triage run over the 24 test tickets.

Default: a ScriptedLLM stand-in (plumbing only; its usage numbers are our own estimates).
With a key: PYTHONPATH=. python examples/m02_lab.py --real
"""
import argparse
import json

from m02_cache_layout import common_prefix, render
from m02_context import compact_with_summary, importance_ranked
from m02_conversation import conversation, facts_kept
from m02_prompts import triage_messages

from supportdesk.data import load_tickets
from supportdesk.llm import Usage, chat, resolve
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages

WINDOWS = {"openai/gpt-oss-120b": 131_072, "gemini-3.5-flash": 1_048_576, "qwen3:8b": 4_096}
MAX_TOKENS = 512  # room for the JSON label plus hidden reasoning on reasoning models


def gold_responder(messages, kwargs):
    """Stand-in that answers with the gold label: tests the loop, says nothing about model quality."""
    ticket = next(t for t in load_tickets("test") if t.text in messages[-1]["content"])
    return json.dumps({"category": ticket.gold["category"], "priority": ticket.gold["priority"],
                       "language": ticket.language, "summary": f"About: {ticket.subject}",
                       "needs_human": not ticket.gold["answerable"]})


def triage_run(llm, model: str) -> None:
    window = WINDOWS.get(model, 8_192)
    total, estimated, previous, shared, rendered, valid = Usage(), 0, None, 0, 0, 0
    for ticket in load_tickets("test"):
        messages = triage_messages(ticket)
        estimate = count_messages(messages)
        if estimate + MAX_TOKENS > window * 0.95:
            print(f"  {ticket.id}: skipped, {estimate} tokens does not fit {window}")
            continue
        result = llm(messages, max_tokens=MAX_TOKENS, temperature=0)
        estimated += estimate
        for name in ("input_tokens", "output_tokens", "cached_tokens", "reasoning_tokens"):
            setattr(total, name, getattr(total, name) + getattr(result.usage, name))
        tokens = render(messages)
        shared += common_prefix(previous, tokens) if previous else 0
        previous, rendered = tokens, rendered + len(tokens)
        try:
            valid += json.loads(result.text)["category"] is not None
        except (json.JSONDecodeError, KeyError, TypeError):
            pass
    print(f"  calls: 24   valid JSON labels: {valid}/24")
    print(f"  input tokens: estimated {estimated:,}, reported {total.input_tokens:,} "
          f"(ratio {total.input_tokens / estimated:.2f})")
    print(f"  output tokens: {total.output_tokens:,} (reasoning {total.reasoning_tokens:,})   "
          f"cached reported: {total.cached_tokens:,}   shareable prefix: {shared:,} ({shared / rendered:.0%})")
    if model in PRICES:
        print(f"  cost: {cost_usd(total, model):.5f} USD live, {cost_usd(total, model, batch=True):.5f} USD batch, "
              f"{cost_usd(total, model) / 24 * 1000:.4f} USD per 1,000 tickets")


def compaction_check(summarizer, budget: int = 300) -> None:
    full = conversation()
    compacted = compact_with_summary(full, summarizer)
    facts = facts_kept(compacted)
    lost = [k for k, ok in facts.items() if not ok]
    over = count_messages(compacted) > budget
    if lost or over:
        print(f"  summary rejected ({'lost ' + ', '.join(lost) if lost else 'over budget'}); "
              "falling back to importance-ranked")
        compacted = importance_ranked(full, budget)
    print(f"  final: {count_messages(compacted)} tokens (from {count_messages(full)}), facts: {facts_kept(compacted)}")


if __name__ == "__main__":
    parser = argparse.ArgumentParser()
    parser.add_argument("--real", action="store_true")
    args = parser.parse_args()
    if args.real:
        _, model = resolve()
        llm, summarizer = chat, chat
    else:
        model = "openai/gpt-oss-120b"  # priced as if it were this model
        llm = ScriptedLLM(responder=gold_responder)
        summarizer = ScriptedLLM(replies=["Customer asked about Slack, exports, and automations; all resolved."])
    print(f"Triage run on {model} ({'real' if args.real else 'stand-in, usage is estimated'}):")
    triage_run(llm, model)
    print("Compaction with a fact-survival guard:")
    compaction_check(summarizer)

Code explained

  • In simple words: a cost- and context-aware triage run you can point at any provider, plus a compaction step that refuses to lose key facts.
  • What happens: triage_run() loops over the 24 test tickets. For each it estimates tokens, skips any request that would not fit 95% of the window after reserving MAX_TOKENS, calls the model, adds up reported usage (including cached and reasoning tokens), tracks the shareable prefix with Part D's render() and common_prefix(), and checks that the reply parses as JSON. It then prints the estimate-to-reported ratio and costs via cost_usd. compaction_check() summarizes the chat, runs facts_kept(), and falls back to importance_ranked() if any fact is lost or the result is over budget. By default the model is a ScriptedLLM that returns each ticket's gold label and the summarizer is a vague scripted summary: plumbing only. With --real, both use llm.chat with your configured provider.
  • Comes out (stand-in, captured)
text
  Triage run on openai/gpt-oss-120b (stand-in, usage is estimated):
    calls: 24   valid JSON labels: 24/24
    input tokens: estimated 15,929, reported 15,929 (ratio 1.00)
    output tokens: 870 (reasoning 0)   cached reported: 0   shareable prefix: 15,361 (92%)
    cost: 0.00291 USD live, 0.00146 USD batch, 0.1213 USD per 1,000 tickets
  Compaction with a fact-survival guard:
    summary rejected (lost invoice id, plan and seats, customer's ask); falling back to importance-ranked
    final: 295 tokens (from 558), facts: {'invoice id': True, 'plan and seats': True, "customer's ask": True}

Everything here is plumbing: 24 of 24 labels are valid because the stand-in returns gold labels, the ratio is 1.00 because the stand-in reports our own estimate, and no tokens are cached because no provider is involved. The fallback path is exercised for real: the vague summary lost all three facts, the guard caught it, and importance ranking delivered a 295-token context with every fact intact. Run it with --real to get your provider's numbers: expect a ratio a little above 1, reasoning tokens if you use gpt-oss-120b, and cached tokens only if the prefix clears your provider's minimum.

Project Milestone

After this module, your supportdesk working copy contains:

  • supportdesk/tokens.py and supportdesk/pricing.py, read in full and used for every estimate (canonical; unchanged).
  • examples/m02_prompts.py: the cache-friendly triage request (stable prefix, ticket last).
  • examples/m02_conversation.py and examples/m02_context.py: a long support chat, a fact checker, and four context strategies.
  • examples/m02_budget.py, m02_overflow.py, m02_tinylm_limit.py: the budget calculator, the overflow guard, and the real TinyLM truncation demo.
  • examples/m02_needle.py: a position-effects harness verified against a planted blind spot, ready for --real.
  • examples/m02_cost.py, m02_cache_layout.py, m02_cost_per_task.py, m02_long_vs_retrieval.py: the cost model, the cache layout measurement, and cost per task.
  • examples/m02_lab.py: the combined run; tests/test_m02_tokens_context_cost.py: 9 passing tests.

Maya, the support lead, can now get a straight answer to "what would it cost to triage every ticket automatically?": about 0.30 USD per 1,000 tickets on the default Groq model under our stated reasoning assumption, confirmed or corrected by one --real run of the lab.

Interview Questions

1. Why can the same prompt cost different amounts on two models with the same per-token price? Because each model family has its own tokenizer, and the same text becomes a different number of tokens. On Brightlane's triage request, four tokenizers gave 652 to 723 tokens, an 11% spread, and for Hindi text the spread was 1.6x to 7.7x the English count. Compare cost per task with real usage, not price per token.

2. What is BPE and why does it make ids and numbers expensive? Byte pair encoding starts from bytes and repeatedly merges the most frequent adjacent pair into a new token. Frequent words become single tokens; rare strings stay split. Invoice ids and long numbers are rare and unpredictable, so they break into many pieces: INV-2026-004512 is 7 tokens in o200k_base, about 2 characters per token versus 5 for prose.

3. Why do models struggle to count letters or add large numbers? They see token ids, not characters. strawberry arrives as st|raw|berry, and 12345 + 67890 as 123|45| +| |678|90, with digit groups that do not align with place value. The fix is to do character and arithmetic work in code and give the model results, then validate anything the model returns.

4. How do you estimate tokens before sending a request, and how accurate is it? Count each message's content with a tokenizer close to the model's, add a few tokens per message for the chat template, and add the priming tokens (count_messages does this). Accuracy depends on tokenizer match and hidden template text; calibrate with a few real calls by dividing reported prompt_tokens by your estimate, and use that ratio plus a margin.

5. Write the context budget equation and explain the output reserve. System + tools + history + retrieved + user + output reserve + safety margin must not exceed the window. The output reserve is max_tokens: the window covers input and output together, and reasoning models spend part of it on hidden thinking, so a reasoning model needs a reserve in the thousands, not the length of the visible answer.

6. What happens when a request exceeds the context window? It depends on the stack. Hosted APIs usually reject it with HTTP 400 (often code context_length_exceeded). Local servers such as Ollama may silently keep only the newest tokens, and TinyLM's generate() does the same, dropping the start of the prompt with no signal. If only the output runs out, you get finish_reason == "length" and a cut-off reply. Budget before sending, catch the error, trim, retry once, and check finish_reason.

7. What is "lost in the middle", and how does it change prompt layout? Liu et al. (2023, TACL 2024) showed that models use information at the beginning and end of a long context better than information in the middle; in the worst case GPT-3.5-Turbo did worse with the answer document in the middle than with no documents. Put key instructions first and restate the task last, place the most relevant retrieved passage next to the question, and test your own model with a needle harness at your lengths.

8. Why is advertised context length not the same as usable context? The advertised figure is what the API accepts. Benchmarks such as RULER found only about half of models claiming 32K+ performed satisfactorily at 32K, and NoLiMa found 11 of 13 models fell below half their short-context score at 32K once literal word matching was removed. Measure effective context for your task and send the smallest context that contains what is needed.

9. Compare drop-oldest, sliding window, importance-ranked truncation, and summary compaction. Drop-oldest and sliding windows are cheap and predictable but lose early facts: on our 23-turn chat both lost the invoice id, plan, and request. Pinning the first message kept two of three. Importance ranking kept all three with no extra call because the scorer knew the fact types. Summary compaction compresses most but can silently lose facts, so verify each summary with an automatic fact check and fall back when it fails.

10. Why is output usually the expensive half of a bill? Output tokens are generated one at a time, each needing a pass through the model, while input tokens are processed in parallel, so providers price output higher: 6x input on gemini-3.5-flash and 4x on gpt-oss-120b. A triage call that is 6% output tokens spent 29% of its cost on output; a draft reply at 17% output spent 56%. Reasoning tokens are billed as output too.

11. How do you lay out a prompt to benefit from prompt caching? Put everything stable first in a fixed, byte-identical order (instructions, tool definitions, examples, reference text) and everything variable last (ticket, ids, timestamps). On 24 triage requests that made 92% of tokens shareable, versus 1% when a single metadata line sat at the top. Then check the provider's minimum prefix: a 667-token prefix caches nothing under a 1,024 minimum, and Gemini 3.5 Flash needs 4,096. Confirm with reported cached_tokens.

12. What is cost per task and why is it the metric that matters? It is total spend divided by successful results, including retries, reasoning tokens, and failures, for example cost per valid triage label. Price per token ignores output length, hidden reasoning, cache hit rates, and retries; cost per task includes them all. Pair it with accuracy from an evaluation to get cost per correct result.

Other Tools and Providers

What we usedAlternativesNotes
tiktoken with vendored o200k_base, cl100k_base, p50k_baseProvider count-token endpoints (Gemini countTokens, Anthropic token counting), Hugging Face tokenizers / transformers AutoTokenizer for open-weight modelsExact counts need the model's own tokenizer; endpoints cost a network call
tokenizers BPE trainerSentencePiece (BPE and unigram), tiktoken's educational BPE moduleSentencePiece is common in Llama, Gemma, and T5 families
count_messages estimateProvider-reported usage, observability tools such as Langfuse, Helicone, or OpenTelemetry GenAI tracesEstimates for budgeting; reported usage for billing (Module 13)
Hand-written truncation and compactionFramework memory utilities (LangChain, LlamaIndex chat memory), provider-side conversation state or compaction featuresFrameworks save code; check what they drop and add a fact check
Our needle harnessRULER, NoLiMa, LongBench, Greg Kamradt's original needle-in-a-haystack testPublic suites for comparing models; your harness for your data
pricing.py tableProvider pricing pages, LiteLLM's model cost map, OpenRouter's model pagesKeep one table in code and verify it on a schedule
Provider-side prompt cachingExplicit caches (Gemini cached contents, Anthropic cache_control), self-hosted prefix caching (vLLM automatic prefix caching, SGLang RadixAttention)Self-hosting makes caching your job and your saving
Batch tiers (Groq, Gemini)OpenAI Batch API, Anthropic Message BatchesTypically half price for asynchronous jobs

Coming Up in Module 3

You now know what a request is made of and what it costs. Module 3, Inference Behaviour and Decoding Control, looks at how the response is produced: prefill and decode as separate phases (and why latency grows with output length, which is the other half of "output is expensive"), time to first token and streaming, what temperature, top-p, top-k, and min-p actually change in TinyLM's probability distributions, stop sequences and max_tokens, constrained decoding, and how reasoning models spend thinking tokens, with budgets and effort levels you can measure.