Part D: Cost mechanics
The canonical price table: supportdesk/pricing.py
Every dollar figure in the course comes from one table, so a price change is a one-line edit.<strong>supportdesk/pricing.py</strong>
# supportdesk/pricing.py
"""Token prices per million tokens, used to turn usage into dollars.
Prices change often. These were checked on 21 September 2026 from the providers'
pricing pages (Groq figures via a third-party summary of Groq's page, so treat
them as directional). Verify current prices before relying on them, and edit
this table rather than hard-coding prices anywhere else.
"""
from __future__ import annotations
from dataclasses import dataclass
from supportdesk.llm import Usage
@dataclass(frozen=True)
class Price:
input: float # USD per 1M input tokens
output: float # USD per 1M output tokens (reasoning tokens bill as output)
cached_input: float # USD per 1M cached input tokens
batch_discount: float = 0.5
PRICES: dict[str, Price] = {
"openai/gpt-oss-120b": Price(input=0.15, output=0.60, cached_input=0.15),
"openai/gpt-oss-20b": Price(input=0.075, output=0.30, cached_input=0.075),
"llama-3.1-8b-instant": Price(input=0.05, output=0.08, cached_input=0.05),
"llama-3.3-70b-versatile": Price(input=0.59, output=0.79, cached_input=0.59),
"gemini-3.5-flash": Price(input=1.50, output=9.00, cached_input=0.15),
"gemini-3.5-flash-lite": Price(input=0.30, output=2.50, cached_input=0.03),
"qwen3:8b": Price(input=0.0, output=0.0, cached_input=0.0, batch_discount=0.0),
}
def cost_usd(usage: Usage, model: str, batch: bool = False) -> float:
"""Dollar cost of one call. Cached input tokens are billed at the cached rate."""
price = PRICES[model]
uncached = max(usage.input_tokens - usage.cached_tokens, 0)
total = (
uncached * price.input
+ usage.cached_tokens * price.cached_input
+ usage.output_tokens * price.output
) / 1_000_000
return total * (1 - price.batch_discount) if batch else totalCode explained
- In simple words: a price list per model and a function that turns a call's token usage into dollars.
- What happens (per part):
- Module docstring: prices were checked on 21 September 2026 and need verifying before you rely on them.
Price: a frozen dataclass with USD per 1 million tokens forinput,output, andcached_input, plusbatch_discount(0.5 means half price). The comment records that reasoning tokens bill as output.PRICES: one entry per model id, keyed by the exact stringllm.chatreports inChatResult.model. Groq models, two Gemini models, andqwen3:8bat zero because Ollama runs on your own hardware (your electricity and hardware are not in the table).cost_usd(usage, model, batch): splits input into uncached and cached tokens, multiplies each count by its price, adds output, divides by one million, and applies the batch discount to the whole bill ifbatch=True.Usage.output_tokensalready includes reasoning tokens, because providers report reasoning inside the completion count, so they are priced as output without extra code.- Comes out: nothing by itself. Two notes from checking the table against provider pages on 21 Sep 2026. First, Groq's model pages list cached input for
openai/gpt-oss-120bat 0.075 USD and foropenai/gpt-oss-20bat 0.037 USD per million (a 50% discount), while the table has the full input price; Part D corrects this locally where it matters. Second, Groq states that its batch discount does not stack with caching (batch tokens bill at 50% regardless of cache status), socost_usd(..., batch=True)slightly underestimates a Groq batch with cached tokens. Gemini 3.5 Flash's figures match Google's pricing page: 1.50 USD input, 9.00 USD output including thinking tokens, 0.15 USD cached input, batch at half price.
Input, output, and why output is the expensive half
# examples/m02_cost.py
"""Module 2: turn usage into dollars, and see where the money goes."""
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, cost_usd
# Realistic shapes for three Brightlane calls (token counts measured earlier in this module
# for triage; the reply and agent shapes are typical sizes, not measurements).
CALLS = {
"triage (JSON label)": Usage(input_tokens=664, output_tokens=45),
"draft reply": Usage(input_tokens=1_200, output_tokens=250),
"draft reply, reasoning": Usage(input_tokens=1_200, output_tokens=1_050, reasoning_tokens=800),
}
print(f"{'model':26}" + "".join(f"{name:>26}" for name in CALLS))
for model in PRICES:
print(f"{model:26}" + "".join(f"{cost_usd(u, model) * 1000:>21.4f} USD" for u in CALLS.values()))
print("(cost per 1,000 calls)\n")
model = "gemini-3.5-flash"
p = PRICES[model]
print(f"{model}: output costs {p.output / p.input:.0f}x input per token")
for name, u in CALLS.items():
out_share = u.output_tokens * p.output / (u.input_tokens * p.input + u.output_tokens * p.output)
print(f" {name:24} {u.output_tokens / (u.input_tokens + u.output_tokens):4.0%} of tokens are output, "
f"{out_share:4.0%} of the cost")
u = CALLS["draft reply, reasoning"]
visible = Usage(input_tokens=u.input_tokens, output_tokens=u.output_tokens - u.reasoning_tokens)
print(f"\nReasoning tokens bill as output: {cost_usd(u, model) / cost_usd(visible, model):.1f}x the cost of the "
f"same reply without {u.reasoning_tokens} hidden reasoning tokens")
print(f"Batch tier for 10,000 draft replies on {model}: "
f"{cost_usd(CALLS['draft reply'], model) * 10_000:.2f} USD live, "
f"{cost_usd(CALLS['draft reply'], model, batch=True) * 10_000:.2f} USD batch")Code explained
- In simple words: price three realistic Brightlane calls on every model, then see which half of the bill is input and which is output.
- What happens: the triage shape (664 in, 45 out) comes from Part A's measurement; the draft-reply shapes are typical sizes, not measurements. The reasoning variant has 800 hidden reasoning tokens inside its 1,050 output tokens, the same way providers report them. We print cost per 1,000 calls per model, then for gemini-3.5-flash the output share of tokens versus cost, the reasoning multiplier, and live versus batch pricing.
- Comes out:
model triage (JSON label) draft reply draft reply, reasoning
openai/gpt-oss-120b 0.1266 USD 0.3300 USD 0.8100 USD
openai/gpt-oss-20b 0.0633 USD 0.1650 USD 0.4050 USD
llama-3.1-8b-instant 0.0368 USD 0.0800 USD 0.1440 USD
llama-3.3-70b-versatile 0.4273 USD 0.9055 USD 1.5375 USD
gemini-3.5-flash 1.4010 USD 4.0500 USD 11.2500 USD
gemini-3.5-flash-lite 0.3117 USD 0.9850 USD 2.9850 USD
qwen3:8b 0.0000 USD 0.0000 USD 0.0000 USD
(cost per 1,000 calls)
gemini-3.5-flash: output costs 6x input per token
triage (JSON label) 6% of tokens are output, 29% of the cost
draft reply 17% of tokens are output, 56% of the cost
draft reply, reasoning 47% of tokens are output, 84% of the cost
Reasoning tokens bill as output: 2.8x the cost of the same reply without 800 hidden reasoning tokens
Batch tier for 10,000 draft replies on gemini-3.5-flash: 40.50 USD live, 20.25 USD batchOn gemini-3.5-flash an output token costs 6 times an input token. So a triage call whose tokens are only 6% output spends 29% of its money on output, and a draft reply that is 17% output spends 56% on it. Output is priced higher because it is produced one token at a time in the decode phase, each step a full pass through the model, while input is processed in parallel in the prefill phase (Module 3 measures both on TinyLM). On Groq the ratio is 4 for gpt-oss and 1.3 to 1.6 for the Llama models, so the effect is smaller but has the same direction.
Reasoning tokens are the hidden thinking a reasoning model generates before its visible answer. You never see them in ChatResult.text, but you pay for them as output, and they count against max_tokens. In this example 800 reasoning tokens make the reply 2.8 times as expensive as the same visible answer without them. llm.py exposes them as usage.reasoning_tokens when the provider reports them. Controlling reasoning effort is Module 3's topic; the cost lesson is to measure reasoning tokens per task before you choose a reasoning model for a high-volume job.
| Situation | Use this | Why |
|---|---|---|
| Classification or extraction | Short structured output, low or no reasoning | Output tokens are the expensive ones |
| Long drafts | Ask for the length you need; set max_tokens | Every extra sentence is billed at the output rate |
| Reasoning model on a simple task | Low reasoning effort, or a non-reasoning model | Hidden tokens can multiply the bill |
| Comparing models | Compare cost per task with real usage | Per-token prices hide output length and reasoning |
Prompt caching and cache-aware layout
Prompt caching lets a provider reuse the work it did on a prompt's beginning when the next request starts with exactly the same tokens. Cached input tokens are billed at a discount and processed faster. Caching works on prefixes: a request hits the cache only for the tokens that are identical from the very first token up to the first difference. One changed character early in the prompt breaks the cache for everything after it. Provider details differ (all checked 21 Sep 2026):
| Provider | How it caches | Discount on cached input | Minimum prefix |
|---|---|---|---|
| Groq (gpt-oss models) | Automatic | 50% | 128 to 1,024 tokens depending on model; expires after 2 hours unused |
| Gemini (2.5 and newer) | Implicit, automatic; explicit caches optional | 90% on 3.5 Flash (0.15 vs 1.50 USD), explicit caches also pay storage per hour | 4,096 tokens for 3.5 Flash |
| Anthropic | You mark breakpoints with cache_control | Reads at 0.1x input; writes cost 1.25x (5-minute cache) or 2x (1-hour) | 512 to 4,096 tokens by model |
| OpenAI | Automatic | Varies by model | 1,024 tokens, reused in 128-token blocks |
So the layout rule is: stable content first, variable content last. Instructions, tool definitions, examples, and reference articles go at the top in a fixed order; the ticket, timestamps, and user-specific data go at the end. Let's measure how much of each triage request is shareable across the 24 test tickets under four layouts.
# examples/m02_cache_layout.py
"""Module 2: measure how much of each triage request a prefix cache could reuse, for three layouts."""
from dataclasses import replace
from m02_prompts import EXAMPLES, TRIAGE_INSTRUCTIONS, triage_messages
from supportdesk.data import load_articles, load_tickets
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.tokens import encode
tickets = load_tickets("test") # 24 tickets, sent one after another
KB = "\n\n".join(f"[{a.id}] {a.title}\n{a.body}" for a in load_articles())
def render(messages: list[dict]) -> list[int]:
"""Approximate the provider's chat template: role header, content, separator."""
return encode("".join(f"<|{m['role']}|>\n{m['content']}\n<|end|>\n" for m in messages))
def good(t):
return triage_messages(t)
def good_with_kb(t):
messages = triage_messages(t)
messages[0] = {"role": "system", "content": TRIAGE_INSTRUCTIONS + "\n\nHelp-center articles:\n" + KB}
return messages
def bad_metadata_first(t):
header = f"Ticket {t.id}, received 2026-09-21, customer tier {t.customer_tier}.\n"
messages = triage_messages(t)
messages[0] = {"role": "system", "content": header + TRIAGE_INSTRUCTIONS}
return messages
def bad_ticket_first(t):
shots = "\n\n".join(f"Example ticket:\n{u}\nExample answer:\n{a}" for u, a in EXAMPLES)
return [{"role": "user", "content": f"{t.text}\n\n---\n{TRIAGE_INSTRUCTIONS}\n\n{shots}"}]
def common_prefix(a: list[int], b: list[int]) -> int:
n = 0
for x, y in zip(a, b):
if x != y:
break
n += 1
return n
def openai_style(prefix: int, minimum: int = 1024, block: int = 128) -> int:
"""Cached tokens under a rule like OpenAI's: nothing below the minimum, then whole 128-token blocks."""
return 0 if prefix < minimum else prefix // block * block
if __name__ == "__main__":
print(f"{'layout':22}{'input tokens':>13}{'shared prefix':>15}{'share':>7}{'cacheable (1024 min, 128 blocks)':>34}")
for layout in (good, good_with_kb, bad_metadata_first, bad_ticket_first):
requests = [render(layout(t)) for t in tickets]
total = sum(len(r) for r in requests)
shared = [common_prefix(requests[i - 1], r) if i else 0 for i, r in enumerate(requests)]
cacheable = sum(openai_style(s) for s in shared)
print(f"{layout.__name__:22}{total:>13,}{sum(shared):>15,}{sum(shared) / total:>7.0%}{cacheable:>34,}")
if layout is good:
print(f"{'':22}per request: {len(requests[0])} tokens, of which {shared[1]} are identical to the previous request")
# What caching is worth for the good_with_kb layout over the 24 tickets.
# Groq's model page (checked 21 Sep 2026) lists cached input for gpt-oss-120b at 0.075 USD per 1M,
# a 50% discount; the canonical table has 0.15. Correct it locally for this estimate.
PRICES["openai/gpt-oss-120b"] = replace(PRICES["openai/gpt-oss-120b"], cached_input=0.075)
requests = [render(good_with_kb(t)) for t in tickets]
shared = [common_prefix(requests[i - 1], r) if i else 0 for i, r in enumerate(requests)]
total_in, total_out = sum(len(r) for r in requests), 24 * 45 # about 45 output tokens per triage JSON
print(f"\ngood_with_kb over 24 tickets: {total_in:,} input tokens, {total_out:,} output tokens")
# Minimum cacheable prefix: Groq documents 128 to 1024 by model (1024 assumed); Gemini 3.5 Flash: 4,096.
for model, minimum in [("openai/gpt-oss-120b", 1024), ("gemini-3.5-flash", 4096), ("gemini-3.5-flash-lite", 4096)]:
cached = sum(s if s >= minimum else 0 for s in shared)
plain = cost_usd(Usage(input_tokens=total_in, output_tokens=total_out), model)
warm = cost_usd(Usage(input_tokens=total_in, output_tokens=total_out, cached_tokens=cached), model)
print(f" {model:22} minimum {minimum:>5} cached {cached:>6,} no cache {plain:.5f} USD with cache {warm:.5f} USD "
f"({1 - warm / plain:.0%} saved)")Code explained
- In simple words: we line up 24 consecutive requests and count how many tokens at the start of each are identical to the one before, which is the most any prefix cache could reuse.
- What happens:
render()approximates the provider's chat template by joining role headers and contents, then tokenizes witho200k_base. The four layouts are:good(instructions and examples first, ticket last),good_with_kb(the same with all 12 help-center articles in the system prompt),bad_metadata_first(a ticket id and date line at the top of the system prompt, a very common habit), andbad_ticket_first(the ticket before the instructions).common_prefix()counts identical leading tokens between consecutive requests;openai_style()applies a 1,024-token minimum and 128-token blocks. The second half prices thegood_with_kbrun with and without caching, after correcting gpt-oss-120b's cached price to Groq's published 0.075 USD in a local copy of the table. - Comes out:
layout input tokens shared prefix share cacheable (1024 min, 128 blocks)
good 16,721 15,361 92% 0
per request: 697 tokens, of which 667 are identical to the previous request
good_with_kb 47,153 44,525 94% 44,160
bad_metadata_first 17,177 200 1% 0
bad_ticket_first 15,713 161 1% 0
good_with_kb over 24 tickets: 47,153 input tokens, 1,080 output tokens
openai/gpt-oss-120b minimum 1024 cached 44,525 no cache 0.00772 USD with cache 0.00438 USD (43% saved)
gemini-3.5-flash minimum 4096 cached 0 no cache 0.08045 USD with cache 0.08045 USD (0% saved)
gemini-3.5-flash-lite minimum 4096 cached 0 no cache 0.01685 USD with cache 0.01685 USD (0% saved)The good layout makes 92% of all input tokens shareable (667 of each request's 697 tokens). Putting one line of ticket metadata at the top drops that to 1%, and so does putting the ticket first: the same content, reordered, turns a cacheable workload into an uncacheable one. But note the last column: the good layout's 667-token prefix is below a 1,024-token minimum, so under that rule nothing is cached at all. Adding the help center to the prefix makes each request larger (1,965 tokens) and cacheable, and on gpt-oss-120b that saves 43% of the bill even though the prompt nearly tripled. On Gemini 3.5 Flash the shared prefix (about 1,936 tokens per request) is still under the 4,096-token minimum, so it saves nothing. Caching is a property of your layout and your provider's rules together; measure both.
These are upper bounds from our own tokenizer. Real hit rates also depend on timing (caches expire), on requests reaching the same server, and on the provider's block size. Read the provider's reported usage.cached_tokens (exposed by llm.py) to see what you actually got; the lab prints it.
| Situation | Use this | Why |
|---|---|---|
| Stable instructions, variable ticket | Instructions first, ticket last | 92% shareable versus 1% with the order reversed |
| Per-request metadata (ids, dates, tier) | Put it in the final user message | Any change near the top breaks the whole prefix |
| Prefix below the provider's minimum | Accept no caching, or add useful stable context | 667 tokens caches nothing under a 1,024 minimum |
| Tool definitions and examples | Fixed order, byte-identical every call | Key reordering or whitespace changes break the prefix |
Batch and async pricing tiers
A batch tier accepts a file of requests and returns results later, in exchange for a discount. Groq charges 50% of the on-demand rate with a completion window you choose from 24 hours to 7 days, and the discount does not stack with caching; Gemini's batch prices for 3.5 Flash are half its standard rates (0.75 USD input, 4.50 USD output per million). cost_usd(usage, model, batch=True) applies each model's batch_discount.
| Situation | Use this | Why |
|---|---|---|
| A customer is waiting for the reply | Live (synchronous) calls | Batch can take hours |
| Nightly re-triage, backfills, evaluation runs (Module 11) | Batch tier | Half price and no rate-limit juggling |
| Large cached prefix on Groq | Compare both | Batch does not stack with cache discounts there |
Cost per task: triaging all 72 tickets
Price per token is an input to the metric that matters: cost per task, what it costs to produce one usable result. For triage, a task is one ticket that ends with a valid label, including retries.
# examples/m02_cost_per_task.py
"""Module 2: what triaging all 72 Brightlane tickets costs under each model in PRICES."""
import json
from m02_prompts import triage_messages
from supportdesk.data import load_tickets
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.tokens import count_messages, count_tokens
tickets = load_tickets()
input_tokens = [count_messages(triage_messages(t)) for t in tickets]
def expected_reply(t) -> str:
"""A reply shaped like a correct answer, to size the output (the summary is a stand-in sentence)."""
g = t.gold
return json.dumps({"category": g["category"], "priority": g["priority"], "language": t.language,
"summary": f"Customer writes about: {t.subject}.", "needs_human": not g["answerable"]},
separators=(",", ":"))
output_tokens = [count_tokens(expected_reply(t)) for t in tickets]
print(f"{len(tickets)} tickets: input {sum(input_tokens):,} tokens (mean {sum(input_tokens) / len(tickets):.0f}), "
f"output {sum(output_tokens):,} tokens (mean {sum(output_tokens) / len(tickets):.0f})")
REASONING = {"openai/gpt-oss-120b": 300, "openai/gpt-oss-20b": 300} # assumed hidden tokens per call; measure yours
SUCCESS = 0.95 # assumed share of replies that parse and validate the first time; failures are retried once
attempts = 1 + (1 - SUCCESS) # expected calls per ticket with one retry
succeeded = 1 - (1 - SUCCESS) ** 2 # share of tickets that end with a valid label
print(f"\n{'model':26}{'all 72, live':>14}{'all 72, batch':>15}{'per 1,000 tickets':>19}{'per 1,000 labels':>18}")
for model in PRICES:
extra = REASONING.get(model, 0)
live = sum(cost_usd(Usage(input_tokens=i, output_tokens=o + extra, reasoning_tokens=extra), model)
for i, o in zip(input_tokens, output_tokens))
batch = sum(cost_usd(Usage(input_tokens=i, output_tokens=o + extra, reasoning_tokens=extra), model, batch=True)
for i, o in zip(input_tokens, output_tokens))
per_ticket = live / len(tickets)
per_label = per_ticket * attempts / succeeded # retries cost money, and a few tickets still fail
note = " +300 reasoning/call (assumed)" if extra else ""
print(f"{model:26}{live:>10.4f} USD{batch:>11.4f} USD{per_ticket * 1000:>15.4f} USD{per_label * 1000:>14.4f} USD{note}")Code explained
- In simple words: price the whole triage job on every model in the table, then convert it into cost per usable label.
- What happens: input tokens are measured with
count_messages()for all 72 real triage requests. Output is sized by counting a compact JSON reply built from each ticket's gold labels, with a stand-in summary sentence. Two numbers are assumptions, marked in the code: 300 hidden reasoning tokens per call for the gpt-oss models (measure yours with the calibration script) and a 95% first-try validity rate with one retry.per 1,000 labelsdivides the expected spend (1.05 calls per ticket) by the share of tickets that end with a valid label (99.75%). - Comes out:
72 tickets: input 47,877 tokens (mean 665), output 2,279 tokens (mean 32)
model all 72, live all 72, batch per 1,000 tickets per 1,000 labels
openai/gpt-oss-120b 0.0215 USD 0.0108 USD 0.2987 USD 0.3145 USD +300 reasoning/call (assumed)
openai/gpt-oss-20b 0.0108 USD 0.0054 USD 0.1494 USD 0.1572 USD +300 reasoning/call (assumed)
llama-3.1-8b-instant 0.0026 USD 0.0013 USD 0.0358 USD 0.0377 USD
llama-3.3-70b-versatile 0.0300 USD 0.0150 USD 0.4173 USD 0.4393 USD
gemini-3.5-flash 0.0923 USD 0.0462 USD 1.2823 USD 1.3498 USD
gemini-3.5-flash-lite 0.0201 USD 0.0100 USD 0.2786 USD 0.2933 USD
qwen3:8b 0.0000 USD 0.0000 USD 0.0000 USD 0.0000 USDTriaging all 72 tickets costs between a quarter of a cent and 9 cents depending on the model, so the dataset itself is cheap; the per-1,000 column is the one to plan with. Input dominates volume (665 tokens in, 32 out per ticket), which is why caching the system prompt matters more for triage than trimming the output. The reasoning assumption decides the gpt-oss ranking: without its 300 hidden tokens per call, gpt-oss-120b would cost about 0.12 USD per 1,000 tickets instead of 0.30. Gemini 3.5 Flash is the most expensive here mainly because of its input price; it also bills any thinking tokens as output, which this table sets to zero. The local qwen3:8b shows zero, but your hardware and electricity are not free and are not in the table.
Cost is only half of the decision. A cheaper model that labels 20% of tickets wrongly costs you agent time, which is far more expensive than tokens. Module 11 builds the evaluation harness that measures accuracy per model; combine its accuracy with this table to get cost per correct label.
Module Lab
The lab joins the module's pieces into one run: count every triage request before sending, check it against the model's window, send it, accumulate reported usage, compare estimate with reported usage, measure the shareable prefix, price the run live and in batch, and compact a long chat with a fact-survival guard that falls back to importance ranking.
# examples/m02_lab.py
"""Module 2 lab: a token-aware, budget-checked, cost-reported triage run over the 24 test tickets.
Default: a ScriptedLLM stand-in (plumbing only; its usage numbers are our own estimates).
With a key: PYTHONPATH=. python examples/m02_lab.py --real
"""
import argparse
import json
from m02_cache_layout import common_prefix, render
from m02_context import compact_with_summary, importance_ranked
from m02_conversation import conversation, facts_kept
from m02_prompts import triage_messages
from supportdesk.data import load_tickets
from supportdesk.llm import Usage, chat, resolve
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages
WINDOWS = {"openai/gpt-oss-120b": 131_072, "gemini-3.5-flash": 1_048_576, "qwen3:8b": 4_096}
MAX_TOKENS = 512 # room for the JSON label plus hidden reasoning on reasoning models
def gold_responder(messages, kwargs):
"""Stand-in that answers with the gold label: tests the loop, says nothing about model quality."""
ticket = next(t for t in load_tickets("test") if t.text in messages[-1]["content"])
return json.dumps({"category": ticket.gold["category"], "priority": ticket.gold["priority"],
"language": ticket.language, "summary": f"About: {ticket.subject}",
"needs_human": not ticket.gold["answerable"]})
def triage_run(llm, model: str) -> None:
window = WINDOWS.get(model, 8_192)
total, estimated, previous, shared, rendered, valid = Usage(), 0, None, 0, 0, 0
for ticket in load_tickets("test"):
messages = triage_messages(ticket)
estimate = count_messages(messages)
if estimate + MAX_TOKENS > window * 0.95:
print(f" {ticket.id}: skipped, {estimate} tokens does not fit {window}")
continue
result = llm(messages, max_tokens=MAX_TOKENS, temperature=0)
estimated += estimate
for name in ("input_tokens", "output_tokens", "cached_tokens", "reasoning_tokens"):
setattr(total, name, getattr(total, name) + getattr(result.usage, name))
tokens = render(messages)
shared += common_prefix(previous, tokens) if previous else 0
previous, rendered = tokens, rendered + len(tokens)
try:
valid += json.loads(result.text)["category"] is not None
except (json.JSONDecodeError, KeyError, TypeError):
pass
print(f" calls: 24 valid JSON labels: {valid}/24")
print(f" input tokens: estimated {estimated:,}, reported {total.input_tokens:,} "
f"(ratio {total.input_tokens / estimated:.2f})")
print(f" output tokens: {total.output_tokens:,} (reasoning {total.reasoning_tokens:,}) "
f"cached reported: {total.cached_tokens:,} shareable prefix: {shared:,} ({shared / rendered:.0%})")
if model in PRICES:
print(f" cost: {cost_usd(total, model):.5f} USD live, {cost_usd(total, model, batch=True):.5f} USD batch, "
f"{cost_usd(total, model) / 24 * 1000:.4f} USD per 1,000 tickets")
def compaction_check(summarizer, budget: int = 300) -> None:
full = conversation()
compacted = compact_with_summary(full, summarizer)
facts = facts_kept(compacted)
lost = [k for k, ok in facts.items() if not ok]
over = count_messages(compacted) > budget
if lost or over:
print(f" summary rejected ({'lost ' + ', '.join(lost) if lost else 'over budget'}); "
"falling back to importance-ranked")
compacted = importance_ranked(full, budget)
print(f" final: {count_messages(compacted)} tokens (from {count_messages(full)}), facts: {facts_kept(compacted)}")
if __name__ == "__main__":
parser = argparse.ArgumentParser()
parser.add_argument("--real", action="store_true")
args = parser.parse_args()
if args.real:
_, model = resolve()
llm, summarizer = chat, chat
else:
model = "openai/gpt-oss-120b" # priced as if it were this model
llm = ScriptedLLM(responder=gold_responder)
summarizer = ScriptedLLM(replies=["Customer asked about Slack, exports, and automations; all resolved."])
print(f"Triage run on {model} ({'real' if args.real else 'stand-in, usage is estimated'}):")
triage_run(llm, model)
print("Compaction with a fact-survival guard:")
compaction_check(summarizer)Code explained
- In simple words: a cost- and context-aware triage run you can point at any provider, plus a compaction step that refuses to lose key facts.
- What happens:
triage_run()loops over the 24 test tickets. For each it estimates tokens, skips any request that would not fit 95% of the window after reservingMAX_TOKENS, calls the model, adds up reported usage (including cached and reasoning tokens), tracks the shareable prefix with Part D'srender()andcommon_prefix(), and checks that the reply parses as JSON. It then prints the estimate-to-reported ratio and costs viacost_usd.compaction_check()summarizes the chat, runsfacts_kept(), and falls back toimportance_ranked()if any fact is lost or the result is over budget. By default the model is aScriptedLLMthat returns each ticket's gold label and the summarizer is a vague scripted summary: plumbing only. With--real, both usellm.chatwith your configured provider. - Comes out (stand-in, captured)
Triage run on openai/gpt-oss-120b (stand-in, usage is estimated):
calls: 24 valid JSON labels: 24/24
input tokens: estimated 15,929, reported 15,929 (ratio 1.00)
output tokens: 870 (reasoning 0) cached reported: 0 shareable prefix: 15,361 (92%)
cost: 0.00291 USD live, 0.00146 USD batch, 0.1213 USD per 1,000 tickets
Compaction with a fact-survival guard:
summary rejected (lost invoice id, plan and seats, customer's ask); falling back to importance-ranked
final: 295 tokens (from 558), facts: {'invoice id': True, 'plan and seats': True, "customer's ask": True}Everything here is plumbing: 24 of 24 labels are valid because the stand-in returns gold labels, the ratio is 1.00 because the stand-in reports our own estimate, and no tokens are cached because no provider is involved. The fallback path is exercised for real: the vague summary lost all three facts, the guard caught it, and importance ranking delivered a 295-token context with every fact intact. Run it with --real to get your provider's numbers: expect a ratio a little above 1, reasoning tokens if you use gpt-oss-120b, and cached tokens only if the prefix clears your provider's minimum.
Project Milestone
After this module, your supportdesk working copy contains:
supportdesk/tokens.pyandsupportdesk/pricing.py, read in full and used for every estimate (canonical; unchanged).examples/m02_prompts.py: the cache-friendly triage request (stable prefix, ticket last).examples/m02_conversation.pyandexamples/m02_context.py: a long support chat, a fact checker, and four context strategies.examples/m02_budget.py,m02_overflow.py,m02_tinylm_limit.py: the budget calculator, the overflow guard, and the real TinyLM truncation demo.examples/m02_needle.py: a position-effects harness verified against a planted blind spot, ready for--real.examples/m02_cost.py,m02_cache_layout.py,m02_cost_per_task.py,m02_long_vs_retrieval.py: the cost model, the cache layout measurement, and cost per task.examples/m02_lab.py: the combined run;tests/test_m02_tokens_context_cost.py: 9 passing tests.
Maya, the support lead, can now get a straight answer to "what would it cost to triage every ticket automatically?": about 0.30 USD per 1,000 tickets on the default Groq model under our stated reasoning assumption, confirmed or corrected by one --real run of the lab.
Interview Questions
1. Why can the same prompt cost different amounts on two models with the same per-token price? Because each model family has its own tokenizer, and the same text becomes a different number of tokens. On Brightlane's triage request, four tokenizers gave 652 to 723 tokens, an 11% spread, and for Hindi text the spread was 1.6x to 7.7x the English count. Compare cost per task with real usage, not price per token.
2. What is BPE and why does it make ids and numbers expensive? Byte pair encoding starts from bytes and repeatedly merges the most frequent adjacent pair into a new token. Frequent words become single tokens; rare strings stay split. Invoice ids and long numbers are rare and unpredictable, so they break into many pieces: INV-2026-004512 is 7 tokens in o200k_base, about 2 characters per token versus 5 for prose.
3. Why do models struggle to count letters or add large numbers? They see token ids, not characters. strawberry arrives as st|raw|berry, and 12345 + 67890 as 123|45| +| |678|90, with digit groups that do not align with place value. The fix is to do character and arithmetic work in code and give the model results, then validate anything the model returns.
4. How do you estimate tokens before sending a request, and how accurate is it? Count each message's content with a tokenizer close to the model's, add a few tokens per message for the chat template, and add the priming tokens (count_messages does this). Accuracy depends on tokenizer match and hidden template text; calibrate with a few real calls by dividing reported prompt_tokens by your estimate, and use that ratio plus a margin.
5. Write the context budget equation and explain the output reserve. System + tools + history + retrieved + user + output reserve + safety margin must not exceed the window. The output reserve is max_tokens: the window covers input and output together, and reasoning models spend part of it on hidden thinking, so a reasoning model needs a reserve in the thousands, not the length of the visible answer.
6. What happens when a request exceeds the context window? It depends on the stack. Hosted APIs usually reject it with HTTP 400 (often code context_length_exceeded). Local servers such as Ollama may silently keep only the newest tokens, and TinyLM's generate() does the same, dropping the start of the prompt with no signal. If only the output runs out, you get finish_reason == "length" and a cut-off reply. Budget before sending, catch the error, trim, retry once, and check finish_reason.
7. What is "lost in the middle", and how does it change prompt layout? Liu et al. (2023, TACL 2024) showed that models use information at the beginning and end of a long context better than information in the middle; in the worst case GPT-3.5-Turbo did worse with the answer document in the middle than with no documents. Put key instructions first and restate the task last, place the most relevant retrieved passage next to the question, and test your own model with a needle harness at your lengths.
8. Why is advertised context length not the same as usable context? The advertised figure is what the API accepts. Benchmarks such as RULER found only about half of models claiming 32K+ performed satisfactorily at 32K, and NoLiMa found 11 of 13 models fell below half their short-context score at 32K once literal word matching was removed. Measure effective context for your task and send the smallest context that contains what is needed.
9. Compare drop-oldest, sliding window, importance-ranked truncation, and summary compaction. Drop-oldest and sliding windows are cheap and predictable but lose early facts: on our 23-turn chat both lost the invoice id, plan, and request. Pinning the first message kept two of three. Importance ranking kept all three with no extra call because the scorer knew the fact types. Summary compaction compresses most but can silently lose facts, so verify each summary with an automatic fact check and fall back when it fails.
10. Why is output usually the expensive half of a bill? Output tokens are generated one at a time, each needing a pass through the model, while input tokens are processed in parallel, so providers price output higher: 6x input on gemini-3.5-flash and 4x on gpt-oss-120b. A triage call that is 6% output tokens spent 29% of its cost on output; a draft reply at 17% output spent 56%. Reasoning tokens are billed as output too.
11. How do you lay out a prompt to benefit from prompt caching? Put everything stable first in a fixed, byte-identical order (instructions, tool definitions, examples, reference text) and everything variable last (ticket, ids, timestamps). On 24 triage requests that made 92% of tokens shareable, versus 1% when a single metadata line sat at the top. Then check the provider's minimum prefix: a 667-token prefix caches nothing under a 1,024 minimum, and Gemini 3.5 Flash needs 4,096. Confirm with reported cached_tokens.
12. What is cost per task and why is it the metric that matters? It is total spend divided by successful results, including retries, reasoning tokens, and failures, for example cost per valid triage label. Price per token ignores output length, hidden reasoning, cache hit rates, and retries; cost per task includes them all. Pair it with accuracy from an evaluation to get cost per correct result.
Other Tools and Providers
| What we used | Alternatives | Notes |
|---|---|---|
tiktoken with vendored o200k_base, cl100k_base, p50k_base | Provider count-token endpoints (Gemini countTokens, Anthropic token counting), Hugging Face tokenizers / transformers AutoTokenizer for open-weight models | Exact counts need the model's own tokenizer; endpoints cost a network call |
tokenizers BPE trainer | SentencePiece (BPE and unigram), tiktoken's educational BPE module | SentencePiece is common in Llama, Gemma, and T5 families |
count_messages estimate | Provider-reported usage, observability tools such as Langfuse, Helicone, or OpenTelemetry GenAI traces | Estimates for budgeting; reported usage for billing (Module 13) |
| Hand-written truncation and compaction | Framework memory utilities (LangChain, LlamaIndex chat memory), provider-side conversation state or compaction features | Frameworks save code; check what they drop and add a fact check |
| Our needle harness | RULER, NoLiMa, LongBench, Greg Kamradt's original needle-in-a-haystack test | Public suites for comparing models; your harness for your data |
pricing.py table | Provider pricing pages, LiteLLM's model cost map, OpenRouter's model pages | Keep one table in code and verify it on a schedule |
| Provider-side prompt caching | Explicit caches (Gemini cached contents, Anthropic cache_control), self-hosted prefix caching (vLLM automatic prefix caching, SGLang RadixAttention) | Self-hosting makes caching your job and your saving |
| Batch tiers (Groq, Gemini) | OpenAI Batch API, Anthropic Message Batches | Typically half price for asynchronous jobs |
Coming Up in Module 3
You now know what a request is made of and what it costs. Module 3, Inference Behaviour and Decoding Control, looks at how the response is produced: prefill and decode as separate phases (and why latency grows with output length, which is the other half of "output is expensive"), time to first token and streaming, what temperature, top-p, top-k, and min-p actually change in TinyLM's probability distributions, stop sequences and max_tokens, constrained decoding, and how reasoning models spend thinking tokens, with budgets and effort levels you can measure.