CourseLarge Language Models · Module 2 : Tokens -context-cost · part 8 of 80
Part 8 · Module 2 : Tokens -context-cost

Part A: Tokenization in practice

30 min read·22 Sept 2026

By the end of this module, you'll have:

  • Measured, with five real tokenizers, how many tokens Brightlane's tickets, invoice ids, JSON, code, and non-English messages cost, and why the same request can cost 1.6 to 7.7 times more in Hindi than in English.
  • Read supportdesk/tokens.py and supportdesk/pricing.py line by line, and used them to count a full triage request before sending it.
  • Built a context budget calculator for the Brightlane assistant, and watched what really happens at and past the limit, both in TinyLM and with a provider's overflow error.
  • A needle-in-a-haystack harness for position effects that runs against any model through llm.chat, and that you have proven can detect a blind spot.
  • Four context management strategies (drop-oldest, sliding window, importance-ranked, summary compaction) compared on a long support chat with an automatic check of which key facts survive.
  • A cost model for the whole triage workload: cost per task under every model in PRICES, output versus input share, reasoning tokens, batch pricing, and a measured cache-aware prompt layout.

Prerequisites: Module 1 (what a model is, the supportdesk repository, llm.chat, ScriptedLLM, TinyLM). Working Python. No API key is needed for any measured output in this module.

Where we are: Module 1 showed that a model reads and writes tokens, one prediction at a time, inside a fixed context window. This module turns those words into numbers you can budget: how many tokens your text becomes, how many fit, which ones the model actually uses, and what each one costs.

How this module is organized

PartWhat it covers
SetupWorking copy, how the examples run
Part A: Tokenization in practiceSubwords and BPE, tokenizer families, fertility on non-English text, code, numbers and ids, artifacts behind counting and arithmetic failures, counting before you send (tokens.py)
Part B: The context windowWhat fills it, the budget equation, TinyLM and providers at the limit, position effects, a needle-in-a-haystack harness, effective versus advertised context
Part C: Managing contextDrop-oldest, sliding window, importance-ranked truncation, summary compaction, what to preserve and when, long context versus retrieval
Part D: Cost mechanicspricing.py, input and output prices, reasoning tokens, prompt caching and cache-aware layout, batch tiers, cost per task
Module LabA token-aware, budget-checked, cost-reported triage run with a compaction guard

Setup

The examples run in order and each one is a complete script. They live in examples/ and import the canonical supportdesk package plus a few small Module 2 helpers (m02_prompts.py, m02_conversation.py, m02_context.py) that sit next to them. Run every command from the repository root.

bash
cp -r supportdesk ~/work/m02 && cd ~/work/m02
source ~/venv/bin/activate          # Python 3.11 with the packages from requirements.txt
pip install -r requirements.txt     # tiktoken==0.14.0, tokenizers==0.23.2, torch==2.14.0, openai==3.16.2, ...
PYTHONPATH=. python examples/m02_tokenizers.py
PYTHONPATH=. python -m pytest -q tests/test_m02_tokens_context_cost.py

Code explained

  • In simple words: make your own copy of the repository, activate the environment, and check that the first example and the module's tests run.
  • What happens: cp -r copies the canonical repository so nothing you do touches the original. PYTHONPATH=. lets Python find the supportdesk package from the repository root. Scripts in examples/ also find their Module 2 neighbors (m02_prompts.py and friends) because Python puts a script's own folder on the import path. The tokenizer files ship in vendor/tokenizers/, so nothing downloads.
  • Comes out: the first example prints token pieces (shown in Part A), and pytest ends with 9 passed. If you see ModuleNotFoundError: supportdesk, you forgot PYTHONPATH=. or ran from another folder.

Two files carry the triage request and the long chat that later examples reuse. The triage prompt is a working draft: Module 4 teaches how to design prompts properly; here it is simply a realistic payload to count and price.

python
# examples/m02_prompts.py
"""Module 2: the triage request used by this module's examples (Module 4 teaches prompt design properly)."""
import json

from supportdesk.data import CATEGORIES, PRIORITIES, Ticket
from supportdesk.schemas import triage_json_schema

TRIAGE_INSTRUCTIONS = f"""You are the triage assistant for Brightlane's support desk.
Brightlane is a project-management SaaS. Plans: Free, Team (12 USD per user per month),
Business (24 USD per user per month), Enterprise (custom pricing).

Read one customer ticket and classify it. Reply with one JSON object and nothing else.

Categories: {", ".join(CATEGORIES)}.
- billing: charges, invoices, prices, payment methods.
- cancellation: cancelling a plan or asking for a refund because they are leaving.
- account_access: login, password, SSO, locked accounts, 2FA.
- bug: something that used to work is broken or erroring.
- how_to: a question about using a feature that exists.
- feature_request: asking for something Brightlane does not do yet.

Priorities: {", ".join(PRIORITIES)}.
- urgent: many users blocked or a security risk right now.
- high: one user blocked, or money charged wrongly.
- normal: needs an answer, nobody is blocked.
- low: a question or an idea.

Write the summary in English even when the ticket is in another language.
Set needs_human to true for refunds, account changes, legal, or security issues.

JSON schema of the reply:
{json.dumps(triage_json_schema(), separators=(",", ":"))}"""

EXAMPLES = [
    ("Subject: Invoice address\n\nCan you add our VAT number to future invoices?",
     '{"category":"billing","priority":"low","language":"en","summary":"Customer wants a VAT number on future invoices.","needs_human":false}'),
    ("Subject: SSO broken\n\nNobody in our company can log in through Okta since this morning.",
     '{"category":"account_access","priority":"urgent","language":"en","summary":"Company-wide SSO login failure through Okta.","needs_human":true}'),
]


def triage_messages(ticket: Ticket) -> list[dict]:
    """Stable parts first (instructions, examples), the ticket last: a cache-friendly layout."""
    messages = [{"role": "system", "content": TRIAGE_INSTRUCTIONS}]
    for user, assistant in EXAMPLES:
        messages += [{"role": "user", "content": user}, {"role": "assistant", "content": assistant}]
    messages.append({"role": "user", "content": f"Customer tier: {ticket.customer_tier}\n{ticket.text}"})
    return messages

Code explained

  • In simple words: this is the request the Brightlane assistant sends to classify one ticket: fixed instructions, two worked examples, then the ticket.
  • What happens: TRIAGE_INSTRUCTIONS holds the plan prices, category and priority definitions, and the JSON schema from schemas.triage_json_schema() (introduced in Module 6; here we only count it). EXAMPLES are two question and answer pairs, a technique called few-shot prompting (Module 4). triage_messages() builds the chat messages with every stable part first and the ticket last. Part D shows, with measurements, why that order matters for cost.
  • Comes out: nothing when run alone; later scripts import it.
python
# examples/m02_conversation.py
"""Module 2: a long multi-turn Brightlane support chat and a checker for the facts that must survive."""
import re

SYSTEM = ("You are Brightlane's support assistant. Answer from the help center, be concise, "
          "and hand refunds and account changes to a human agent.")

TURNS = [
    ("user", "Hi, we were charged twice this month, invoice INV-2026-004512. Please refund the duplicate charge."),
    ("assistant", "Sorry about that. I can pass this to our billing team. Which plan is the workspace on, and how many seats?"),
    ("user", "We are on the Business plan with 40 seats, paid monthly."),
    ("assistant", "Thanks. A billing agent will review it; duplicate charges go back to the original card in 5 to 10 business days."),
    ("user", "While I have you: Slack notifications stopped arriving in our #ops channel yesterday."),
    ("assistant", "Please open Settings > Integrations > Slack and check that the channel is still connected. Reconnecting fixes most cases."),
    ("user", "It says connected, but nothing arrives. We did rename the channel last week."),
    ("assistant", "Renaming breaks the link. Choose the channel again in the Slack integration settings and send a test message."),
    ("user", "That worked, thank you. Another question: can I export a board to CSV with the custom fields?"),
    ("assistant", "Yes. Open the board, choose More > Export > CSV, and tick Include custom fields. Large boards arrive by email."),
    ("user", "The export link in the email says expired."),
    ("assistant", "Export links last 24 hours. Run the export again and download it the same day."),
    ("user", "OK. Our designers also want automations that move cards when a checklist is complete. Possible?"),
    ("assistant", "Yes: Automations > New rule > When checklist completed > Move card to column. Rules run within a minute."),
    ("user", "The rule ran twice on one card this morning."),
    ("assistant", "That can happen when two rules match the same card. Check for an older rule with the same trigger and disable one."),
    ("user", "Found it, there was an old copy. Does the mobile app support offline mode?"),
    ("assistant", "The mobile app caches boards you opened recently and syncs your edits when you reconnect."),
    ("user", "Great. Last thing on mobile: push notifications are delayed by an hour on Android."),
    ("assistant", "Android battery optimization can delay them. Exclude the Brightlane app from battery optimization in system settings."),
    ("user", "Done, they arrive now. Can guests see private boards?"),
    ("assistant", "No. Guests only see boards they are explicitly invited to, and private boards stay hidden from them."),
    ("user", "Good to know. So, where are we on my original request?"),
]

FACTS = {
    "invoice id": lambda text: "INV-2026-004512" in text,
    "plan and seats": lambda text: "Business" in text and re.search(r"\b40 seats\b", text) is not None,
    # "refund" and "duplicate" close together, so the system prompt's "refunds" alone does not count
    "customer's ask": lambda text: re.search(r"refund.{0,40}duplicate|duplicate.{0,40}refund", text, re.I) is not None,
}


def conversation() -> list[dict]:
    """System message plus the 23 turns above, as chat messages."""
    return [{"role": "system", "content": SYSTEM}] + [{"role": r, "content": c} for r, c in TURNS]


def facts_kept(messages: list[dict]) -> dict[str, bool]:
    """Which key facts are still present anywhere in the messages we would send."""
    text = "\n".join(m["content"] for m in messages)
    return {name: check(text) for name, check in FACTS.items()}

Code explained

  • In simple words: a 23-turn support chat in which the important facts come early and the question that needs them comes last, plus a checker that says whether those facts are still in whatever we send.
  • What happens: the customer gives an invoice id, the plan and seat count, and the request (refund the duplicate charge) in the first three turns. Then the chat wanders through Slack, exports, automations, and mobile, and finally asks "where are we on my original request?". FACTS defines three checks by exact text or regular expression; facts_kept() joins the message contents and runs each check. None of the later turns repeat those facts, so a strategy only passes if it really kept them.
  • Comes out: nothing when run alone. The full chat is 24 messages and 558 tokens (measured in Part C).

Part A: Tokenization in practice

What a token is

A token is the unit a model reads and writes: a chunk of text that has an integer id in the model's vocabulary (its fixed list of known chunks). A tokenizer is the program that turns text into token ids and back. Frequent words usually become one token; rare words, numbers, and ids break into several subword pieces. Every price, limit, and speed figure you will see is quoted in tokens, not characters or words.

Let's look at one real Brightlane ticket through five tokenizers: three OpenAI families that ship with tiktoken (p50k_base from the GPT-3 era, cl100k_base from GPT-3.5 and GPT-4, o200k_base from GPT-4o onward), an older Anthropic tokenizer (claude-legacy), and TinyLM's own tokenizer from Module 1.

python
# examples/m02_tokenizers.py
"""Module 2: see one Brightlane ticket through five tokenizers, then count the whole dataset."""
from tokenizers import Tokenizer

from supportdesk.data import load_tickets
from supportdesk.tinylm import MODEL_DIR
from supportdesk.tokens import ENCODINGS, count_tokens, pieces

tiny = Tokenizer.from_file(str(MODEL_DIR / "tokenizer.json"))  # TinyLM's own byte-level BPE


def tiny_pieces(text: str) -> list[str]:
    return [tiny.decode([i]) for i in tiny.encode(text).ids]


tickets = load_tickets()
first = tickets[0]
print(f"{first.id}: {first.body}\n")
for name in ENCODINGS:
    p = pieces(first.body, name)
    print(f"{name:13} {len(p):3} tokens  {'|'.join(p)}")
p = tiny_pieces(first.body)
print(f"{'tinylm':13} {len(p):3} tokens  {'|'.join(p)}")

print("\nAll 72 tickets (subject + body):")
chars = sum(len(t.text) for t in tickets)
words = sum(len(t.text.split()) for t in tickets)
print(f"{'characters':13} {chars:6}   words {words}")
for name in ENCODINGS:
    n = sum(count_tokens(t.text, name) for t in tickets)
    print(f"{name:13} {n:6} tokens   {chars / n:4.2f} chars/token   {n / words:4.2f} tokens/word")
n = sum(len(tiny.encode(t.text).ids) for t in tickets)
print(f"{'tinylm':13} {n:6} tokens   {chars / n:4.2f} chars/token   {n / words:4.2f} tokens/word")

Code explained

  • In simple words: we hold the same sentence up to five different "rulers" and see where each one draws its tick marks.
  • What happens: pieces() from supportdesk/tokens.py returns the text of each token. For TinyLM we load models/tinylm-base/tokenizer.json with the tokenizers library and decode each id alone. The second half sums characters, words, and tokens over all 72 tickets (subject plus body, as Ticket.text formats them).
  • Comes out:
text
  T-1001: Hi, my card was charged 288 USD twice on 3 September for the Team plan (invoice INV-2026-004512). Please refund the duplicate.

  p50k_base      33 tokens  Hi|,| my| card| was| charged| 288| USD| twice| on| 3| September| for| the| Team| plan| (|inv|oice| INV|-|20|26|-|00|45|12|).| Please| refund| the| duplicate|.
  cl100k_base    33 tokens  Hi|,| my| card| was| charged| |288| USD| twice| on| |3| September| for| the| Team| plan| (|invoice| INV|-|202|6|-|004|512|).| Please| refund| the| duplicate|.
  o200k_base     33 tokens  Hi|,| my| card| was| charged| |288| USD| twice| on| |3| September| for| the| Team| plan| (|invoice| INV|-|202|6|-|004|512|).| Please| refund| the| duplicate|.
  claude-legacy  33 tokens  Hi|,| my| card| was| charged| 288| USD| twice| on| 3| September| for| the| Team| plan| (|invoice| IN|V|-|20|26|-|00|45|12|).| Please| refund| the| duplicate|.
  tinylm         35 tokens  H|i|,| my| card| was| charged| 2|8|8| USD| twice| on| 3| S|ep|tember| for| the| Team| plan| (|in|voice| INV|-|2026|-|004512|).| Please| refund| the| duplicate|.

  All 72 tickets (subject + body):
  characters      7498   words 1266
  p50k_base       2004 tokens   3.74 chars/token   1.58 tokens/word
  cl100k_base     1818 tokens   4.12 chars/token   1.44 tokens/word
  o200k_base      1725 tokens   4.35 chars/token   1.36 tokens/word
  claude-legacy   1914 tokens   3.92 chars/token   1.51 tokens/word
  tinylm          3149 tokens   2.38 chars/token   2.49 tokens/word

Read the pieces first. Common words (charged, refund, duplicate) are single tokens in every tokenizer, and most tokens carry their leading space ( card, not card). The invoice id INV-2026-004512 falls apart into 5 to 9 pieces. TinyLM, trained only on Brightlane text, splits Hi into letters and September into three pieces, but has learned 2026 and even 004512 as single tokens because its corpus repeats them. On the whole dataset the newest OpenAI tokenizer needs 14% fewer tokens than the oldest (1,725 versus 2,004), and TinyLM's tiny vocabulary needs 83% more than o200k_base. The rule of thumb "about 4 characters per token for English" holds here (3.7 to 4.4), but only for English prose.

.

How subwords form: train a tiny BPE

Almost every modern tokenizer is built with byte pair encoding (BPE). Training starts from single bytes (256 symbols, enough to spell any text in any language) and repeatedly merges the most frequent adjacent pair into a new symbol, until the vocabulary reaches a target size. Frequent words end up as one token; rare strings stay split. Watching this happen explains most of the surprises in this part, so let's train a few BPE tokenizers on Brightlane's own corpus.

python
# examples/m02_train_bpe.py
"""Module 2: train byte-level BPE tokenizers of growing size and watch subword merges form."""
import json

from tokenizers import Tokenizer, decoders, models, pre_tokenizers, trainers

from supportdesk.data import DATA_DIR

corpus = str(DATA_DIR / "corpus.txt")
# "_" in the output marks a space that belongs to the token.
WORDS = [" refund", " duplicate", " Brightlane", " unsubscribed", " INV-2026-004512"]


def train(vocab_size: int) -> Tokenizer:
    tok = Tokenizer(models.BPE())
    tok.pre_tokenizer = pre_tokenizers.ByteLevel(add_prefix_space=False)
    tok.decoder = decoders.ByteLevel()
    trainer = trainers.BpeTrainer(vocab_size=vocab_size, show_progress=False,
                                  initial_alphabet=pre_tokenizers.ByteLevel.alphabet())
    tok.train([corpus], trainer)
    return tok


for size in (256, 280, 400, 1000, 4000):
    tok = train(size)
    shown = ["|".join(tok.decode([i]).replace(" ", "_") for i in tok.encode(w).ids) for w in WORDS]
    print(f"vocab {tok.get_vocab_size():5}: " + "   ".join(shown))

big = train(4000)
merges = json.loads(big.to_str())["model"]["merges"]
print(f"\nFirst 12 of {len(merges)} merges (_ marks a leading space):")
print("  " + ", ".join(" + ".join(m).replace("\u0120", "_") for m in merges[:12]))
print(f"Asked for 4000, got {big.get_vocab_size()}: training stops when no pair is left to merge.")

Code explained

  • In simple words: we build the same kind of tokenizer TinyLM uses, five times with bigger and bigger vocabularies, and watch words fuse from letters into whole tokens.
  • What happens: models.BPE() is the BPE model from the tokenizers library. pre_tokenizers.ByteLevel first splits text into words (keeping the leading space with the word) and maps bytes to printable symbols; BpeTrainer learns merges up to vocab_size. The initial_alphabet guarantees all 256 bytes are in the vocabulary, so no text is ever unrepresentable. We print five probe words at each size, then the first merges learned and the final size. This mirrors train_tokenizer() in scripts/pretrain_tinylm.py, which built TinyLM's tokenizer.
  • Comes out (about 4 seconds):
text
  vocab   256: _|r|e|f|u|n|d   _|d|u|p|l|i|c|a|t|e   _|B|r|i|g|h|t|l|a|n|e   _|u|n|s|u|b|s|c|r|i|b|e|d   _|I|N|V|-|2|0|2|6|-|0|0|4|5|1|2
  vocab   280: _|re|f|u|n|d   _|d|u|p|l|i|c|at|e   _|B|r|i|g|h|t|l|an|e   _|u|n|s|u|b|s|c|r|i|b|e|d   _|I|N|V|-|2|0|2|6|-|0|0|4|5|1|2
  vocab   400: _refund   _d|up|l|ic|at|e   _B|ri|g|h|t|lan|e   _|un|s|u|b|s|c|ri|b|ed   _I|N|V|-|2|0|2|6|-|0|0|4|5|1|2
  vocab  1000: _refund   _duplicate   _Brightlane   _un|s|ub|scri|b|ed   _I|N|V|-|20|26|-|00|4|5|1|2
  vocab  1502: _refund   _duplicate   _Brightlane   _un|s|ub|scri|b|ed   _INV|-|2026|-|004512

  First 12 of 1246 merges (_ marks a leading space):
    e + r, a + n, t + h, i + n, u + s, _ + (, ) + :, t + o, e + n, _ + th, _ + a, to + m
  Asked for 4000, got 1502: training stops when no pair is left to merge.

At 256 entries every word is spelled byte by byte. The first merges are the corpus's most frequent pairs: e + r, a + n, t + h, and, tellingly, _ + ( and ) + :, because the corpus is full of lines like Customer (Ben):. By 400 entries refund is one token; by 1,000 duplicate and Brightlane are too. unsubscribed never becomes one token: the word is rare in this corpus, so its pieces never win a merge. Finally, we asked for 4,000 entries and got 1,502. BPE stops when every word in the training text is already a single token, and this templated corpus has only about 1,250 distinct pairs worth merging. That is also why TinyLM's tokenizer has 1,503 entries (1,502 plus the <|endoftext|> special token) although its model reserves 2,048 embedding rows (vocab_size in models/tinylm-base/config.json); 545 rows are never used.

The lesson for everything that follows: what a tokenizer merges depends on what it was trained on. A tokenizer trained mostly on English web text learns English words, so other scripts and unusual strings stay expensive.

The canonical token counter: supportdesk/tokens.py

This course counts tokens through one small file. Module 2 introduces it; later modules import it.

<strong>supportdesk/tokens.py</strong>

python
# supportdesk/tokens.py
"""Count tokens offline with real tokenizers from several model families.

Encodings available without network access:
  p50k_base    GPT-3 era (about 50k vocabulary)
  cl100k_base  GPT-3.5 / GPT-4 era (about 100k vocabulary)
  o200k_base   GPT-4o and later OpenAI models (about 200k vocabulary)
  claude-legacy  an older Anthropic tokenizer (about 65k vocabulary)
Other providers (Llama, Gemini, Qwen) use their own tokenizers; counts differ
by model, so treat any single count as an estimate for a model it was not built for.
"""
from __future__ import annotations

import os
from functools import lru_cache
from pathlib import Path

VENDOR = Path(__file__).resolve().parents[1] / "vendor" / "tokenizers"
os.environ.setdefault("TIKTOKEN_CACHE_DIR", str(VENDOR))

import tiktoken  # noqa: E402  (must import after the cache directory is set)
from tokenizers import Tokenizer  # noqa: E402

ENCODINGS = ("p50k_base", "cl100k_base", "o200k_base", "claude-legacy")
DEFAULT_ENCODING = "o200k_base"


@lru_cache(maxsize=None)
def _encoder(name: str):
    if name == "claude-legacy":
        return Tokenizer.from_file(str(VENDOR / "anthropic_tokenizer.json"))
    if name not in ENCODINGS:
        raise ValueError(f"Unknown encoding {name!r}. Choose one of {ENCODINGS}.")
    return tiktoken.get_encoding(name)


def encode(text: str, encoding: str = DEFAULT_ENCODING) -> list[int]:
    enc = _encoder(encoding)
    if encoding == "claude-legacy":
        return enc.encode(text).ids
    return enc.encode(text, disallowed_special=())


def count_tokens(text: str, encoding: str = DEFAULT_ENCODING) -> int:
    """Number of tokens `text` becomes under one tokenizer."""
    return len(encode(text, encoding))


def pieces(text: str, encoding: str = DEFAULT_ENCODING) -> list[str]:
    """The text each token stands for, useful for seeing where a tokenizer splits."""
    enc = _encoder(encoding)
    if encoding == "claude-legacy":
        return [enc.decode([i]) for i in enc.encode(text).ids]
    return [enc.decode_single_token_bytes(i).decode("utf-8", errors="replace") for i in encode(text, encoding)]


def count_messages(messages: list[dict], encoding: str = DEFAULT_ENCODING, per_message_overhead: int = 4) -> int:
    """Estimate prompt tokens for a chat request.

    Each message costs its content tokens plus a few tokens of role and
    separator markup; 4 is a common approximation. Providers report the exact
    number in the response usage, so use this for budgeting before you send.
    """
    total = 0
    for m in messages:
        content = m.get("content") or ""
        if not isinstance(content, str):
            content = str(content)
        total += count_tokens(content, encoding) + per_message_overhead
    return total + 3  # the assistant turn the model is primed to write

Code explained

  • In simple words: one place to count tokens offline with several real tokenizers, so every budget and cost estimate in the course uses the same ruler.
  • What happens (per part):
  • Module docstring and VENDOR: the tokenizer files live in vendor/tokenizers/. Setting TIKTOKEN_CACHE_DIR before import tiktoken makes tiktoken load them from disk instead of downloading, which is why the import sits below the environment line.
  • ENCODINGS and DEFAULT_ENCODING: the four names you can pass. o200k_base is the default because the course's default model, openai/gpt-oss-120b, uses o200k_harmony, a superset of o200k_base with extra special tokens for its chat format, so content counts match closely.
  • _encoder(name): builds each tokenizer once (lru_cache remembers it) and returns either a tiktoken Encoding or a tokenizers.Tokenizer for claude-legacy. An unknown name raises ValueError listing the valid ones.
  • encode(text, encoding): returns token ids. disallowed_special=() tells tiktoken to treat text like <|endoftext|> inside a ticket as ordinary characters instead of raising an error, which matters when you count untrusted customer text.
  • count_tokens(text, encoding): the length of encode(). This is the function you will call most.
  • pieces(text, encoding): decodes each id on its own so you can see split points. A single token can be part of a multi-byte character (common in Japanese and Hindi); errors="replace" shows such fragments as the replacement character instead of crashing.
  • count_messages(messages, encoding, per_message_overhead): estimates a chat request: content tokens plus about 4 tokens per message for role markers and separators, plus 3 for the assistant turn the model is primed to write. Non-string content (for example a list of content parts) is converted with str(), which is only a rough estimate.
  • Comes out: nothing by itself. The docstring's warning is the important output: counts from one tokenizer are estimates for any model built with another.

Tokenizer differences across model families

Each model family trains its own tokenizer, so the same text has different token counts, and therefore different prices and limits, depending on the model. The table below uses the measured totals from the first example.

TokenizerVocabulary sizeUsed byTokens for all 72 tickets
p50k_base50,281GPT-3 era OpenAI models2,004
cl100k_base100,277GPT-3.5, GPT-41,818
o200k_base200,019GPT-4o and later; gpt-oss via o200k_harmony1,725
claude-legacy65,000older Anthropic models1,914
TinyLM BPE1,503TinyLM only3,149

Vocabulary sizes were read from the loaded tokenizers (tiktoken.get_encoding(name).n_vocab, Tokenizer.get_vocab_size()). Two practical consequences:

  • Bigger vocabularies usually mean fewer tokens, especially outside English, because more words and scripts earned their own merges. They also mean a bigger embedding table in the model; that trade-off is the model builder's, not yours.
  • You cannot count exactly for a model whose tokenizer you do not have. Gemini, Llama, and Qwen use their own tokenizers, and current Claude models do not use claude-legacy. Providers offer exact counting (Gemini's countTokens, Anthropic's token counting endpoint) and always report exact usage after a call. Use local counts for budgeting, with a safety margin, and calibrate against reported usage (later in this part).

Fertility: non-English text

Fertility is how many tokens a tokenizer needs per unit of text: per word, per character, or, most usefully for cost, relative to the same meaning in English. High fertility means a message costs more, fills the context faster, and is generated more slowly. The ticket dataset has only 8 non-English tickets, so we also measure a parallel set: the same three support requests written in each language.

python
# examples/m02_fertility.py
"""Module 2: how many tokens the same support request costs in five languages."""
from tokenizers import Tokenizer

from supportdesk.data import load_tickets
from supportdesk.tinylm import MODEL_DIR
from supportdesk.tokens import ENCODINGS, count_tokens

tiny = Tokenizer.from_file(str(MODEL_DIR / "tokenizer.json"))
NAMES = ENCODINGS + ("tinylm",)


def count(text: str, name: str) -> int:
    return len(tiny.encode(text).ids) if name == "tinylm" else count_tokens(text, name)


# The same three requests, written in each language (hand translations for this course).
PARALLEL = {
    "en": ["Hi, I was charged twice for the Team plan this month. Please refund the duplicate charge.",
           "I forgot my password and the reset link has expired. What can I do?",
           "How do I export a board as a CSV file?"],
    "es": ["Hola, me cobraron dos veces el plan Team este mes. Por favor, reembolsen el cargo duplicado.",
           "Olvidé mi contraseña y el enlace para restablecerla ha caducado. ¿Qué puedo hacer?",
           "¿Cómo exporto un tablero como archivo CSV?"],
    "de": ["Hallo, mir wurde der Team-Plan diesen Monat zweimal berechnet. Bitte erstatten Sie die doppelte Abbuchung.",
           "Ich habe mein Passwort vergessen und der Link zum Zurücksetzen ist abgelaufen. Was kann ich tun?",
           "Wie exportiere ich ein Board als CSV-Datei?"],
    "ja": ["こんにちは。今月、Teamプランの料金が二重に請求されました。重複した請求分を返金してください。",
           "パスワードを忘れてしまい、リセット用のリンクの有効期限が切れています。どうすればよいですか?",
           "ボードをCSVファイルとしてエクスポートするにはどうすればよいですか?"],
    "hi": ["नमस्ते, इस महीने Team प्लान के लिए मुझसे दो बार शुल्क लिया गया। कृपया दोहरा शुल्क वापस करें।",
           "मैं अपना पासवर्ड भूल गया हूँ और रीसेट लिंक की समय सीमा समाप्त हो गई है। मैं क्या करूँ?",
           "मैं किसी बोर्ड को CSV फ़ाइल के रूप में कैसे एक्सपोर्ट करूँ?"],
}

print("Parallel set: total tokens for the same 3 requests, and the ratio to English")
print(f"{'lang':5}{'chars':>6}{'bytes':>6}" + "".join(f"{n:>15}" for n in NAMES))
english = {n: sum(count(s, n) for s in PARALLEL["en"]) for n in NAMES}
for lang, texts in PARALLEL.items():
    chars = sum(len(s) for s in texts)
    nbytes = sum(len(s.encode("utf-8")) for s in texts)
    cells = []
    for n in NAMES:
        total = sum(count(s, n) for s in texts)
        cells.append(f"{total:>6} ({total / english[n]:.1f}x)")
    print(f"{lang:5}{chars:>6}{nbytes:>6}" + "".join(f"{c:>15}" for c in cells))

print("\nReal tickets: tokens per character by language (o200k_base and tinylm)")
tickets = load_tickets()
for lang in ("en", "es", "de", "ja", "hi"):
    group = [t for t in tickets if t.language == lang]
    chars = sum(len(t.text) for t in group)
    o200 = sum(count(t.text, "o200k_base") for t in group)
    tl = sum(count(t.text, "tinylm") for t in group)
    words = sum(len(t.text.split()) for t in group)
    per_word = "no spaces" if lang == "ja" else f"{o200 / words:.2f}/word"
    print(f"{lang}: n={len(group):2}  chars={chars:5}  o200k={o200:5} ({o200 / chars:.2f}/char, "
          f"{per_word})  tinylm={tl:5} ({tl / chars:.2f}/char)")

Code explained

  • In simple words: we say the same three things in five languages and ask each tokenizer what it charges.
  • What happens: PARALLEL holds hand translations of a duplicate-charge complaint, a password reset question, and a CSV export question. For each language and tokenizer we total the tokens and divide by the English total. We also print characters and UTF-8 bytes: Latin letters are 1 byte, accented letters 2, Japanese and Devanagari 3. The second table measures the real tickets by language, per character and per word (Japanese is written without spaces, so "per word" does not apply).
  • Comes out:
text
  Parallel set: total tokens for the same 3 requests, and the ratio to English
  lang  chars bytes      p50k_base    cl100k_base     o200k_base  claude-legacy         tinylm
  en      194   194      46 (1.0x)      46 (1.0x)      46 (1.0x)      46 (1.0x)      52 (1.0x)
  es      216   222      78 (1.7x)      61 (1.3x)      56 (1.2x)      66 (1.4x)     130 (2.5x)
  de      245   246      89 (1.9x)      67 (1.5x)      59 (1.3x)      74 (1.6x)     138 (2.7x)
  ja      129   373     163 (3.5x)     113 (2.5x)      84 (1.8x)     115 (2.5x)     369 (7.1x)
  hi      237   599     353 (7.7x)     233 (5.1x)      73 (1.6x)     250 (5.4x)    592 (11.4x)

  Real tickets: tokens per character by language (o200k_base and tinylm)
  en: n=64  chars= 6767  o200k= 1508 (0.22/char, 1.29/word)  tinylm= 2442 (0.36/char)
  es: n= 2  chars=  250  o200k=   67 (0.27/char, 1.81/word)  tinylm=  147 (0.59/char)
  de: n= 3  chars=  306  o200k=   71 (0.23/char, 1.65/word)  tinylm=  162 (0.53/char)
  ja: n= 2  chars=  104  o200k=   61 (0.59/char, no spaces)  tinylm=  250 (2.40/char)
  hi: n= 1  chars=   71  o200k=   18 (0.25/char, 1.64/word)  tinylm=  148 (2.08/char)

The same meaning costs 1.2 to 1.9 times more tokens in Spanish and German, 1.8 to 3.5 times more in Japanese, and 1.6 to 7.7 times more in Hindi, depending on the tokenizer. The newest tokenizer (o200k_base) narrows the gap dramatically for Hindi (7.7x down to 1.6x) because its larger vocabulary includes many Devanagari merges; older ones fall back to several byte-level tokens per character. TinyLM, which saw almost no non-English text, needs 11.4 times more tokens for Hindi: it is spelling UTF-8 bytes one or two at a time. The real-ticket table agrees in direction, but with n = 1 to 3 tickets per non-English language it is anecdote, not measurement; the parallel set is the fairer comparison, and it is still only three sentences.

What this means for Brightlane: a Japanese customer's ticket costs about 1.8 times as much to process as the English equivalent on a current tokenizer, fills the context window 1.8 times faster, and a reply in Japanese takes about 1.8 times as many decode steps. If your traffic is multilingual, measure fertility on your own messages with your model's tokenizer before you set budgets.

Fertility: code, numbers, ids, and JSON

Support tickets are not just prose. They contain invoice ids, amounts, URLs, pasted code, and the assistant itself produces JSON.

python
# examples/m02_hard_strings.py
"""Module 2: code, numbers, ids, and JSON under five tokenizers, plus the pieces behind classic failures."""
import json

from tokenizers import Tokenizer

from supportdesk.schemas import Triage
from supportdesk.tinylm import MODEL_DIR
from supportdesk.tokens import ENCODINGS, count_tokens, encode, pieces

tiny = Tokenizer.from_file(str(MODEL_DIR / "tokenizer.json"))
NAMES = ENCODINGS + ("tinylm",)


def count(text: str, name: str) -> int:
    return len(tiny.encode(text).ids) if name == "tinylm" else count_tokens(text, name)


triage = Triage(category="billing", priority="high", language="en",
                summary="Customer was charged twice and wants the duplicate refunded.", needs_human=True)
SAMPLES = {
    "prose": "Please refund the duplicate charge on our Team plan.",
    "invoice ids": "INV-2026-004512, INV-2026-004871, INV-2025-019934",
    "amounts": "288.00 USD, 1,234.56 EUR, 0.0075 USD per token",
    "long number": "Workspace 81736450921 has 3141592653 events.",
    "python code": "if ticket.priority == 'urgent':\n    notify(oncall, ticket.id)\n",
    "json (pretty)": triage.model_dump_json(indent=2),
    "json (compact)": json.dumps(triage.model_dump(), separators=(",", ":")),
    "url + email": "https://status.brightlane.example/incidents?id=4411 billing@brightlane.example",
}
print(f"{'sample':15}{'chars':>6}" + "".join(f"{n:>14}" for n in NAMES) + "   chars/token (o200k)")
for label, text in SAMPLES.items():
    counts = [count(text, n) for n in NAMES]
    print(f"{label:15}{len(text):>6}" + "".join(f"{c:>14}" for c in counts) + f"   {len(text) / counts[2]:.2f}")

print("\nWhere the pieces fall (o200k_base):")
for text in ["INV-2026-004512", "3141592653", "1,234.56", "    notify(oncall, ticket.id)"]:
    print(f"  {text!r:34} -> {pieces(text)}")

print("\nWhy letter counting, spelling, and arithmetic are hard (o200k_base):")
for text in ["Brightlane", "strawberry", " refund", " unsubscribed", "288 / 2 = 144", "14 * 24 = 336", "12345 + 67890"]:
    print(f"  {text!r:18} -> {pieces(text)}")
print("  'r' in 'strawberry':", "strawberry".count("r"), "(the model never sees letters, only the pieces above)")

print("\nSame word, different tokens (o200k_base ids):")
for text in ["refund", " refund", " Refund", " REFUND"]:
    print(f"  {text!r:10} -> ids {encode(text)}  pieces {pieces(text)}")

Code explained

  • In simple words: we measure the strings that are not ordinary words and look at exactly where they break.
  • What happens: SAMPLES holds realistic Brightlane strings, including a real Triage object from supportdesk/schemas.py serialized two ways: pretty (indented) and compact (no spaces). The table prints token counts per tokenizer and characters per token for o200k_base (higher is cheaper). Then pieces() shows the split points for ids, numbers, code, and the classic failure cases, and encode() shows that one word gets different ids depending on spacing and case.
  • Comes out:
text
  sample          chars     p50k_base   cl100k_base    o200k_base claude-legacy        tinylm   chars/token (o200k)
  prose              52            10            10            10            10            11   5.20
  invoice ids        49            27            23            23            28            29   2.13
  amounts            46            19            21            21            19            27   2.19
  long number        44            14            14            14            13            26   3.14
  python code        62            21            15            15            19            42   4.13
  json (pretty)     169            53            45            46            47            85   3.67
  json (compact)    148            33            30            31            33            74   4.77
  url + email        78            22            18            18            22            29   4.33

  Where the pieces fall (o200k_base):
    'INV-2026-004512'                  -> ['INV', '-', '202', '6', '-', '004', '512']
    '3141592653'                       -> ['314', '159', '265', '3']
    '1,234.56'                         -> ['1', ',', '234', '.', '56']
    '    notify(oncall, ticket.id)'    -> ['   ', ' notify', '(on', 'call', ',', ' ticket', '.id', ')']

  Why letter counting, spelling, and arithmetic are hard (o200k_base):
    'Brightlane'       -> ['Bright', 'lane']
    'strawberry'       -> ['st', 'raw', 'berry']
    ' refund'          -> [' refund']
    ' unsubscribed'    -> [' unsub', 'scribed']
    '288 / 2 = 144'    -> ['288', ' /', ' ', '2', ' =', ' ', '144']
    '14 * 24 = 336'    -> ['14', ' *', ' ', '24', ' =', ' ', '336']
    '12345 + 67890'    -> ['123', '45', ' +', ' ', '678', '90']
    'r' in 'strawberry': 3 (the model never sees letters, only the pieces above)

  Same word, different tokens (o200k_base ids):
    'refund'   -> ids [148482]  pieces ['refund']
    ' refund'  -> ids [18376]  pieces [' refund']
    ' Refund'  -> ids [100598]  pieces [' Refund']
    ' REFUND'  -> ids [76596, 21592]  pieces [' REF', 'UND']

Prose runs at 5.2 characters per token; invoice ids and amounts at about 2.1. Ids are expensive because they are unpredictable: INV-2026-004512 becomes INV, -, 202, 6, -, 004, 512. The cl100k_base and o200k_base tokenizers split digits into groups of at most three, from the left, so 3141592653 becomes 314|159|265|3 and the groups do not line up with thousands. Pretty-printed JSON costs 46 tokens where compact JSON costs 31 for the same data: indentation and spaces are tokens too, a 33% saving on every structured reply if you ask for compact output. Python code is fairly efficient in the newer tokenizers (runs of spaces merge into one token) and poor in p50k_base.

Tokenization artifacts: counting, spelling, and arithmetic

The last part of that output explains several failures from Module 1's list of things models are unreliable at.

  • Counting letters. The model receives st|raw|berry, three ids. It never sees the letters r, r, r. To count them it must have learned, from training text, which letters each token contains. It often has, but not reliably, which is why "how many r's in strawberry" became a famous failure.
  • Spelling and character edits. refund, Refund, and refund are three unrelated ids (18376, 100598, 148482), and REFUND is two tokens. Tasks like "reverse this invoice id", "is this id's check digit correct", or "fix the typo in this SKU" operate on characters the model cannot see directly.
  • Arithmetic. 12345 + 67890 arrives as 123|45| +| |678|90. Column addition needs digits aligned by place value, but the pieces are aligned by the tokenizer's left-to-right three-digit grouping. The model has to reconstruct place values from chunk boundaries that mean nothing numerically.

The engineering response is the same in each case: do not ask the model to do character-level or arithmetic work you can do in code. For Brightlane that means computing refund amounts, seat totals, and prorations in Python, extracting invoice ids with a regular expression, and validating them in code (Module 6), then giving the model the results.

SituationUse thisWhy
Count, reverse, or validate characters in an idPython string operations or a regexThe model sees tokens, not characters
Compute a refund or a seat totalPython arithmetic, then pass the result to the modelDigit chunks do not align with place value
Ask the model to return numbers or idsCopy them from input, then check them in codeCopying is reliable, transformation is not
Structured repliesCompact JSONAbout a third fewer tokens than indented JSON here

Counting before you send

Before any call, you want to know: will this request fit, and what will it cost? count_messages() gives the estimate. Here it is on a full triage request.

python
# examples/m02_count_request.py
"""Module 2: count a full triage request before sending it, message by message and across tokenizers."""
from m02_prompts import triage_messages

from supportdesk.data import load_tickets
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import ENCODINGS, count_messages, count_tokens

ticket = load_tickets("test")[0]
messages = triage_messages(ticket)

print(f"Triage request for {ticket.id} ({len(messages)} messages), o200k_base:")
for m in messages:
    first_line = m["content"].splitlines()[0][:48]
    print(f"  {m['role']:9} {count_tokens(m['content']):4} tokens  {first_line!r}")
print(f"  count_messages total: {count_messages(messages)} (content + 4 per message + 3 priming)")

print("\nSame request under each tokenizer:")
for name in ENCODINGS:
    print(f"  {name:13} {count_messages(messages, encoding=name):5}")

reply = '{"category":"cancellation","priority":"normal","language":"en","summary":"Customer wants to cancel their plan.","needs_human":true}'
fake = ScriptedLLM(replies=[reply])  # plumbing only: not a model, and its usage is itself an estimate
result = fake(messages, max_tokens=200)
print(f"\nScriptedLLM usage (estimated with count_messages, not reported by a provider): {result.usage}")

Code explained

  • In simple words: we weigh the request message by message before it leaves the building.
  • What happens: we build the triage request for the first test ticket, count each message's content, and call count_messages() for the total. Then we count the same request with every tokenizer. Finally we pass it through ScriptedLLM, which is not a model: it returns our scripted JSON and fills usage using the same count_messages() estimate, so its numbers prove only that the plumbing carries usage through.
  • Comes out:
text
  Triage request for T-1003 (6 messages), o200k_base:
    system     512 tokens  "You are the triage assistant for Brightlane's su"
    user        15 tokens  'Subject: Invoice address'
    assistant   30 tokens  '{"category":"billing","priority":"low","language'
    user        20 tokens  'Subject: SSO broken'
    assistant   32 tokens  '{"category":"account_access","priority":"urgent"'
    user        28 tokens  'Customer tier: team'
    count_messages total: 664 (content + 4 per message + 3 priming)

  Same request under each tokenizer:
    p50k_base       723
    cl100k_base     652
    o200k_base      664
    claude-legacy   713

  ScriptedLLM usage (estimated with count_messages, not reported by a provider): Usage(input_tokens=664, output_tokens=29, cached_tokens=0, reasoning_tokens=0)

The system prompt is 77% of the request (512 of 664 tokens), and it is identical for every ticket; the ticket itself is 28 tokens. Keep that ratio in mind: it drives the caching result in Part D. The four tokenizers disagree by up to 11% (652 to 723) on the same request, which is the size of error to expect when you count with the wrong tokenizer. The 29 output tokens are for the compact JSON label.

A provider's reported usage is the ground truth, and it is usually a little higher than a content-only estimate because the chat template adds tokens you never wrote (role headers, a default system preamble for some models, tool definitions rendered as text). Calibrate once per model with a few real calls:

python
# examples/m02_calibrate.py
"""Module 2: compare our pre-send estimate with the usage a provider reports.

PYTHONPATH=. python examples/m02_calibrate.py --real   (needs a provider key)
"""
import argparse

from m02_prompts import triage_messages

from supportdesk.data import load_tickets
from supportdesk.llm import chat
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages

parser = argparse.ArgumentParser()
parser.add_argument("--real", action="store_true")
args = parser.parse_args()
llm = chat if args.real else ScriptedLLM(responder=lambda messages, kwargs: '{"category":"billing"}')

ratios = []
for ticket in load_tickets("test")[:8]:
    messages = triage_messages(ticket)
    estimate = count_messages(messages)
    result = llm(messages, max_tokens=512)
    ratios.append(result.usage.input_tokens / estimate)
    print(f"{ticket.id} {ticket.language}: estimated {estimate:4}  reported {result.usage.input_tokens:4}  "
          f"output {result.usage.output_tokens:4} (reasoning {result.usage.reasoning_tokens})")
print(f"reported / estimated: min {min(ratios):.3f}  max {max(ratios):.3f}  "
      f"({'real provider' if args.real else 'stand-in: it reports our own estimate, so 1.000 by construction'})")

Code explained

  • In simple words: send eight real requests, compare what we predicted with what the provider billed, and learn our correction factor.
  • What happens: for eight test tickets, estimate with count_messages(), call the model, and divide result.usage.input_tokens (read by llm.py from the provider's usage.prompt_tokens) by the estimate. Without --real it uses ScriptedLLM, whose "reported" usage is our own estimate.
  • Comes out (stand-in, captured):
text
  T-1003 en: estimated  664  reported  664  output    5 (reasoning 0)
  T-1006 en: estimated  668  reported  668  output    5 (reasoning 0)
  T-1009 en: estimated  662  reported  662  output    5 (reasoning 0)
  T-1012 en: estimated  660  reported  660  output    5 (reasoning 0)
  T-1015 en: estimated  663  reported  663  output    5 (reasoning 0)
  T-1018 en: estimated  662  reported  662  output    5 (reasoning 0)
  T-1021 en: estimated  669  reported  669  output    5 (reasoning 0)
  T-1024 en: estimated  662  reported  662  output    5 (reasoning 0)
  reported / estimated: min 1.000  max 1.000  (stand-in: it reports our own estimate, so 1.000 by construction)

With --real, you get the provider's numbers. Illustrative sample run (not captured in this build; produced for teaching). Your output will differ:

text
  T-1003 en: estimated  664  reported  741  output  212 (reasoning 171)
  T-1006 en: estimated  668  reported  745  output  198 (reasoning 158)
  ...
  reported / estimated: min 1.112  max 1.118

The shape is what to look for: a stable ratio a bit above 1 (template overhead), and output far larger than the 30-token label because a reasoning model spends hidden reasoning tokens first. Use your measured ratio as the safety margin in your budget.