Part F: Capabilities and Limits, Measured
What models are reliably good at
Across providers and sizes, instruction-tuned models are reliably good at tasks where the answer is mostly in the prompt or is a common pattern in text:
- Rewriting, summarizing, and changing tone of text you provide.
- Classifying text into categories you define, and extracting fields from it (with validation, Module 6).
- Translating between widely written languages.
- Writing common code and explaining code you paste in.
- Drafting replies grounded in documents you retrieve and include (Module 7).
They are unreliable at tasks that need exact symbol manipulation, precise recall of rare facts, facts newer than their training data, or accurate reports about themselves. The next sections measure each of these with tools you can run.
Tokens hide letters and digits
The model never sees characters. It sees token ids. Counting letters or doing long arithmetic means reasoning about pieces it cannot look inside.
"""Module 1: what the model actually sees. Why counting letters and long arithmetic are hard."""
from supportdesk.tinylm import load
from supportdesk.tokens import ENCODINGS, encode, pieces
_, tiny_tok = load()
words = ["strawberry", "Brightlane", "INV-2026-004512", "288 x 14 = 4032", "12345678901234"]
for text in words:
print(f"\n{text!r} (letters r: {text.count('r')}, characters: {len(text)})")
for enc in ENCODINGS:
print(f" {enc:14} {len(encode(text, enc)):2} tokens {pieces(text, enc)}")
tiny = tiny_tok.encode(text)
print(f" {'TinyLM':14} {len(tiny.ids):2} tokens {[tiny_tok.decode([i]) for i in tiny.ids]}")
print("\nWhat o200k_base hands the model for 'strawberry':", encode("strawberry", "o200k_base"))Code explained
- In simple words: show how five strings are cut into tokens by four real production tokenizers and by TinyLM's, so you can see what the model actually receives.
- What happens:
tokens.pieces(text, encoding)returns the text of each token.ENCODINGSholds four tokenizers stored offline invendor/tokenizers:p50k_base(GPT-3 era),cl100k_base(GPT-3.5 and GPT-4),o200k_base(GPT-4o and later OpenAI models, and the base of gpt-oss's tokenizer), and an older Anthropic tokenizer.tokens.pyis taught in full in Module 2; here we only borrowpiecesandencode. TinyLM's tokenizer is used directly through thetokenizerslibrary. - Comes out: real output.
'strawberry' (letters r: 3, characters: 10)
p50k_base 3 tokens ['st', 'raw', 'berry']
cl100k_base 3 tokens ['str', 'aw', 'berry']
o200k_base 3 tokens ['st', 'raw', 'berry']
claude-legacy 3 tokens ['st', 'raw', 'berry']
TinyLM 6 tokens ['st', 'r', 'a', 'w', 'ber', 'ry']
'Brightlane' (letters r: 1, characters: 10)
p50k_base 2 tokens ['Bright', 'lane']
cl100k_base 2 tokens ['Bright', 'lane']
o200k_base 2 tokens ['Bright', 'lane']
claude-legacy 2 tokens ['Bright', 'lane']
TinyLM 1 tokens ['Brightlane']
'INV-2026-004512' (letters r: 0, characters: 15)
p50k_base 9 tokens ['IN', 'V', '-', '20', '26', '-', '00', '45', '12']
cl100k_base 7 tokens ['INV', '-', '202', '6', '-', '004', '512']
o200k_base 7 tokens ['INV', '-', '202', '6', '-', '004', '512']
claude-legacy 8 tokens ['INV', '-', '20', '26', '-', '00', '45', '12']
TinyLM 6 tokens ['I', 'NV', '-', '2026', '-', '004512']
'288 x 14 = 4032' (letters r: 0, characters: 15)
p50k_base 6 tokens ['288', ' x', ' 14', ' =', ' 40', '32']
cl100k_base 8 tokens ['288', ' x', ' ', '14', ' =', ' ', '403', '2']
o200k_base 8 tokens ['288', ' x', ' ', '14', ' =', ' ', '403', '2']
claude-legacy 6 tokens ['288', ' x', ' 14', ' =', ' 40', '32']
TinyLM 13 tokens ['2', '8', '8', ' ', 'x', ' 14', ' ', '=', ' ', '4', '0', '3', '2']
'12345678901234' (letters r: 0, characters: 14)
p50k_base 6 tokens ['123', '45', '67', '89', '01', '234']
cl100k_base 5 tokens ['123', '456', '789', '012', '34']
o200k_base 5 tokens ['123', '456', '789', '012', '34']
claude-legacy 3 tokens ['123456789', '01', '234']
TinyLM 11 tokens ['12', '3', '45', '6', '7', '8', '9', '0', '12', '3', '4']
What o200k_base hands the model for 'strawberry': [302, 1618, 19772]Now the famous failures make sense:
- "How many r's in strawberry?" Production tokenizers hand the model
st,raw,berry(ids[302, 1618, 19772]under o200k_base). The answer depends on knowing the spelling of each token, which the model only learned indirectly from text that happens to spell things out. Models often get it right now because the question became famous, but the same failure appears on less famous words. - Arithmetic:
4032becomes403+2in o200k_base but40+32in p50k_base. A 14-digit number becomes 3 to 6 chunks depending on the tokenizer, and the chunks do not line up with place value. Adding two numbers means aligning digits the model cannot see individually. - Identifiers: the invoice number
INV-2026-004512is 6 to 9 tokens depending on the tokenizer. Copying it exactly is usually fine; transforming it (reversing, incrementing) is error-prone.
The fix is architectural, not a better prompt: let code do exact work. For Brightlane, the assistant should look up invoice ids and compute refund amounts with tools (Module 6), not in its head. Reasoning models are much better at arithmetic because they write out intermediate steps, but "much better" is not "exact"; verify.
Sensitivity to phrasing
A human support agent gives the same answer to "I'm locked out" and "my account is locked". Does the model? Measure it on TinyLM by asking the same question five ways and reading the probability of the correct reply's first token.
"""Module 1: sensitivity to phrasing. Same question, different wording, measured on TinyLM."""
import torch
from supportdesk.tinylm import load
torch.set_num_threads(1) # one thread is plenty for a 1M-parameter model
model, tokenizer = load()
@torch.no_grad()
def p_token(prompt: str, token_text: str) -> tuple[float, str]:
"""Probability of one specific next token, plus the model's actual top choice."""
target = tokenizer.encode(token_text).ids
assert len(target) == 1, f"{token_text!r} is not a single token"
probs = torch.softmax(model(torch.tensor([tokenizer.encode(prompt).ids]))[0, -1], dim=-1)
return probs[target[0]].item(), tokenizer.decode([int(probs.argmax())])
INTENTS = {
# correct first token of the right reply : paraphrases (the first matches the training template)
" After": ["Hello, my account is locked after too many attempts.",
"Hello, I'm locked out after typing my password wrong a few times.",
"Hello, it says my account is locked. Why?",
"HELLO, MY ACCOUNT IS LOCKED AFTER TOO MANY ATTEMPTS.",
"account locked too many attempts"],
" Sorry": ["Hi, I was charged twice for the Team plan this month. Can you refund the duplicate?",
"Hi, you billed my card two times this month for Team, please refund one.",
"Hi, there is a duplicate charge on my card.",
"Hi, can you refund the duplicate? I was charged twice for the Team plan this month.",
"charged twice refund"],
" Export": ["Hey there, how do I export my board data?",
"Hey there, how can I download my board as a spreadsheet?",
"Hey there, I need a CSV of my board.",
"Hey there, how do I EXPORT my BOARD data?",
"export board"],
}
hits = {"training wording": [0, 0], "reworded": [0, 0]}
for correct, variants in INTENTS.items():
print(f"\ncorrect first token {correct!r}")
for i, text in enumerate(variants):
p, top = p_token(f"Customer (Maya): {text}\nAgent (Dara):", correct)
bucket = hits["training wording" if i == 0 else "reworded"]
bucket[0] += top == correct
bucket[1] += 1
flag = "ok " if top == correct else "MISS"
print(f" {p:6.3f} {flag} top={top!r:14} {text}")
print()
for bucket, (ok, n) in hits.items():
print(f"{bucket:17} top-1 correct {ok}/{n}")Code explained
- In simple words: for three customer intents, try the exact training wording plus four paraphrases (reworded, reordered, shouted in capitals, or cut to keywords), and check whether the correct reply still comes out on top.
- What happens:
p_tokenencodes the prompt, runs the model once, and returns the probability of one specific token plus the model's actual top choice. The assertion guards against a "correct token" that is really several tokens. The script tallies top-1 hits separately for the training wording (first variant) and the rewordings. - Comes out: real output.
okmeans the correct first token is the top choice.
correct first token ' After'
0.998 ok top=' After' Hello, my account is locked after too many attempts.
0.015 MISS top=' Reset' Hello, I'm locked out after typing my password wrong a few times.
0.029 MISS top=' Sorry' Hello, it says my account is locked. Why?
0.007 MISS top=' Hey' HELLO, MY ACCOUNT IS LOCKED AFTER TOO MANY ATTEMPTS.
0.995 ok top=' After' account locked too many attempts
correct first token ' Sorry'
0.997 ok top=' Sorry' Hi, I was charged twice for the Team plan this month. Can you refund the duplicate?
0.065 MISS top=' Thank' Hi, you billed my card two times this month for Team, please refund one.
0.922 ok top=' Sorry' Hi, there is a duplicate charge on my card.
0.976 ok top=' Sorry' Hi, can you refund the duplicate? I was charged twice for the Team plan this month.
0.487 ok top=' Sorry' charged twice refund
correct first token ' Export'
0.997 ok top=' Export' Hey there, how do I export my board data?
0.032 MISS top=' Hey' Hey there, how can I download my board as a spreadsheet?
0.023 MISS top=' Thanks' Hey there, I need a CSV of my board.
0.797 ok top=' Export' Hey there, how do I EXPORT my BOARD data?
0.118 MISS top=' Hey' export board
training wording top-1 correct 3/3
reworded top-1 correct 5/12The training wording works every time (3 of 3, probability above 0.99). Rewordings work 5 of 12 times, and the results are not intuitive: the keyword-only "account locked too many attempts" works (0.995) while the natural "it says my account is locked. Why?" fails (0.029); lowercase-to-capitals flips lockout from 0.998 to 0.007 but barely hurts export (0.797). With n = 12 rewordings, the 95 percent interval on "5 of 12" is wide (roughly 0.19 to 0.68), so treat the exact rate as noisy; the gap between 3/3 and 5/12 is the finding.
Large models are far more robust than TinyLM, but the same effect is measurable in them: accuracy moves with wording, with the order of examples and options, and with formatting (Markdown vs plain text, where the question sits relative to the context). The engineering consequence runs through the whole course: never judge a prompt on one phrasing of one input. Build a test set of real, varied inputs first (Module 4), and measure changes on it (Module 10).
Calibration and the confident-wrong failure
A model is calibrated if its confidence matches its accuracy: of all the answers it gives with 80 percent confidence, about 80 percent should be right. Calibration is what would let you use confidence to decide when to escalate a ticket to a human. The most dangerous failure is the confident-wrong answer: high probability, wrong fact.
The probe set below takes 32 facts from the help center. For each, the prompt is the start of the fact and the gold answer is the token that correctly comes next. 26 use the corpus's own wording (some of those facts appear in hundreds of templated conversations; the rest appear only once, inside an article), and 6 are reworded.
"""Module 1: calibration. Is TinyLM's confidence a good guide to whether it is right?
Each probe is the start of a fact from the Brightlane help center (data/kb) and the
token that correctly comes next. Some use the corpus's exact wording, some are reworded.
"""
import math
import torch
from supportdesk.tinylm import load
torch.set_num_threads(1) # one thread is plenty for a 1M-parameter model
model, tokenizer = load()
PROBES = [ # (prompt, correct next token, kb article)
("Reset links expire after", " 30", "account-login"),
("After 5 failed attempts an account is locked for", " 15", "account-login"),
("Duplicate charges are always refunded in full within", " 5", "billing-refunds"),
("Annual plans cancelled within", " 14", "billing-refunds"),
("Team costs", " 12", "billing-plans"),
("Business costs", " 24", "billing-plans"),
("SSO is available on", " Business", "account-sso"),
("During an incident, updates are posted every", " 30", "status-incidents"),
("Limits: Team plans get", " 250", "boards-automations"),
("Offline mode caches the last", " 20", "mobile-app"),
("Deleted cards stay in Trash for", " 30", "data-privacy"),
("backups are purged within", " 90", "data-privacy"),
("Free supports up to", " 3", "billing-plans"),
("SCIM user provisioning is available on", " Enterprise", "account-sso"),
("Automations that call webhooks retry", " 3", "boards-automations"),
("Brightlane has apps for iOS", " 17", "mobile-app"),
("on Android 15, attachments larger than", " 25", "mobile-app"),
("A fix is planned for version", " 8", "mobile-app"),
("Enterprise pricing is custom and includes a", " 99", "billing-plans"),
("can request service credits within", " 30", "status-incidents"),
("Invoices are emailed to the", " billing", "billing-invoices"),
("arrive by email as a ZIP within", " 24", "exports-data"),
("use one of your", " 10", "account-login"),
("bank transfer (ACH or SEPA) for annual plans over", " 5", "billing-invoices"),
("The product team reviews the top-voted ideas every", " quarter", "feature-requests"),
("Moving an existing workspace between regions is available on", " Enterprise", "data-privacy"),
# reworded: same facts, wording the corpus never used
("A password reset link is valid for", " 30", "account-login"),
("How long is an account locked? It is locked for", " 15", "account-login"),
("The monthly automation run limit on the Business plan is", " 5", "boards-automations"),
("Refunds for duplicate charges arrive within", " 5", "billing-refunds"),
("The Team plan price per user per month is", " 12", "billing-plans"),
("Trash keeps deleted cards for", " 30", "data-privacy"),
]
@torch.no_grad()
def top_guess(prompt: str) -> tuple[str, float]:
probs = torch.softmax(model(torch.tensor([tokenizer.encode(prompt).ids]))[0, -1], dim=-1)
p, i = probs.max(dim=-1)
return tokenizer.decode([int(i)]), p.item()
rows = []
for prompt, answer, _ in PROBES:
first_answer_token = tokenizer.decode(tokenizer.encode(answer).ids[:1])
guess, confidence = top_guess(prompt)
rows.append((prompt, answer, guess, confidence, guess == first_answer_token))
print(f"{'conf':>5} {'ok':3} {'guess':12} {'gold':12} prompt")
for prompt, answer, guess, confidence, ok in sorted(rows, key=lambda r: -r[3]):
print(f"{confidence:5.2f} {'yes' if ok else 'NO':3} {guess!r:12} {answer!r:12} {prompt[:58]}")
BINS = [(0.0, 0.5), (0.5, 0.8), (0.8, 0.95), (0.95, 1.0001)]
print("\nReliability table")
print(f"{'confidence bin':16} {'n':>3} {'mean conf':>10} {'accuracy':>9} 95% interval (Wilson)")
ece = 0.0
for lo, hi in BINS:
in_bin = [r for r in rows if lo <= r[3] < hi]
if not in_bin:
continue
n, k = len(in_bin), sum(r[4] for r in in_bin)
conf = sum(r[3] for r in in_bin) / n
z = 1.96
centre = (k / n + z * z / (2 * n)) / (1 + z * z / n)
half = z * math.sqrt(k / n * (1 - k / n) / n + z * z / (4 * n * n)) / (1 + z * z / n)
ece += n / len(rows) * abs(conf - k / n)
print(f"[{lo:.2f}, {min(hi, 1):.2f}) {n:>3} {conf:10.2f} {k / n:9.2f} {max(centre - half, 0):.2f} to {min(centre + half, 1):.2f}")
acc = sum(r[4] for r in rows) / len(rows)
print(f"\noverall: n={len(rows)}, accuracy {acc:.2f}, mean confidence {sum(r[3] for r in rows) / len(rows):.2f}, "
f"expected calibration error {ece:.3f}")Code explained
- In simple words: for each fact, record the model's top guess and how confident it was, then group guesses by confidence and check how often each group was right.
- What happens:
top_guessreturns the highest-probability next token and its probability. Correctness compares it with the first token of the gold answer.- The reliability table puts each guess into a confidence bin and compares the bin's mean confidence with its accuracy. A calibrated model has the two columns roughly equal.
- The Wilson interval is a confidence interval for a proportion that behaves well at small n. With 5 or 6 items per bin, intervals are wide; that is the honest message.
- Expected calibration error (ECE) is the weighted average gap between confidence and accuracy across bins: 0 is perfect.
- Comes out: real output.
conf ok guess gold prompt
1.00 yes ' 5' ' 5' Duplicate charges are always refunded in full within
1.00 yes ' 14' ' 14' Annual plans cancelled within
0.99 yes ' 30' ' 30' Reset links expire after
0.98 NO ',' ' 8' A fix is planned for version
0.97 yes ' 30' ' 30' During an incident, updates are posted every
0.96 yes ' 15' ' 15' After 5 failed attempts an account is locked for
0.92 yes ' Business' ' Business' SSO is available on
0.92 yes ' 15' ' 15' How long is an account locked? It is locked for
0.85 NO ' reached' ' 5' The monthly automation run limit on the Business plan is
0.84 yes ' quarter' ' quarter' The product team reviews the top-voted ideas every
0.80 NO ' 10' ' 5' bank transfer (ACH or SEPA) for annual plans over
0.74 yes ' billing' ' billing' Invoices are emailed to the
0.66 NO ' available' ' 12' The Team plan price per user per month is
0.66 yes ' Enterprise' ' Enterprise' Moving an existing workspace between regions is available
0.66 NO ' Business' ' Enterprise' SCIM user provisioning is available on
0.58 yes ' 30' ' 30' Deleted cards stay in Trash for
0.55 yes ' 12' ' 12' Team costs
0.48 yes ' 30' ' 30' can request service credits within
0.42 NO ' a' ' 250' Limits: Team plans get
0.42 NO ' 12' ' 24' Business costs
0.41 yes ' 3' ' 3' Automations that call webhooks retry
0.40 yes ' 90' ' 90' backups are purged within
0.40 yes ' 5' ' 5' Refunds for duplicate charges arrive within
0.38 yes ' 24' ' 24' arrive by email as a ZIP within
0.36 NO ' identity' ' 10' use one of your
0.31 NO ' named' ' 99' Enterprise pricing is custom and includes a
0.29 NO ' 50' ' 25' on Android 15, attachments larger than
0.19 NO ' workspace' ' 3' Free supports up to
0.16 NO ' your' ' 17' Brightlane has apps for iOS
0.16 NO ' everyone' ' 30' A password reset link is valid for
0.08 NO '?' ' 20' Offline mode caches the last
0.07 NO ' unreleased' ' 30' Trash keeps deleted cards for
Reliability table
confidence bin n mean conf accuracy 95% interval (Wilson)
[0.00, 0.50) 15 0.30 0.33 0.15 to 0.58
[0.50, 0.80) 6 0.64 0.67 0.30 to 0.90
[0.80, 0.95) 5 0.86 0.60 0.23 to 0.88
[0.95, 1.00) 6 0.98 0.83 0.44 to 0.97
overall: n=32, accuracy 0.53, mean confidence 0.58, expected calibration error 0.089How to read it:
- Frequent facts are right and confident. Facts TinyLM saw hundreds of times in their exact wording (refund window, 14-day rule, reset links, incident updates, lockout time, SSO plans) come out at 0.92 to 1.00, correctly.
- Rare facts are mostly wrong, and usually (not always) with low confidence. Facts that appear once, like offline caching of 20 boards or the iOS 17 requirement, get 0.08 and 0.16.
- Confident-wrong exists even here. "A fix is planned for version" gets
','at 0.98. "The monthly automation run limit on the Business plan is" gets' reached'at 0.85: the model continues with a plausible phrase from the automations article ("when the limit is reached") instead of the number. "Business costs" gets' 12'(the Team price) at 0.42 because the corpus always states Team's price first. - Overall, TinyLM is roughly calibrated on this set (accuracy 0.53, mean confidence 0.58, ECE 0.089), which fits the GPT-4 report's finding that pretrained models tend to be calibrated. But the bins overlap within their intervals and n = 32, so the only safe conclusions are qualitative: confidence carries signal, and it is not a guarantee.
For a hosted assistant the same experiment needs token probabilities (not all APIs expose them) or a different confidence signal such as agreement across several samples (Module 5). Two rules follow for Brightlane. First, a threshold on confidence reduces errors but never removes them, so irreversible actions (refunds, account changes) always need a check that does not depend on the model's self-assessment (Module 8). Second, any threshold must be recalibrated on the real model and real tickets; TinyLM's numbers say nothing about gpt-oss-120b.
Knowledge cutoffs and stale facts
A model's knowledge cutoff is the date after which its training data stops. It knows nothing later, and it does not know what it does not know. Published cutoffs for the models in this module:
| Model | Knowledge cutoff | Source |
|---|---|---|
| Llama 3.1 | December 2023 | Llama 3.1 model card |
| gpt-oss-120b / 20b | June 2024 | gpt-oss model card |
| Gemini 3.5 Flash | January 2025 | Gemini API docs |
Today (September 2026) all three are 20 or more months out of date. For Brightlane that matters less than it sounds, because no public model knows Brightlane anyway. What matters is the pattern inside our own help center:
mobile-app.mdsays "Known issue (September 2026): on Android 15, attachments larger than 25 MB fail to upload. A fix is planned for version 8.4."feature-requests.mdsays "Currently planned for Q4 2026: Gantt view dependencies and a native Microsoft Teams integration."
Both will be wrong within months. If facts like these were baked into a fine-tuned model's weights, every change would need retraining, and the old answer would linger in the weights. TinyLM demonstrates the trap: it has these sentences memorized (the Gantt line even leaks into unrelated prompts, as you saw in Part E) and has no way to learn that version 8.4 shipped.
The design rule: facts that change live in retrieved documents, not in weights. The model supplies language skill; the help center supplies facts at request time (Module 7). Fine-tuning is for behavior and format, not for freshness (Module 9).
Self-knowledge
Models are unreliable narrators of themselves. Asked "what is your knowledge cutoff?", "how many parameters do you have?", or "can you browse the web?", a model produces a plausible answer from its training data, which often describes an older model or a different product. It can also misreport its own reasoning (Module 3 covers reading reasoning traces critically). Get these facts from the model card and the provider's documentation, and measure capability with your own evaluation. This is why every number in Part D came from a config file or a model card, never from asking a model.