Part C: What Pretraining Learned
Reading the training log
Pretraining is the first and largest training stage: show the model windows of text, ask it to predict every next token, measure the loss (cross-entropy: low when the model gave high probability to the actual next token), and nudge every parameter slightly in the direction that lowers the loss. A step is one such nudge on one batch (TinyLM uses 16 windows of 128 tokens per step). The script held out 5 percent of the corpus as a validation set: text the model never trains on, used only to check whether what it learned generalizes.
TinyLM's pretraining script saved its loss curve. Read it.
"""Module 1: read what pretraining did, from TinyLM's saved training log."""
import json
from pathlib import Path
log = json.loads(Path("models/tinylm-base/training_log.json").read_text())
best = min(log, key=lambda row: row["val_loss"])
print("step train val gap val curve (each # = 0.05 loss)")
for row in log:
gap = row["val_loss"] - row["train_loss"]
bar = "#" * min(round(row["val_loss"] / 0.05), 40)
mark = " <- lowest validation loss" if row is best else ""
print(f"{row['step']:>4} {row['train_loss']:6.3f} {row['val_loss']:6.3f} {gap:6.3f} {bar}{mark}")
print(f"\nstart: val loss {log[0]['val_loss']:.3f} (uniform guess over 2048 tokens would be {__import__('math').log(2048):.3f})")
print(f"best: step {best['step']}, val loss {best['val_loss']:.3f}")
print(f"end: step {log[-1]['step']}, train {log[-1]['train_loss']:.3f}, val {log[-1]['val_loss']:.3f}")
print(f"tokens seen: {log[-1]['tokens_seen']:,} in {log[-1]['seconds']:.0f} s")Code explained
- In simple words: print the training and validation loss every 100 steps as a table with a crude text bar chart, and mark the best validation point.
- What happens:
training_log.jsonwas written byscripts/pretrain_tinylm.pywhile training.gapis validation minus training loss: how much worse the model does on unseen text than on text it trained on. The bar length is proportional to validation loss. The script also compares the starting loss withln(2048), the loss of a model that spreads probability evenly over 2,048 vocabulary rows. - Comes out: real numbers from the saved log (they do not change between runs because they were recorded once).
step train val gap val curve (each # = 0.05 loss)
1 7.619 7.494 -0.125 ########################################
100 0.601 0.960 0.359 ###################
200 0.352 0.720 0.368 ##############
300 0.472 0.680 0.207 ##############
400 0.304 0.655 0.352 ############# <- lowest validation loss
500 0.284 0.679 0.394 ##############
600 0.287 0.669 0.382 #############
700 0.286 0.673 0.387 #############
800 0.269 0.687 0.418 ##############
900 0.288 0.678 0.390 ##############
1000 0.277 0.673 0.397 #############
1100 0.281 0.685 0.404 ##############
1200 0.275 0.682 0.407 ##############
1300 0.269 0.673 0.403 #############
1400 0.267 0.679 0.412 ##############
1500 0.306 0.679 0.373 ##############
1600 0.283 0.680 0.397 ##############
1700 0.276 0.678 0.401 ##############
1800 0.275 0.680 0.404 ##############
1900 0.278 0.680 0.402 ##############
2000 0.264 0.680 0.416 ##############
start: val loss 7.494 (uniform guess over 2048 tokens would be 7.625)
best: step 400, val loss 0.655
end: step 2000, train 0.264, val 0.680
tokens seen: 4,096,000 in 223 s- How to read it:
- Step 1: validation loss 7.49, close to 7.62 for a uniform guess. An untrained model knows nothing.
- Steps 1 to 400: both losses collapse. The model learns the vocabulary of support replies, the conversation format, and then specific answers.
- Step 400 onward: validation loss stops improving (best 0.655 at step 400, 0.680 at the end), while training loss keeps drifting down to about 0.26. The gap widens from about 0.35 to about 0.42. This is overfitting: the model keeps getting better at the exact text it trains on without getting better at unseen text. A validation loss of 0.68 means perplexity
exp(0.68) = 1.97: on held-out text the model is, on average, about as unsure as a choice between two tokens. - Why it happened: the run saw 4.1 million tokens but the training set has only 156,620, so it read every token about 26 times (26 epochs). With heavy repetition and a repetitive corpus, memorizing is the cheapest way to lower training loss.
The practical lessons carry straight to large models. Frontier labs train on so much text that they rarely repeat it more than a few times, and they watch held-out loss for exactly this reason. When you fine-tune in Module 9, with a few hundred examples, you will be in TinyLM's situation, and you will stop training by watching the validation curve the same way.
Memorization you can measure
If the model memorized, its outputs should contain long verbatim copies of the training text. Measure it.
"""Module 1: what pretraining stored (memorization) and what it cannot know (hallucination)."""
from pathlib import Path
import torch
from supportdesk.tinylm import SamplingParams, generate, load, next_token_distribution, perplexity
torch.set_num_threads(1) # one thread is plenty for a 1M-parameter model
model, tokenizer = load()
corpus = Path("data/corpus.txt").read_text(encoding="utf-8")
greedy = SamplingParams(max_new_tokens=40, temperature=0)
def longest_verbatim(text: str) -> str:
"""Longest prefix of `text` that appears word for word somewhere in the training corpus."""
lo, hi = 0, len(text)
while lo < hi:
mid = (lo + hi + 1) // 2
lo, hi = (mid, hi) if text[:mid] in corpus else (lo, mid - 1)
return text[:lo]
print("== Memorization: how much of each greedy reply is copied from the corpus? ==")
for question in ["Can I get a refund on my annual Team plan?", "Where can I find my invoices?",
"Can I pay by bank transfer?", "How many automation runs does Business get?"]:
prompt = f"Customer (Maya): Hi, {question[0].lower() + question[1:]}\nAgent (Dara):"
reply = generate(model, tokenizer, prompt, greedy).text.split("\n")[0]
copied = longest_verbatim(reply)
first_token, p_first = next_token_distribution(model, tokenizer, prompt, top=1)[0]
print(f"Q: {question}\nA:{reply}\n p(first token {first_token!r}) = {p_first:.2f}; "
f"first {len(copied)} of {len(reply)} characters appear verbatim in corpus.txt\n")
print("== Hallucination: a question the corpus never covers ==")
for prompt in ["The capital of France is", "Brightlane was founded in the year"]:
top = next_token_distribution(model, tokenizer, prompt, top=5)
print(prompt, "->", ", ".join(f"{t!r} {p:.2f}" for t, p in top))
print(" greedy:", repr(generate(model, tokenizer, prompt, greedy).text[:90]))
print("\n== Perplexity: how familiar does text look to the model? ==")
for label, text in [
("in-domain, seen", "Reset links expire after 30 minutes and can be used once."),
("in-domain, reworded", "A password reset link is valid for half an hour and works one time."),
("general knowledge", "Paris is the capital and largest city of France."),
]:
print(f"{label:22} {perplexity(model, tokenizer, text):10.1f}")Code explained
- In simple words: ask four questions, generate the agent's reply greedily, and check how many characters of each reply appear word for word in the training corpus. Then ask questions the corpus never covers, and finally score three sentences for how familiar they look.
- What happens:
longest_verbatimbinary-searches for the longest prefix of the reply that is a substring ofcorpus.txt. If the whole reply is found, the model reproduced training text exactly.- Two questions match corpus templates; two do not (bank transfer is only mentioned once, inside an article, never in a conversation; automation run counts likewise). The script also prints the probability of the reply's first token.
- The hallucination section feeds prompts about things the corpus never mentions and prints the top five next tokens and the greedy continuation.
perplexityscores whole sentences: lower means "this looks like my training data".- Comes out: real output.
== Memorization: how much of each greedy reply is copied from the corpus? ==
Q: Can I get a refund on my annual Team plan?
A: Annual plans cancelled within 14 days of purchase or renewal get a full refund. After 14 days they stay active until the end of the term.
p(first token ' Annual') = 1.00; first 138 of 138 characters appear verbatim in corpus.txt
Q: Where can I find my invoices?
A: Invoices are emailed to the billing contact and are available under Settings > Billing > Invoices.
p(first token ' Invoices') = 1.00; first 99 of 99 characters appear verbatim in corpus.txt
Q: Can I pay by bank transfer?
A: Sorry about the duplicate charge. Duplicate charges are always refunded in full within 5 to 10 business days to the original payment method.
p(first token ' Sorry') = 0.14; first 141 of 141 characters appear verbatim in corpus.txt
Q: How many automation runs does Business get?
A: SSO is available on Business and Enterprise plans. It supports SAML 2.0 with Okta, Entra ID, and Google Workspace.
p(first token ' SSO') = 0.20; first 115 of 115 characters appear verbatim in corpus.txt
== Hallucination: a question the corpus never covers ==
The capital of France is -> ' not' 0.20, ' stored' 0.13, ' available' 0.12, ' created' 0.11, ' locked' 0.08
greedy: ' not loading. Is there an outage?\nAgent (Nia): Please check status.brightlane.example for '
Brightlane was founded in the year -> '.' 0.08, ' dates' 0.05, ',' 0.05, ' cards' 0.04, ' ideas' 0.04
greedy: '.\nWhen the board URL, the browser or app is not loading. Is there an outage?\nAgent (Ana): '
== Perplexity: how familiar does text look to the model? ==
in-domain, seen 9.3
in-domain, reworded 4015.1
general knowledge 11918.6Four observations, each of which reappears later in the course:
- Every reply is 100 percent verbatim training text. Even the wrong ones. TinyLM does not compose answers; it retrieves memorized ones. Large models memorize far less in proportion, but they do memorize (Module 11 covers training-data extraction as a security risk).
- Wrong answers are fluent. Asked about bank transfer, TinyLM confidently apologizes for a duplicate charge. Asked about automation limits, it recites the SSO answer. Nothing in the text signals the error.
- The warning sign is in the numbers you usually never see. Right answers started with probability 1.00; wrong ones with 0.14 and 0.20. Hosted APIs can expose token probabilities (often called logprobs) for some models, and you will use signals like this in Modules 5 and 10. But as Part F shows, confidence is an imperfect guide.
- Rewording is as foreign as a new topic. The reworded reset-link sentence has perplexity 4,015 versus 9.3 for the memorized wording. TinyLM learned strings, not meanings. Larger models generalize across wording far better, but the direction of the effect (familiar phrasing works better) persists, which is why Part F measures phrasing sensitivity.
Hallucination is a property of the objective
"The capital of France is" has one obvious continuation, and TinyLM's top guess is ' not' (20 percent), continuing into a memorized outage reply. It is not broken. It is doing exactly what it was trained to do: produce the most plausible continuation given its training data, and its training data is a support desk. There is no "I don't know" in the objective. A distribution over tokens must put its probability somewhere, and generation always emits a token.
A hallucination is fluent output that is not supported by facts or by the provided context. Large models hallucinate less often than TinyLM because they have read vastly more, and post-training (Part E) teaches them to say "I am not sure" in some situations. But the root cause is identical and cannot be patched out of the objective: the model is rewarded for plausible text, not for true text, and for rare facts plausible and true diverge. That is why this course grounds answers in retrieved help-center text (Module 7), constrains outputs (Module 6), and evaluates rather than trusts (Module 10).