Part D: Reading a Model Spec
The four numbers on every spec sheet
When a model is released, its model card and config.json list a handful of numbers. Four of them drive most practical decisions:
- Parameters: how many learned numbers the model has. More parameters can store more knowledge and skill, and cost more memory and compute to run. For a mixture-of-experts (MoE) model, two counts matter: total parameters (all must sit in memory) and active parameters (the subset used for each token, which sets compute per token and therefore much of the speed).
- Layers: how many transformer blocks are stacked. Together with the hidden size (
d_model, the width of each token's vector), layers set how much computation happens per token. - Context window: the maximum number of tokens (prompt plus output) the model can attend to at once. Anything beyond it is invisible to the model. Some models list a "native" window and a longer one reached with a position-scaling technique such as YaRN; quality in the extended range is worth testing yourself (Module 2).
- Vocabulary: how many distinct tokens the tokenizer can produce. Larger vocabularies usually mean fewer tokens per sentence (cheaper and more room in the context window), especially for non-English text, at the cost of a bigger embedding table.
Counting parameters, then checking a published config
"""Module 1: read a model spec. Count TinyLM's parameters, check the arithmetic on a
published config, and turn parameter counts into memory."""
import os
from collections import defaultdict
import torch
from supportdesk.tinylm import load
torch.set_num_threads(1) # one thread is plenty for a 1M-parameter model
model, tokenizer = load()
cfg = model.cfg
# 1. Where TinyLM's parameters live (tied weights are counted once by .parameters()).
groups: dict[str, int] = defaultdict(int)
for name, p in model.named_parameters():
part = "token embeddings (shared with output head)" if name.startswith("tok_emb") else \
"position embeddings" if name.startswith("pos_emb") else \
"attention (4 layers)" if ".qkv." in name or ".proj." in name else \
"feed-forward MLP (4 layers)" if ".mlp." in name else "layer norms"
groups[part] += p.numel()
total = model.num_parameters()
print(f"TinyLM config: {cfg}")
for part, n in sorted(groups.items(), key=lambda kv: -kv[1]):
print(f" {part:44} {n:>9,} {n / total:6.1%}")
print(f" {'total':44} {total:>9,}")
print(f" tokenizer vocabulary: {tokenizer.get_vocab_size()}")
# 2. The same arithmetic for a modern dense model, from the numbers in its config.json.
def dense_params(vocab: int, d: int, layers: int, heads: int, kv_heads: int, head_dim: int,
ffn: int, tied: bool) -> int:
"""Approximate parameters of a Llama/Qwen-style decoder (ignores norms and biases)."""
attention = d * heads * head_dim * 2 + d * kv_heads * head_dim * 2 # Q and O, then K and V
mlp = 3 * d * ffn # gated MLP: gate, up, down
embeddings = vocab * d * (1 if tied else 2) # input table (+ output head)
return layers * (attention + mlp) + embeddings
published = {
# name: (config.json values, published headline)
"Qwen3-8B": (dict(vocab=151936, d=4096, layers=36, heads=32, kv_heads=8, head_dim=128, ffn=12288, tied=False), "8.2B"),
"Llama-3.1-8B": (dict(vocab=128256, d=4096, layers=32, heads=32, kv_heads=8, head_dim=128, ffn=14336, tied=False), "8B"),
}
print("\nmodel estimate from config published")
for name, (c, headline) in published.items():
print(f"{name:14} {dense_params(**c) / 1e9:8.2f}B {headline}")
# 3. Memory for the weights alone = parameters x bytes per parameter.
BYTES = {"fp32": 4, "bf16": 2, "int8": 1, "4-bit": 0.5}
sizes = {"Qwen3-8B": 8.2e9, "gpt-oss-20b": 20.91e9, "gpt-oss-120b": 116.83e9}
print("\nweights only, GB (1e9 bytes)")
print(f"{'model':14}" + "".join(f"{k:>9}" for k in BYTES))
for name, n in sizes.items():
print(f"{name:14}" + "".join(f"{n * b / 1e9:9.1f}" for b in BYTES.values()))
print(f"{'TinyLM (MB)':14}" + "".join(f"{total * b / 1e6:9.2f}" for b in BYTES.values()))
print(f"\nTinyLM model.pt on disk: {os.path.getsize('models/tinylm-base/model.pt') / 1e6:.2f} MB")
# 4. Long context costs memory too: the KV cache stores keys and values for every token.
print("\nKV cache at bf16 (2 bytes): 2 (K and V) x layers x kv_heads x head_dim x 2 bytes per token")
for name, (c, _) in published.items():
per_token = 2 * c["layers"] * c["kv_heads"] * c["head_dim"] * 2
print(f"{name:14} {per_token:,} bytes/token; 32,768 tokens = {per_token * 32768 / 1e9:.1f} GB; "
f"131,072 tokens = {per_token * 131072 / 1e9:.1f} GB")Code explained
- In simple words: count where TinyLM's parameters live, then use the same arithmetic on the numbers in two published
config.jsonfiles to see if we can reproduce their headline sizes, and finally convert parameter counts into memory. - What happens:
- Section 1 walks
model.named_parameters()and groups each tensor by the part of the architecture it belongs to. PyTorch reports the shared (tied) embedding and output weights once. - Section 2's
dense_paramsis the standard back-of-envelope formula for a Llama- or Qwen-style decoder. Attention has four projection matrices (query, key, value, output); with grouped-query attention (GQA) the key and value projections are smaller because several query heads share one key/value head. The MLP in these models is "gated", with three matrices. Embeddings arevocab x d, twice if the output head is not tied. The config values come from each model's publishedconfig.jsonon Hugging Face (Qwen3-8B, and Llama 3.1 8B's, mirrored at unsloth/Meta-Llama-3.1-8B-Instruct because Meta's own repository requires sign-in). - Section 3 multiplies parameters by bytes per parameter for four common precisions (how many bits store each number): fp32 (4 bytes), bf16 (2), int8 (1), and 4-bit (half a byte). Storing weights at lower precision is called quantization.
- Section 4 estimates the KV cache: the keys and values that attention stores for every token in the context. It grows linearly with context length.
- Comes out: real output.
TinyLM config: TinyConfig(vocab_size=2048, context=128, d_model=128, n_layers=4, n_heads=4, dropout=0.0)
feed-forward MLP (4 layers) 526,848 49.2%
attention (4 layers) 264,192 24.6%
token embeddings (shared with output head) 262,144 24.5%
position embeddings 16,384 1.5%
layer norms 2,304 0.2%
total 1,071,872
tokenizer vocabulary: 1503
model estimate from config published
Qwen3-8B 8.19B 8.2B
Llama-3.1-8B 8.03B 8B
weights only, GB (1e9 bytes)
model fp32 bf16 int8 4-bit
Qwen3-8B 32.8 16.4 8.2 4.1
gpt-oss-20b 83.6 41.8 20.9 10.5
gpt-oss-120b 467.3 233.7 116.8 58.4
TinyLM (MB) 4.29 2.14 1.07 0.54
TinyLM model.pt on disk: 4.30 MB
KV cache at bf16 (2 bytes): 2 (K and V) x layers x kv_heads x head_dim x 2 bytes per token
Qwen3-8B 147,456 bytes/token; 32,768 tokens = 4.8 GB; 131,072 tokens = 19.3 GB
Llama-3.1-8B 131,072 bytes/token; 32,768 tokens = 4.3 GB; 131,072 tokens = 17.2 GBWhat this tells you:
- In TinyLM, half the parameters are in the MLPs and a quarter in the embedding table. In an 8B model, embeddings are a smaller share (Qwen3-8B: 1.25B of 8.2B), but the pattern of "most parameters are in the MLPs" holds.
- The formula reproduces published sizes to within 1 percent (8.19B vs 8.2B, 8.03B vs 8B). Qwen's model card also lists 6.95B non-embedding parameters, which is exactly our layer total. You can now sanity-check any dense model's headline size from its config in a minute.
- The config's vocabulary is not always the tokenizer's. TinyLM's config reserves 2,048 embedding rows, but its tokenizer, trained on a tiny corpus, only learned 1,503 tokens, so 545 rows are never used. Production configs often pad the vocabulary too (Qwen3's 151,936 rows is a rounded-up size). Count tokens with the tokenizer, not the config.
- Memory for weights is parameters times bytes. An 8B model needs about 16 GB at bf16 and about 4 GB at 4-bit.
model.pton disk (4.30 MB) matches the fp32 figure for TinyLM (4.29 MB) plus a little file overhead. - Context costs memory too. At bf16, a full 131,072-token context for Llama 3.1 8B needs about 17 GB of KV cache, more than the 16 GB of weights. Grouped-query attention (8 key/value heads instead of 32) is what keeps this from being four times larger. Module 13 turns this into hardware sizing.
gpt-oss-120b is a useful reality check. Our 4-bit estimate is 58.4 GB. OpenAI's model card reports a 60.8 GiB checkpoint (about 65 GB) using the MXFP4 format at 4.25 bits per parameter for the MoE weights, with other weights kept at higher precision (gpt-oss model card). The simple arithmetic lands within about 10 percent, and it tells you instantly that the model fits on one 80 GB GPU but not on a 24 GB consumer card.
Four spec sheets side by side
| TinyLM | Qwen3-8B | Llama 3.1 8B | gpt-oss-120b | |
|---|---|---|---|---|
| Parameters | 1,071,872 | 8.2B (6.95B non-embedding) | 8B | 116.83B total, 5.13B active per token |
| Architecture | Dense | Dense | Dense | Mixture of experts: 128 experts, 4 active per token |
| Layers | 4 | 36 | 32 | 36 |
Hidden size (d_model) | 128 | 4,096 | 4,096 | 2,880 |
| Attention heads (query / key-value) | 4 / 4 | 32 / 8 | 32 / 8 | 64 / 8 |
| Context window | 128 | 32,768 native; 131,072 with YaRN | 131,072 | 131,072 |
| Vocabulary (config rows) | 2,048 (1,503 used) | 151,936 | 128,256 | 201,088 (o200k_harmony tokenizer) |
| Pretraining data | 156,620 tokens | About 36 trillion tokens, 119 languages | About 15 trillion tokens | Not disclosed; about 2.1 million H100 GPU-hours of compute |
| Knowledge cutoff | Whatever is in corpus.txt | Not stated on the model card | December 2023 | June 2024 |
| Licence | Course code | Apache 2.0 | Llama 3.1 Community License | Apache 2.0 |
| Released | This course | April 2025 | 23 July 2024 | 5 August 2025 |
Sources: Qwen3-8B model card and Qwen3 Technical Report; Llama 3.1 model card; Introducing gpt-oss, the gpt-oss model card, and its config.json. Checked 21 September 2026.
Two of these are the course's defaults: qwen3:8b on Ollama and openai/gpt-oss-120b on Groq. Notice what the numbers already predict. gpt-oss-120b stores 14 times more parameters than Qwen3-8B but computes with fewer per token (5.1B active vs 8.2B), so it can be fast on the right hardware while needing far more memory. That is the MoE trade in one line.
| Situation | Use this | Why |
|---|---|---|
| You have one 80 GB datacenter GPU and want the strongest open model that fits | A large MoE at 4-bit (for example gpt-oss-120b) | Total parameters decide memory; a low active count keeps it fast |
| You want to run locally on a 16 to 24 GB laptop or consumer GPU | An 8B dense model at 4 to 8 bits, or a small MoE such as gpt-oss-20b | Weights of about 4 to 13 GB leave room for the KV cache |
| You need very long contexts on your own hardware | Budget KV-cache memory explicitly, and prefer models with few key-value heads | The cache can exceed the weights at 100k+ tokens |
| You only call hosted APIs | Read the spec for context window, knowledge cutoff, and price, not memory | The provider carries the hardware; you pay per token |