CourseLarge Language Models · Module 1: What a Large Language Model Actually Is · part 4 of 80
Part 4 · Module 1: What a Large Language Model Actually Is

Part D: Reading a Model Spec

8 min read·22 Sept 2026

The four numbers on every spec sheet

When a model is released, its model card and config.json list a handful of numbers. Four of them drive most practical decisions:

  • Parameters: how many learned numbers the model has. More parameters can store more knowledge and skill, and cost more memory and compute to run. For a mixture-of-experts (MoE) model, two counts matter: total parameters (all must sit in memory) and active parameters (the subset used for each token, which sets compute per token and therefore much of the speed).
  • Layers: how many transformer blocks are stacked. Together with the hidden size (d_model, the width of each token's vector), layers set how much computation happens per token.
  • Context window: the maximum number of tokens (prompt plus output) the model can attend to at once. Anything beyond it is invisible to the model. Some models list a "native" window and a longer one reached with a position-scaling technique such as YaRN; quality in the extended range is worth testing yourself (Module 2).
  • Vocabulary: how many distinct tokens the tokenizer can produce. Larger vocabularies usually mean fewer tokens per sentence (cheaper and more room in the context window), especially for non-English text, at the cost of a bigger embedding table.

Counting parameters, then checking a published config

python
"""Module 1: read a model spec. Count TinyLM's parameters, check the arithmetic on a
published config, and turn parameter counts into memory."""
import os
from collections import defaultdict

import torch

from supportdesk.tinylm import load

torch.set_num_threads(1)  # one thread is plenty for a 1M-parameter model
model, tokenizer = load()
cfg = model.cfg

# 1. Where TinyLM's parameters live (tied weights are counted once by .parameters()).
groups: dict[str, int] = defaultdict(int)
for name, p in model.named_parameters():
    part = "token embeddings (shared with output head)" if name.startswith("tok_emb") else \
           "position embeddings" if name.startswith("pos_emb") else \
           "attention (4 layers)" if ".qkv." in name or ".proj." in name else \
           "feed-forward MLP (4 layers)" if ".mlp." in name else "layer norms"
    groups[part] += p.numel()
total = model.num_parameters()
print(f"TinyLM config: {cfg}")
for part, n in sorted(groups.items(), key=lambda kv: -kv[1]):
    print(f"  {part:44} {n:>9,}  {n / total:6.1%}")
print(f"  {'total':44} {total:>9,}")
print(f"  tokenizer vocabulary: {tokenizer.get_vocab_size()}")


# 2. The same arithmetic for a modern dense model, from the numbers in its config.json.
def dense_params(vocab: int, d: int, layers: int, heads: int, kv_heads: int, head_dim: int,
                 ffn: int, tied: bool) -> int:
    """Approximate parameters of a Llama/Qwen-style decoder (ignores norms and biases)."""
    attention = d * heads * head_dim * 2 + d * kv_heads * head_dim * 2   # Q and O, then K and V
    mlp = 3 * d * ffn                                                    # gated MLP: gate, up, down
    embeddings = vocab * d * (1 if tied else 2)                          # input table (+ output head)
    return layers * (attention + mlp) + embeddings


published = {
    # name: (config.json values, published headline)
    "Qwen3-8B": (dict(vocab=151936, d=4096, layers=36, heads=32, kv_heads=8, head_dim=128, ffn=12288, tied=False), "8.2B"),
    "Llama-3.1-8B": (dict(vocab=128256, d=4096, layers=32, heads=32, kv_heads=8, head_dim=128, ffn=14336, tied=False), "8B"),
}
print("\nmodel          estimate from config   published")
for name, (c, headline) in published.items():
    print(f"{name:14} {dense_params(**c) / 1e9:8.2f}B              {headline}")

# 3. Memory for the weights alone = parameters x bytes per parameter.
BYTES = {"fp32": 4, "bf16": 2, "int8": 1, "4-bit": 0.5}
sizes = {"Qwen3-8B": 8.2e9, "gpt-oss-20b": 20.91e9, "gpt-oss-120b": 116.83e9}
print("\nweights only, GB (1e9 bytes)")
print(f"{'model':14}" + "".join(f"{k:>9}" for k in BYTES))
for name, n in sizes.items():
    print(f"{name:14}" + "".join(f"{n * b / 1e9:9.1f}" for b in BYTES.values()))
print(f"{'TinyLM (MB)':14}" + "".join(f"{total * b / 1e6:9.2f}" for b in BYTES.values()))
print(f"\nTinyLM model.pt on disk: {os.path.getsize('models/tinylm-base/model.pt') / 1e6:.2f} MB")

# 4. Long context costs memory too: the KV cache stores keys and values for every token.
print("\nKV cache at bf16 (2 bytes): 2 (K and V) x layers x kv_heads x head_dim x 2 bytes per token")
for name, (c, _) in published.items():
    per_token = 2 * c["layers"] * c["kv_heads"] * c["head_dim"] * 2
    print(f"{name:14} {per_token:,} bytes/token; 32,768 tokens = {per_token * 32768 / 1e9:.1f} GB; "
          f"131,072 tokens = {per_token * 131072 / 1e9:.1f} GB")

Code explained

  • In simple words: count where TinyLM's parameters live, then use the same arithmetic on the numbers in two published config.json files to see if we can reproduce their headline sizes, and finally convert parameter counts into memory.
  • What happens:
  • Section 1 walks model.named_parameters() and groups each tensor by the part of the architecture it belongs to. PyTorch reports the shared (tied) embedding and output weights once.
  • Section 2's dense_params is the standard back-of-envelope formula for a Llama- or Qwen-style decoder. Attention has four projection matrices (query, key, value, output); with grouped-query attention (GQA) the key and value projections are smaller because several query heads share one key/value head. The MLP in these models is "gated", with three matrices. Embeddings are vocab x d, twice if the output head is not tied. The config values come from each model's published config.json on Hugging Face (Qwen3-8B, and Llama 3.1 8B's, mirrored at unsloth/Meta-Llama-3.1-8B-Instruct because Meta's own repository requires sign-in).
  • Section 3 multiplies parameters by bytes per parameter for four common precisions (how many bits store each number): fp32 (4 bytes), bf16 (2), int8 (1), and 4-bit (half a byte). Storing weights at lower precision is called quantization.
  • Section 4 estimates the KV cache: the keys and values that attention stores for every token in the context. It grows linearly with context length.
  • Comes out: real output.
text
  TinyLM config: TinyConfig(vocab_size=2048, context=128, d_model=128, n_layers=4, n_heads=4, dropout=0.0)
    feed-forward MLP (4 layers)                    526,848   49.2%
    attention (4 layers)                           264,192   24.6%
    token embeddings (shared with output head)     262,144   24.5%
    position embeddings                             16,384    1.5%
    layer norms                                      2,304    0.2%
    total                                        1,071,872
    tokenizer vocabulary: 1503

  model          estimate from config   published
  Qwen3-8B           8.19B              8.2B
  Llama-3.1-8B       8.03B              8B

  weights only, GB (1e9 bytes)
  model              fp32     bf16     int8    4-bit
  Qwen3-8B           32.8     16.4      8.2      4.1
  gpt-oss-20b        83.6     41.8     20.9     10.5
  gpt-oss-120b      467.3    233.7    116.8     58.4
  TinyLM (MB)        4.29     2.14     1.07     0.54

  TinyLM model.pt on disk: 4.30 MB

  KV cache at bf16 (2 bytes): 2 (K and V) x layers x kv_heads x head_dim x 2 bytes per token
  Qwen3-8B       147,456 bytes/token; 32,768 tokens = 4.8 GB; 131,072 tokens = 19.3 GB
  Llama-3.1-8B   131,072 bytes/token; 32,768 tokens = 4.3 GB; 131,072 tokens = 17.2 GB

What this tells you:

  • In TinyLM, half the parameters are in the MLPs and a quarter in the embedding table. In an 8B model, embeddings are a smaller share (Qwen3-8B: 1.25B of 8.2B), but the pattern of "most parameters are in the MLPs" holds.
  • The formula reproduces published sizes to within 1 percent (8.19B vs 8.2B, 8.03B vs 8B). Qwen's model card also lists 6.95B non-embedding parameters, which is exactly our layer total. You can now sanity-check any dense model's headline size from its config in a minute.
  • The config's vocabulary is not always the tokenizer's. TinyLM's config reserves 2,048 embedding rows, but its tokenizer, trained on a tiny corpus, only learned 1,503 tokens, so 545 rows are never used. Production configs often pad the vocabulary too (Qwen3's 151,936 rows is a rounded-up size). Count tokens with the tokenizer, not the config.
  • Memory for weights is parameters times bytes. An 8B model needs about 16 GB at bf16 and about 4 GB at 4-bit. model.pt on disk (4.30 MB) matches the fp32 figure for TinyLM (4.29 MB) plus a little file overhead.
  • Context costs memory too. At bf16, a full 131,072-token context for Llama 3.1 8B needs about 17 GB of KV cache, more than the 16 GB of weights. Grouped-query attention (8 key/value heads instead of 32) is what keeps this from being four times larger. Module 13 turns this into hardware sizing.

gpt-oss-120b is a useful reality check. Our 4-bit estimate is 58.4 GB. OpenAI's model card reports a 60.8 GiB checkpoint (about 65 GB) using the MXFP4 format at 4.25 bits per parameter for the MoE weights, with other weights kept at higher precision (gpt-oss model card). The simple arithmetic lands within about 10 percent, and it tells you instantly that the model fits on one 80 GB GPU but not on a 24 GB consumer card.

Four spec sheets side by side

TinyLMQwen3-8BLlama 3.1 8Bgpt-oss-120b
Parameters1,071,8728.2B (6.95B non-embedding)8B116.83B total, 5.13B active per token
ArchitectureDenseDenseDenseMixture of experts: 128 experts, 4 active per token
Layers4363236
Hidden size (d_model)1284,0964,0962,880
Attention heads (query / key-value)4 / 432 / 832 / 864 / 8
Context window12832,768 native; 131,072 with YaRN131,072131,072
Vocabulary (config rows)2,048 (1,503 used)151,936128,256201,088 (o200k_harmony tokenizer)
Pretraining data156,620 tokensAbout 36 trillion tokens, 119 languagesAbout 15 trillion tokensNot disclosed; about 2.1 million H100 GPU-hours of compute
Knowledge cutoffWhatever is in corpus.txtNot stated on the model cardDecember 2023June 2024
LicenceCourse codeApache 2.0Llama 3.1 Community LicenseApache 2.0
ReleasedThis courseApril 202523 July 20245 August 2025

Sources: Qwen3-8B model card and Qwen3 Technical Report; Llama 3.1 model card; Introducing gpt-oss, the gpt-oss model card, and its config.json. Checked 21 September 2026.

Two of these are the course's defaults: qwen3:8b on Ollama and openai/gpt-oss-120b on Groq. Notice what the numbers already predict. gpt-oss-120b stores 14 times more parameters than Qwen3-8B but computes with fewer per token (5.1B active vs 8.2B), so it can be fast on the right hardware while needing far more memory. That is the MoE trade in one line.

SituationUse thisWhy
You have one 80 GB datacenter GPU and want the strongest open model that fitsA large MoE at 4-bit (for example gpt-oss-120b)Total parameters decide memory; a low active count keeps it fast
You want to run locally on a 16 to 24 GB laptop or consumer GPUAn 8B dense model at 4 to 8 bits, or a small MoE such as gpt-oss-20bWeights of about 4 to 13 GB leave room for the KV cache
You need very long contexts on your own hardwareBudget KV-cache memory explicitly, and prefer models with few key-value headsThe cache can exceed the weights at 100k+ tokens
You only call hosted APIsRead the spec for context window, knowledge cutoff, and price, not memoryThe provider carries the hardware; you pay per token