CourseLarge Language Models · Module 3: Inference Behaviour and Decoding Control · part 14 of 80
Part 14 · Module 3: Inference Behaviour and Decoding Control

Part C: Constrained generation

13 min read·22 Sept 2026

Forcing a format token by token

Prompting asks the model for a format. Constrained decoding enforces it: at every step, tokens that would break the format get probability zero before sampling. The model can only choose among continuations that stay valid. TinyLM's generate supports this through its allowed argument, a function that receives the tokens generated so far and returns the set of token ids permitted next. An empty set means "done".

For a fixed list of answers, the right data structure is a trie over token sequences: tokenize every valid answer, and at step n allow exactly the tokens that continue some answer whose first n tokens match what was generated. Our first target is the six Brightlane categories. TinyLM was never trained to classify tickets, which makes this an honest test of what a constraint can and cannot do.

examples/m03_constrained.py

python
"""Constrained decoding with generate(allowed=...): TinyLM can only emit one of the six categories."""
import math
from collections import Counter

import torch

from examples.m03_setup import MODEL, TOK
from supportdesk.data import CATEGORIES, load_tickets
from supportdesk.tinylm import SamplingParams, generate


def trie_constraint(options: list[str]):
    """Build an `allowed` function: only token sequences that spell one of `options` may be generated."""
    sequences = [TOK.encode(o).ids for o in options]

    def allowed(generated: list[int]) -> set[int]:
        n = len(generated)
        return {seq[n] for seq in sequences if len(seq) > n and seq[:n] == generated}   # empty set = finished
    return allowed


def prompt_for(ticket) -> str:
    return f"Ticket: {ticket.subject}\n{ticket.body}\nCategory:"


@torch.no_grad()
def score_options(prompt: str, options: list[str]) -> dict[str, float]:
    """Total log-probability the raw model assigns to each option as the continuation of `prompt`."""
    ids = TOK.encode(prompt).ids[-100:]                      # keep room inside the 128-token context
    scores = {}
    for option in options:
        cont = TOK.encode(option).ids
        logp = torch.log_softmax(MODEL(torch.tensor([ids + cont]))[0], dim=-1)
        scores[option] = sum(logp[len(ids) - 1 + i, t].item() for i, t in enumerate(cont))
    return scores


def main() -> None:
    OPTIONS = [" " + c for c in CATEGORIES]
    allowed = trie_constraint(OPTIONS)
    dev = load_tickets("dev")
    first_tokens = sorted({TOK.encode(o).ids[0] for o in OPTIONS})

    free, sampled, greedy, scored, mass = [], [], [], [], []
    for t in dev:
        p = prompt_for(t)
        free.append(generate(MODEL, TOK, p, SamplingParams(max_new_tokens=6, temperature=0)).text)
        sampled.append(generate(MODEL, TOK, p, SamplingParams(max_new_tokens=8, temperature=1.0, seed=0), allowed=allowed).text)
        greedy.append(generate(MODEL, TOK, p, SamplingParams(max_new_tokens=8, temperature=0), allowed=allowed).text)
        s = score_options(p, OPTIONS)
        scored.append((max(s, key=s.get), s))
        with torch.no_grad():
            ids = TOK.encode(p).ids[-100:]
            probs = torch.softmax(MODEL(torch.tensor([ids]))[0, -1], dim=-1)
        mass.append(probs[first_tokens].sum().item())          # how much the model "wanted" any category at all

    gold = [t.gold["category"] for t in dev]
    n = len(dev)


    def report(label, outputs):
        parsed = [o.strip() for o in outputs]
        ok = sum(o in CATEGORIES for o in parsed)
        acc = sum(o == g for o, g in zip(parsed, gold))
        print(f"{label:>30} | parses {ok:>2}/{n} | correct {acc:>2}/{n} | most common: {Counter(parsed).most_common(1)[0]}")


    print("first ticket, unconstrained greedy:", repr(free[0]))
    report("unconstrained greedy", free)
    report("constrained, sampled (T=1)", sampled)
    report("constrained, greedy (T=0)", greedy)
    report("constrained, scored (argmax)", [c for c, _ in scored])
    majority = Counter(gold).most_common(1)[0][0]
    per_token = [max(s, key=lambda o: s[o] / len(TOK.encode(o).ids)) for _, s in scored]   # length-normalized
    report("scored per token (argmax)", per_token)
    report(f"majority baseline ({majority})", [majority] * n)

    best_logp = [s[c] for c, s in scored]
    print(f"\nprobability mass on any category's first token: median {sorted(mass)[n // 2]:.5f}, max {max(mass):.5f}")
    print(f"log-probability of the chosen category: median {sorted(best_logp)[n // 2]:.1f} "
          f"(probability about {math.exp(sorted(best_logp)[n // 2]):.1e})")
    c, s = scored[0]
    print("ticket 1 option scores:", {k.strip(): round(v, 1) for k, v in sorted(s.items(), key=lambda kv: -kv[1])})


if __name__ == "__main__":
    main()

Code explained

  • In simple words: we build a "you may only spell one of these six words" rule, run it on all 48 dev tickets several ways, and check both whether the answer parses and whether it is right.
  • What happens: trie_constraint tokenizes each option (for example feature_request becomes six tokens: feat, ure, _, re, qu, est) and returns an allowed function. prompt_for builds Ticket: ...\nCategory:. For each ticket we run: unconstrained greedy; constrained sampling at temperature 1 with a seed; constrained greedy; and scoring, which computes the total log-probability of each of the six options and takes the best (with a per-token variant that divides by length). mass records how much probability the model put on the first token of any category before we forced anything. A majority-class baseline sits at the bottom for comparison.
  • Comes out: PYTHONPATH=. python examples/m03_constrained.py:
text
first ticket, unconstrained greedy: ' Gantt view dependencies and a workspace'
          unconstrained greedy | parses  0/48 | correct  0/48 | most common: ('Team costs 12 USD per user', 29)
    constrained, sampled (T=1) | parses 48/48 | correct 10/48 | most common: ('billing', 43)
     constrained, greedy (T=0) | parses 48/48 | correct 12/48 | most common: ('how_to', 48)
  constrained, scored (argmax) | parses 48/48 | correct 11/48 | most common: ('billing', 48)
     scored per token (argmax) | parses 48/48 | correct 11/48 | most common: ('billing', 48)
    majority baseline (how_to) | parses 48/48 | correct 12/48 | most common: ('how_to', 48)

probability mass on any category's first token: median 0.00072, max 0.00161
log-probability of the chosen category: median -8.1 (probability about 3.0e-04)
ticket 1 option scores: {'billing': -7.9, 'bug': -10.2, 'cancellation': -20.0, 'how_to': -28.5, 'account_access': -41.9, 'feature_request': -58.7}

Read it top to bottom. Unconstrained, TinyLM writes support-desk text and never produces a category: 0 of 48 parse. With the constraint, 48 of 48 parse, every time, by construction. That is what constraints are for.

Correctness is another matter. Constrained sampling gets 10 of 48 right and scoring gets 11 of 48, against 12 of 48 for always answering how_to. With n=48, a 95 percent Wilson interval around 11/48 runs from about 13 to 37 percent, so none of these differ from the baseline beyond noise. The constraint guaranteed the shape of the answer and contributed nothing to its accuracy. TinyLM simply does not know the task.

The diagnostic that tells you this without gold labels is the probability mass line. Before forcing, TinyLM put a median of 0.07 percent of its probability on the first token of any category. The constraint then renormalized that sliver into 100 percent. Whenever the allowed tokens hold a tiny share of the model's probability, the format is doing all the work, and the answer is closer to a coin flip than to a judgment. Scoring shows the same thing: the chosen category's log-probability has a median of -8.1 (about 3 in 10,000), and on ticket 1 the ranking follows token count ( billing is one token, feature_request six) more than meaning. Dividing by length did not change a single answer here.

A failure to diagnose: why is every greedy answer how_to? The constrained, greedy (T=0) row answered how_to for all 48 tickets, and scored 12/48, identical to the majority baseline. That looks like a sensible model with a prior. It is a bug. Read generate from Part A: at temperature 0, apply_sampling returns a one-hot vector on the model's overall favorite token. If that token is not allowed, multiplying by the mask leaves all zeros, so the code falls back to mask / mask.sum(), a uniform distribution over the allowed tokens, and then takes argmax, which returns the first maximum: the lowest token id. The allowed first tokens are ids 540 ( billing), 619, 576, 1240, 417 ( how), and 1200, so how wins every time. The model's preferences among the allowed tokens were thrown away.

The lesson generalizes: with constraints, apply the mask to the model's logits or full distribution first, then pick greedily among what remains. Never pick greedily and then check. In this module we work around it by sampling within the constraint (temperature above 0) or by scoring the options directly, which is also how many classification systems use LLMs in practice. Real grammar engines apply the mask to logits before any sampling transform, so they do not have this problem.

A small JSON schema, enforced

Categories are one field. Real outputs are structured objects. The triage object {"category": ..., "priority": ...} has 6 x 4 = 24 valid values, so the same trie enforces a complete JSON "schema".

examples/m03_json_constrained.py:

python
"""A tiny JSON 'schema' enforced token by token: every output parses, whatever the model prefers."""
import json
import math

from examples.m03_constrained import prompt_for, trie_constraint
from examples.m03_setup import MODEL, TOK, token_probs
from supportdesk.data import CATEGORIES, PRIORITIES, load_tickets
from supportdesk.tinylm import SamplingParams, generate

# The schema {"category": one of 6, "priority": one of 4} has exactly 24 valid outputs.
VALID = [f' {{"category": "{c}", "priority": "{p}"}}' for c in CATEGORIES for p in PRIORITIES]
allowed = trie_constraint(VALID)
dev = load_tickets("dev")

free_ok = forced_ok = 0
for i, t in enumerate(dev):
    prompt = prompt_for(t).replace("Category:", "Triage JSON:")
    free = generate(MODEL, TOK, prompt, SamplingParams(max_new_tokens=24, temperature=1.0, seed=i))
    forced = generate(MODEL, TOK, prompt, SamplingParams(max_new_tokens=40, temperature=1.0, seed=i), allowed=allowed)
    for text, kind in ((free.text, "free"), (forced.text, "forced")):
        try:
            obj = json.loads(text)
            good = obj.get("category") in CATEGORIES and obj.get("priority") in PRIORITIES
        except (json.JSONDecodeError, AttributeError):
            good = False
        if kind == "free":
            free_ok += good
        else:
            forced_ok += good
    if i == 0:
        print("ticket 1 free:  ", repr(free.text[:60]))
        print("ticket 1 forced:", repr(forced.text), "| stop_reason:", forced.stop_reason)
        probs = token_probs(prompt, forced.token_ids)
        pieces = [TOK.decode([tid]) for tid in forced.token_ids]
        print("raw model probability of each forced token:")
        print("  " + "  ".join(f"{piece!r}:{p:.0e}" for piece, p in zip(pieces, probs)))
        print(f"  total log-probability {sum(math.log(p) for p in probs):.1f} over {len(probs)} tokens")

print(f"\nvalid triage JSON: free {free_ok}/{len(dev)}, constrained {forced_ok}/{len(dev)}")

Code explained

  • In simple words: list every valid JSON answer, let TinyLM spell only those, and compare with letting it write freely.
  • What happens: VALID holds 24 complete JSON strings. For each dev ticket the script samples freely and with the constraint, then checks with json.loads plus a field check. For ticket 1 it prints the raw model probability of every forced token.
  • Comes out:
text
ticket 1 free:   ' use one of each monthDara): Hi, how much does the Enterpris'
ticket 1 forced: ' {"category": "cancellation", "priority": "high"}' | stop_reason: constraint
raw model probability of each forced token:
  ' ':6e-05  '{':5e-05  '"':2e-04  'c':8e-05  'ate':2e-05  'g':5e-05  'or':4e-05  'y':5e-05  '"':4e-04  ':':1e-03  ' "':3e-04  'c':9e-05  'an':3e-05  'c':5e-05  'ell':3e-05  'ation':1e-04  '"':4e-04  ',':2e-04  ' "':2e-04  'p':7e-05  'ri':8e-05  'or':5e-05  'ity':7e-05  '"':3e-04  ':':2e-03  ' "':5e-04  'h':6e-05  'ig':2e-05  'h':7e-05  '"':7e-04  '}':7e-05
  total log-probability -282.0 over 31 tokens

valid triage JSON: free 0/48, constrained 48/48

Free generation never produced valid JSON (0/48); constrained generation always did (48/48). But look at the per-token probabilities: every structural token ({, ", category, :) had a raw probability between 0.00002 and 0.002. The whole object has a log-probability of -282, which is to say the model found this text astronomically unlikely. Notice also how the constraint tokenizes category as c, ate, g, or, y: the trie follows one fixed tokenization of each valid string, which may not be the one the model would naturally use. Production engines handle this properly by tracking which tokens are consistent with the grammar at the character level.

Real LLMs are nowhere near this bad at JSON; they have seen enormous amounts of it. The point transfers anyway: when a constraint forces tokens the model considers unlikely, the model is "writing in a straitjacket", and research has found that strict format restrictions can reduce reasoning quality on some tasks (Tam et al., "Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models", 2024). The practical pattern in Module 6 is to let the model reason in free text where needed and constrain only the final answer.

Logit bias: nudge, ban, or force

Logit bias adds a fixed number to specific token scores before sampling. A large negative value bans a token; a large positive value forces it; small values nudge. It is the bluntest constraint there is, because it applies at every step and knows nothing about context.

examples/m03_logit_bias.py:

python
"""Logit bias: nudge, ban, or force specific tokens before sampling."""
from examples.m03_setup import MODEL, TOK, UNCERTAIN, distribution, top_tokens
from supportdesk.tinylm import SamplingParams, generate

how = TOK.encode(" how").ids[0]
refund_q = TOK.encode(" can").ids[0]
print("token ids:", {" how": how, " can": refund_q})

for label, bias in [("no bias", {}), ("ban ' how' (-100)", {how: -100.0}),
                    ("nudge ' can' (+1)", {refund_q: 1.0}), ("force ' can' (+100)", {refund_q: 100.0})]:
    probs = distribution(UNCERTAIN, SamplingParams(temperature=1.0, logit_bias=bias))
    g = generate(MODEL, TOK, UNCERTAIN, SamplingParams(max_new_tokens=16, temperature=0, stop=["\n"], logit_bias=bias))
    print(f"{label:>20} | top 3 {top_tokens(probs, 3)} | greedy: {g.text!r}")

# The bias applies at EVERY step, not just the first one.
g = generate(MODEL, TOK, "Customer (Ana): Hi, how do I export my board data?\nAgent (Lena):",
             SamplingParams(max_new_tokens=30, temperature=0, stop=["\n"], logit_bias={TOK.encode(" CSV").ids[0]: -100.0}))
print("\nagent reply with ' CSV' banned:", repr(g.text))

Code explained

  • In simple words: push individual tokens up or down and watch both the next-token distribution and a greedy continuation change.
  • What happens: we look up token ids for how and can, then apply four biases: none, ban how (-100), nudge can (+1), force can (+100). Finally we ban CSV from an agent reply about exporting.
  • Comes out:
text
token ids: {' how': 417, ' can': 309}
             no bias | top 3 [(' how', 0.27), (' I', 0.153), (' can', 0.108)] | greedy: ' how do I export my invoice data?'
   ban ' how' (-100) | top 3 [(' I', 0.21), (' can', 0.147), (' where', 0.125)] | greedy: ' I was charged twice for the Free plan this month. Can you refund the duplicate'
   nudge ' can' (+1) | top 3 [(' can', 0.247), (' how', 0.228), (' I', 0.129)] | greedy: ' can I get a refund on my annual Team plan?'
 force ' can' (+100) | top 3 [(' can', 1.0), (' how', 0.0), (' I', 0.0)] | greedy: ' can can can can can can can can can can can can can can can can'

agent reply with ' CSV' banned: ' Export a board to purchase or renewal get a full refund. After 14 days they stay active until the end of the term.'

Banning how promotes the next option and yields a different memorized question. A +1 nudge is enough to flip the top choice from how (0.27) to can (0.25). Forcing can with +100 is a disaster: the bias applies at every step, so the model says can sixteen times. And banning CSV from the export answer does not produce a paraphrase; TinyLM's memorized sentence breaks and it splices in the refund policy. Logit bias is useful for single-token decisions (for example, restrict a yes/no classifier's first token to the ids for "yes" and "no") and for banning a specific token everywhere. For anything structured, use a real constraint.

Provider notes, checked 21 September 2026: logit bias works on token ids of the provider's own tokenizer, so ids from supportdesk.tokens.encode only apply to models using that encoding. Groq's documentation lists logit_bias as unsupported and returning a 400 error.

Tradeoffs between constraint and quality

Here is what our measurements support:

What we measuredUnconstrainedConstrainedTakeaway
Category parses (48 dev tickets)0/4848/48Constraints guarantee format
Category correct0/48 (no label at all)10 to 11/48 vs 12/48 majority baselineConstraints do not create knowledge
JSON valid (48 dev tickets)0/4848/48Same for structure
Probability the model gave the allowed tokensn/amedian 0.07 percent at the first stepLow mass means the constraint is doing all the work: route to review
SituationUse thisWhy
Output feeds code (triage JSON, tool arguments)Provider structured outputs with strict schema, or a grammar engine when self-hostingParsing never fails; retries and repair loops disappear
Choosing from a small fixed listScore each option (log-probabilities) or constrain to the listDeterministic, and the scores double as a confidence signal
Task needs reasoning before the answerFree-text reasoning field first, constrained answer last (Module 6)Strict formats can hurt reasoning quality
Ban a single token (a competitor's name, a profanity token)Logit bias of -100, where supportedCheap, but applies everywhere and only to exact token ids
Model has low mass on every allowed optionFlag for human review, do not trust the forced answerThe constraint turned a guess into something that looks confident

The production versions of this idea:

ToolWhat it isWhere you meet it
OpenAI-style structured outputs (response_format with json_schema, strict: true)Schema-constrained decoding on the provider's serversGroq supports strict: true on GPT-OSS 20B, GPT-OSS 120B, and Qwen 3.8 27B, with "100% schema adherence"; best-effort mode on more models
Gemini structured outputResponse JSON schema in the native API; response_format through the OpenAI-compatible endpointllm.chat(..., response_format=...) with LLM_PROVIDER=gemini
Ollama structured outputsA JSON schema passed as format (native) or response_formatLocal models such as qwen3:8b
Anthropic structured outputsJSON schema output and strict tool use on the Claude APIIf you add Anthropic as a provider
XGrammar (MLC)Fast grammar engine; compiles JSON schemas and grammars into token masksA structured-output backend in vLLM and SGLang
llguidance (guidance-ai)Low-latency grammar engine behind the guidance libraryvLLM's guidance backend; llama.cpp
Outlines (dottxt)Library for regex, JSON schema, and grammar-constrained generationSelf-hosted models via Transformers, vLLM, and others
llama.cpp GBNF grammarsA grammar format for constraining local modelsllama-server and tools built on llama.cpp

vLLM's structured output feature accepts choice, regex, json, grammar, and structural_tag constraints and picks a backend automatically by default. Module 6 builds the Brightlane triage schema on top of provider structured outputs.