CourseLarge Language Models · Module 2 : Tokens -context-cost · part 9 of 80
Part 9 · Module 2 : Tokens -context-cost

Part B: The context window

19 min read·22 Sept 2026

What fills it

The context window is the maximum number of tokens a model can attend to in one call, and it covers the input and the output together. For Brightlane's reply assistant, the input side has five parts, and the output needs its own reserve:

.

The budget equation is: system + tools + history + retrieved + user + output reserve + safety margin must not exceed the window. The output reserve is the max_tokens you pass; if you do not reserve it, a long input leaves no room for the answer. The safety margin covers the error of counting with a tokenizer that is not the model's (up to 11% in Part A) and template overhead.

python
# examples/m02_budget.py
"""Module 2: a context budget for one Brightlane assistant request, checked against several windows."""
import json
from dataclasses import dataclass, field

from m02_conversation import SYSTEM, conversation

from supportdesk.data import get_article
from supportdesk.kb_search import KBSearch
from supportdesk.tokens import count_messages, count_tokens

TOOLS = [
    {"type": "function", "function": {
        "name": "lookup_invoice", "description": "Fetch one invoice by id, with amount, date, and payment status.",
        "parameters": {"type": "object", "properties": {"invoice_id": {"type": "string", "pattern": "^INV-\\d{4}-\\d{6}$"}},
                       "required": ["invoice_id"]}}},
    {"type": "function", "function": {
        "name": "create_refund_request", "description": "Open a refund request for a human billing agent to approve.",
        "parameters": {"type": "object", "properties": {"invoice_id": {"type": "string"}, "reason": {"type": "string"},
                                                        "amount_usd": {"type": "number"}},
                       "required": ["invoice_id", "reason"]}}},
]


@dataclass
class ContextBudget:
    """window >= system + tools + history + retrieved + user + output_reserve + safety margin."""
    window: int
    output_reserve: int
    safety: float = 0.05  # our tokenizer is not the model's tokenizer: keep 5% slack
    parts: dict[str, int] = field(default_factory=dict)

    def add(self, name: str, tokens: int) -> None:
        self.parts[name] = self.parts.get(name, 0) + tokens

    @property
    def usable(self) -> int:
        return int(self.window * (1 - self.safety)) - self.output_reserve

    @property
    def used(self) -> int:
        return sum(self.parts.values())

    def report(self, label: str) -> str:
        status = "fits" if self.used <= self.usable else f"OVER by {self.used - self.usable}"
        return f"{label:28} window {self.window:>9,}  usable input {self.usable:>9,}  used {self.used:>6,}  {status}"


messages = conversation()
history, question = messages[1:-1], messages[-1]["content"]
hits = KBSearch().search("duplicate charge refund invoice", k=3)
retrieved = "\n\n".join(f"[{h.article_id}]\n{get_article(h.article_id).body}" for h in hits)

parts = {
    "system": count_tokens(SYSTEM) + 4,
    "tools": count_tokens(json.dumps(TOOLS)),
    "history": count_messages(history) - 3,
    "retrieved": count_tokens(retrieved) + 4,
    "user": count_tokens(question) + 4 + 3,
}
print("Request parts (o200k_base estimates):")
for name, n in parts.items():
    print(f"  {name:10} {n:5}")
print(f"  {'total':10} {sum(parts.values()):5}   retrieved articles: {[h.article_id for h in hits]}\n")

for label, window, reserve in [
    ("gpt-oss-120b on Groq", 131_072, 4_096),       # reasoning model: reserve room to think
    ("gemini-3.5-flash", 1_048_576, 4_096),
    ("qwen3:8b on Ollama, 4k", 4_096, 1_024),         # Ollama default below 24 GiB VRAM
    ("a 2k-token small model", 2_048, 1_024),
    ("TinyLM", 128, 40),
]:
    budget = ContextBudget(window=window, output_reserve=reserve)
    for name, n in parts.items():
        budget.add(name, n)
    print(budget.report(label))

per_message = parts["history"] / len(history)
fixed = sum(parts.values()) - parts["history"]
for label, usable in [("qwen3:8b on Ollama, 4k", ContextBudget(4_096, 1_024).usable),
                      ("gpt-oss-120b on Groq", ContextBudget(131_072, 4_096).usable)]:
    print(f"{label}: room for about {int((usable - fixed) / per_message):,} history messages "
          f"at {per_message:.0f} tokens each")

Code explained

  • In simple words: a calculator that adds up every part of one request and checks it against several models' windows.
  • What happens: TOOLS are two tool definitions in the Chat Completions format (Module 10 uses tools for real; here we count their JSON, which is roughly how providers bill them). ContextBudget holds the window, the output reserve, a 5% safety margin, and named parts; usable is the input room left. We measure each part with count_tokens and count_messages, retrieve help-center articles with KBSearch (Module 7), and check five targets: gpt-oss-120b on Groq (131,072 tokens), gemini-3.5-flash (1,048,576), qwen3:8b on Ollama at its default 4,096 on machines with under 24 GiB of GPU memory, a 2,048-token small model, and TinyLM (128). The last lines estimate how many more history messages fit.
  • Comes out:
text
  Request parts (o200k_base estimates):
    system        32
    tools        174
    history      505
    retrieved    218
    user          21
    total        950   retrieved articles: ['billing-refunds', 'billing-invoices']

  gpt-oss-120b on Groq         window   131,072  usable input   120,422  used    950  fits
  gemini-3.5-flash             window 1,048,576  usable input   992,051  used    950  fits
  qwen3:8b on Ollama, 4k       window     4,096  usable input     2,867  used    950  fits
  a 2k-token small model       window     2,048  usable input       921  used    950  OVER by 29
  TinyLM                       window       128  usable input        81  used    950  OVER by 869
  qwen3:8b on Ollama, 4k: room for about 105 history messages at 23 tokens each
  gpt-oss-120b on Groq: room for about 5,226 history messages at 23 tokens each

The history is already the biggest part (505 of 950 tokens) after 23 short turns, and it is the only part that grows without limit. On the hosted models this request is tiny. On Ollama's default 4,096 window, the chat can grow to about 105 messages before it overflows, which a real support chat can reach in a day. The 2,048-token model is already over by 29 tokens. Window sizes: Groq's model page for gpt-oss-120b lists 131,072 tokens of context and 65,536 max output; Google's model page lists 1,048,576 input tokens for gemini-3.5-flash; Ollama's documentation gives its default context as 4k below 24 GiB of VRAM, 32k from 24 to 48 GiB, and 256k above (all checked 21 Sep 2026).

SituationUse thisWhy
Deciding max_tokens for a reasoning modelReserve thinking plus answer (thousands)Reasoning tokens come out of the same output budget
Counting for a model whose tokenizer you lackYour count times the calibrated ratio, plus 5 to 10%Up to 11% disagreement between tokenizers here
Local model on OllamaSet the context length explicitly and budget to itThe default may be 4k even if the model supports far more
Hosted model with a huge windowBudget for cost and quality, not only fitFitting is not the same as being used well (next sections)

At and past the limit: TinyLM

TinyLM's window is 128 tokens, so we can watch overflow happen for real. Read the first line of generate() in supportdesk/tinylm.py: ids = tokenizer.encode(prompt).ids[-(model.cfg.context - params.max_new_tokens):]. It keeps only the newest prompt tokens, leaving room for the output.

python
# examples/m02_tinylm_limit.py
"""Module 2: what TinyLM does with a prompt longer than its 128-token context window."""
import torch

from supportdesk.tinylm import SamplingParams, generate, load

torch.manual_seed(0)
torch.set_num_threads(2)
model, tok = load()
window = model.cfg.context

prompt = (
    "Customer (Ana): Hi, I was charged twice for invoice INV-2026-004512 on the Business plan. Please refund it.\n"
    "Agent (Omar): Sorry about that. A billing agent will review the duplicate charge.\n"
    "Customer (Ana): Also, Slack notifications stopped arriving in our channel.\n"
    "Agent (Omar): Choose the channel again in the Slack integration settings.\n"
    "Customer (Ana): Thanks, that worked. How do I export a board to CSV?\n"
    "Agent (Omar): Open the board and choose More, then Export, then CSV.\n"
    "Customer (Ana): And where are we on my refund for the duplicate charge?\n"
    "Agent (Omar):"
)
ids = tok.encode(prompt).ids
params = SamplingParams(max_new_tokens=40, temperature=0)
kept = ids[-(window - params.max_new_tokens):]  # the same slice generate() applies
print(f"prompt tokens: {len(ids)}   window: {window}   reserved for output: {params.max_new_tokens}")
print(f"kept: last {len(kept)} tokens, dropped: first {len(ids) - len(kept)} tokens")
print(f"dropped text: {tok.decode(ids[:len(ids) - len(kept)])!r}")
print(f"kept text starts: {tok.decode(kept)[:60]!r}")
print(f"invoice id still in view: {'INV-2026-004512' in tok.decode(kept)}")

out = generate(model, tok, prompt, params)
print(f"\nTinyLM continues ({out.stop_reason}, {len(out.token_ids)} tokens): {out.text.strip()[:160]!r}")

print("\nAsking for the whole window as output:")
try:
    generate(model, tok, prompt, SamplingParams(max_new_tokens=window, temperature=0))
except IndexError as err:
    print(f"  IndexError: {err}")

Code explained

  • In simple words: we hand TinyLM a conversation that is too long and look at exactly which part it never sees.
  • What happens: the prompt is a support chat in TinyLM's training format. We apply the same slice generate() uses, print what was dropped and kept, then let TinyLM continue with greedy decoding (temperature=0, taught in Module 3). Finally we ask for max_new_tokens=128, the whole window, and catch the result.
  • Comes out:

text
  prompt tokens: 154   window: 128   reserved for output: 40
  kept: last 88 tokens, dropped: first 66 tokens
  dropped text: 'Customer (Ana): Hi, I was charged twice for invoice INV-2026-004512 on the Business plan. Please refund it.\nAgent (Omar): Sorry about that. A billing agent will review the duplicate charge.\nCustomer (Ana): Also, Slack notifications stopped arriving in our channel.\n'
  kept text starts: 'Agent (Omar): Choose the channel again in the Slack integrat'
  invoice id still in view: False

  TinyLM continues (length, 40 tokens): 'Sorry about the sign-in page to 10 business days to the original payment method.\nTwo-factor authentication (2FA) recovery: Team costs 12 USD per user per user p'

  Asking for the whole window as output:
    IndexError: index out of range in self

The prompt is 154 tokens; with 40 reserved for output, only the last 88 are kept, so the first 66 tokens (the invoice id, the plan, and the refund request) are silently gone. Nothing in the returned Generation says truncation happened. TinyLM then produces fluent support-desk text stitched from memorized replies ("to 10 business days to the original payment method"), which sounds relevant but is not grounded in the dropped facts. This is the worst kind of failure: no error, a plausible answer, missing information.

The last line is a real bug in the canonical generate(): when max_new_tokens is 128 or more, the slice ids[-(128 - 128):] becomes ids[-0:], which in Python means the whole list, and the position embedding lookup fails with IndexError: index out of range in self. The workaround is to keep max_new_tokens below model.cfg.context (the module's test pins this behavior); the file is canonical, so we do not edit it here.

At and past the limit: providers

Hosted providers behave in one of three ways at the limit, and your code needs a plan for each:

BehaviorWhere you see itWhat to do
Reject the request with HTTP 400Hosted APIs such as Groq, Gemini, OpenAI when input plus max_tokens exceeds the windowTrim before sending; catch the error, trim harder, retry once
Silently drop the oldest tokensLocal servers such as Ollama when the prompt exceeds the configured contextBudget before sending; you get no error, only worse answers
Stop the output early with finish_reason == "length"Any provider when the reply reaches max_tokensCheck finish_reason; raise max_tokens or ask for less

The wording of overflow errors differs by provider, so detect them by code first and message second, and log the raw error body the first time you meet a new provider.

python
# examples/m02_overflow.py
"""Module 2: stay inside the context window before sending, and recover if the provider still says no."""
import httpx
import openai
from m02_context import drop_oldest
from m02_conversation import conversation

from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages

OVERFLOW_HINTS = ("context_length_exceeded", "context length", "maximum context", "too many tokens",
                  "exceeds the maximum", "reduce the length")


def is_context_overflow(err: Exception) -> bool:
    """Providers word this differently; check the code first, then the message."""
    if not isinstance(err, openai.BadRequestError):
        return False
    text = f"{getattr(err, 'code', '')} {err.message} {err.body}".lower()
    return any(hint in text for hint in OVERFLOW_HINTS)


def send_within_window(messages, llm, window: int, max_tokens: int, safety: float = 0.05):
    """Trim to fit before sending; on an overflow error, trim 20% harder and retry once."""
    budget = int(window * (1 - safety)) - max_tokens
    for attempt in (1, 2):
        trimmed = drop_oldest(messages, budget)
        print(f"  attempt {attempt}: sending {count_messages(trimmed)} estimated tokens "
              f"({len(messages) - len(trimmed)} messages dropped)")
        try:
            result = llm(trimmed, max_tokens=max_tokens)
        except openai.BadRequestError as err:
            if not is_context_overflow(err) or attempt == 2:
                raise
            print(f"  provider refused: {err.body['message']}")
            budget = int(budget * 0.8)
            continue
        if result.finish_reason == "length":
            print("  warning: the reply hit max_tokens and is cut off")
        return result


def fake_provider(limit: int, tokenizer_ratio: float):
    """A ScriptedLLM that behaves like a provider whose tokenizer counts more tokens than ours (plumbing only)."""
    def respond(messages, kwargs):
        real = int(count_messages(messages) * tokenizer_ratio)
        if real + kwargs["max_tokens"] > limit:
            request = httpx.Request("POST", "https://example.invalid/v1/chat/completions")
            raise openai.BadRequestError(
                "Error code: 400", response=httpx.Response(400, request=request),
                body={"message": f"This model's maximum context length is {limit} tokens. "
                                 f"Your request used {real + kwargs['max_tokens']} tokens.",
                      "type": "invalid_request_error", "code": "context_length_exceeded"})
        return "(scripted reply: request accepted)"
    return ScriptedLLM(responder=respond)


chat_log = conversation()
print(f"Conversation: {count_messages(chat_log)} estimated tokens")
print("Provider A (tokenizer counts like ours):")
reply = send_within_window(chat_log, fake_provider(limit=512, tokenizer_ratio=1.0), window=512, max_tokens=128)
print(f"  reply: {reply.text!r}")
print("Provider B (tokenizer counts 30% more than ours):")
reply = send_within_window(chat_log, fake_provider(limit=512, tokenizer_ratio=1.3), window=512, max_tokens=128)
print(f"  reply: {reply.text!r}")

# With a real provider the call is the same; only the window changes, for example:
#   from supportdesk.llm import chat
#   send_within_window(chat_log, chat, window=131_072, max_tokens=1_024)   # gpt-oss-120b on Groq

Code explained

  • In simple words: a guard that trims the chat to fit before sending and, if the provider still refuses, trims harder and tries once more.
  • What happens: is_context_overflow() accepts only openai.BadRequestError (HTTP 400 from any OpenAI-compatible endpoint) and looks for the standard context_length_exceeded code or common phrases. send_within_window() computes the input budget (window minus safety margin minus max_tokens), trims with drop_oldest() (Part C), calls the model, and on an overflow error retries once with a budget 20% smaller. It also warns when finish_reason is length. fake_provider() is a ScriptedLLM that raises a real openai.BadRequestError when its own count exceeds the limit; with tokenizer_ratio=1.3 it simulates a model whose tokenizer counts 30% more than ours. It tests plumbing, not a model.
  • Comes out:
text
  Conversation: 558 estimated tokens
  Provider A (tokenizer counts like ours):
    attempt 1: sending 336 estimated tokens (9 messages dropped)
    reply: '(scripted reply: request accepted)'
  Provider B (tokenizer counts 30% more than ours):
    attempt 1: sending 336 estimated tokens (9 messages dropped)
    provider refused: This model's maximum context length is 512 tokens. Your request used 564 tokens.
    attempt 2: sending 272 estimated tokens (12 messages dropped)
    reply: '(scripted reply: request accepted)'

Provider A accepts the trimmed request. Provider B rejects it (564 tokens by its count against 512), and the retry with a tighter budget succeeds. Notice what the retry cost: 12 of 23 history messages dropped instead of 9. A calibrated safety margin avoids the round trip. With a real key, the same function works with supportdesk.llm.chat, as the comment at the end of the file shows. Illustrative example of a provider's overflow error body (not captured in this build; wording varies by provider): {"message": "Please reduce the length of the messages or completion.", "type": "invalid_request_error", "param": "messages", "code": "context_length_exceeded"}.

Position effects: primacy, recency, and lost in the middle

Fitting in the window does not mean the model uses every token equally. Three effects are well documented:

  • Primacy: information at the start of the context is used relatively well.
  • Recency: information at the end, near the question, is used best of all.
  • Lost in the middle: information in the middle of a long context is used worst.

The standard reference is Liu et al., "Lost in the Middle: How Language Models Use Long Contexts" (arXiv 2307.03172, 2023; published in TACL in 2024). In multi-document question answering, they moved the one document containing the answer among up to 30 distractors. Accuracy traced a U shape over position: highest when the answer document was first or last, lowest in the middle. For GPT-3.5-Turbo the drop was more than 20 points, and in the worst case, performance with 20 or 30 documents fell below its closed-book accuracy of 56.1%, meaning the model did better with no documents at all than with the answer buried in the middle. The effect also appeared in models built for long contexts. Newer models have reduced it, but no one guarantees it is gone for your model and your data, so measure.

Illustration: A line chart with the position of the relevant document (first to last) on the x-axis and answer accuracy on the y-axis, showing a U-shaped curve that is high at both ends and dips in the middle, with a horizontal dashed line for closed-book accuracy crossing above the dip. | Alt text: U-shaped accuracy curve over document position | File: m02-lost-in-the-middle.png

For Brightlane, the practical rules are: put the instructions that matter most at the start (system prompt) and restate the specific task right before the output (the last user message); put the retrieved article most likely to answer the question closest to the question; and do not assume a fact in the middle of a 40-message history will be used just because it is present.

A needle-in-a-haystack harness

A needle-in-a-haystack test hides one fact (the needle) at a controlled depth inside filler text (the haystack) and asks for it. It is a narrow test, but it is cheap and it catches gross position failures in your own stack. This harness builds haystacks from Brightlane help-center paragraphs, hides a random approval code at five depths and three lengths, and scores exact matches.

python
# examples/m02_needle.py
"""Module 2: a needle-in-a-haystack harness for position effects.

Run with a stand-in (default) to test the plumbing, or with a real model:
  PYTHONPATH=. python examples/m02_needle.py --real
"""
import argparse
import random
import re

from supportdesk.data import load_articles
from supportdesk.llm import Usage, chat, resolve
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_tokens

DEPTHS = (0.0, 0.25, 0.5, 0.75, 1.0)
QUESTION = "What is the refund approval code for workspace {ws}? Reply with the code only."
NEEDLE = "Internal billing note: the refund approval code for workspace {ws} is {code}."


def build_haystack(target_tokens: int, rng: random.Random) -> list[str]:
    """Help-center paragraphs, shuffled and repeated until the text reaches about target_tokens."""
    paragraphs = [p for a in load_articles() for p in a.body.split("\n") if p.strip()]
    out, total = [], 0
    while total < target_tokens:
        p = rng.choice(paragraphs)
        out.append(p)
        total += count_tokens(p) + 1
    return out


def make_case(target_tokens: int, depth: float, rng: random.Random) -> tuple[list[dict], str]:
    ws, code = f"ACME-{rng.randint(10, 99)}", f"PLUM-{rng.randint(1000, 9999)}"
    paragraphs = build_haystack(target_tokens, rng)
    paragraphs.insert(round(depth * len(paragraphs)), NEEDLE.format(ws=ws, code=code))
    context = "\n".join(paragraphs)
    messages = [
        {"role": "system", "content": "Answer using only the reference text."},
        {"role": "user", "content": f"Reference text:\n{context}\n\n{QUESTION.format(ws=ws)}"},
    ]
    return messages, code


def run(llm, lengths, trials: int, seed: int = 0) -> dict[tuple[int, float], float]:
    rng = random.Random(seed)
    scores = {}
    for length in lengths:
        for depth in DEPTHS:
            hits = 0
            for _ in range(trials):
                messages, code = make_case(length, depth, rng)
                # Reasoning models spend max_tokens on thinking first, so leave room beyond the short answer.
                hits += code in llm(messages, max_tokens=512, temperature=0).text
            scores[(length, depth)] = hits / trials
    return scores


def perfect_reader(messages, kwargs):
    """Stand-in that finds the code with a regex: it can only prove the plumbing works."""
    found = re.search(r"code for workspace \S+ is (PLUM-\d+)", messages[-1]["content"])
    return found.group(1) if found else "not found"


def head_and_tail_reader(messages, kwargs, keep: float = 0.3):
    """Stand-in with a planted blind spot: it only reads the first and last 30% of the text."""
    text = messages[-1]["content"]
    cut = int(len(text) * keep)
    return perfect_reader([{"role": "user", "content": text[:cut] + text[-cut:]}], kwargs)


def show(title: str, scores: dict, trials: int) -> None:
    lengths = sorted({k[0] for k in scores})
    print(f"{title}\n  share answered correctly, n = {trials} per cell")
    print(f"  {'tokens':>7}" + "".join(f"{f'depth {d:.2f}':>12}" for d in DEPTHS))
    for length in lengths:
        print(f"  {length:>7}" + "".join(f"{scores[(length, d)]:>12.2f}" for d in DEPTHS))


if __name__ == "__main__":
    parser = argparse.ArgumentParser()
    parser.add_argument("--real", action="store_true", help="call supportdesk.llm.chat (needs a provider key)")
    parser.add_argument("--trials", type=int, default=5)
    args = parser.parse_args()
    lengths = (2_000, 8_000, 32_000)
    if args.real:
        calls = len(lengths) * len(DEPTHS) * args.trials
        input_tokens = sum(lengths) * len(DEPTHS) * args.trials
        _, model = resolve()
        estimate = cost_usd(Usage(input_tokens=input_tokens, output_tokens=calls * 512), model) if model in PRICES else float("nan")
        print(f"{calls} calls to {model}, about {input_tokens:,} input tokens, at most {estimate:.2f} USD")
        show("Real model (supportdesk.llm.chat):", run(chat, lengths, args.trials), args.trials)
    else:
        show("Stand-in: perfect regex reader (plumbing check, not a model)",
             run(ScriptedLLM(responder=perfect_reader), lengths, args.trials), args.trials)
        show("Stand-in: planted head-and-tail blind spot (does the harness detect it?)",
             run(ScriptedLLM(responder=head_and_tail_reader), lengths, args.trials), args.trials)

Code explained

  • In simple words: hide a code at the start, quarter, middle, three quarters, and end of a long reference text, ask for it, and tabulate how often the answer is right.
  • What happens: build_haystack() draws shuffled help-center paragraphs until the text reaches the target token count. make_case() generates a random workspace and code (so no answer can be memorized), inserts the needle at the chosen depth, and builds the messages. run() loops over lengths, depths, and trials and records the share of correct answers. max_tokens=512 leaves room for a reasoning model's hidden thinking; with a tiny max_tokens, a reasoning model can spend the whole budget thinking and return an empty answer. Two stand-ins validate the harness: perfect_reader finds the code with a regex, and head_and_tail_reader has a planted blind spot (it reads only the first and last 30% of the text). With --real, the same harness calls llm.chat and first prints the number of calls, input tokens, and a cost ceiling.
  • Comes out (stand-ins, captured):
text
  Stand-in: perfect regex reader (plumbing check, not a model)
    share answered correctly, n = 5 per cell
     tokens  depth 0.00  depth 0.25  depth 0.50  depth 0.75  depth 1.00
       2000        1.00        1.00        1.00        1.00        1.00
       8000        1.00        1.00        1.00        1.00        1.00
      32000        1.00        1.00        1.00        1.00        1.00
  Stand-in: planted head-and-tail blind spot (does the harness detect it?)
    share answered correctly, n = 5 per cell
     tokens  depth 0.00  depth 0.25  depth 0.50  depth 0.75  depth 1.00
       2000        1.00        1.00        0.00        1.00        1.00
       8000        1.00        1.00        0.00        1.00        1.00
      32000        1.00        1.00        0.00        1.00        1.00

The perfect reader scores 1.00 everywhere, so the plumbing works (needles are inserted, codes are checked). The planted blind spot shows up exactly where it should, at depth 0.50, at all three lengths. That second check matters: a harness you have never seen fail cannot be trusted to detect failure. Both tables are stand-in output, not model behavior.

With --real and the default Groq model, the preamble reads 75 calls to openai/gpt-oss-120b, about 1,050,000 input tokens, at most 0.18 USD (computed, not sent). Illustrative sample run (not captured in this build; produced for teaching). Your output will differ:

text
     tokens  depth 0.00  depth 0.25  depth 0.50  depth 0.75  depth 1.00
       2000        1.00        1.00        1.00        1.00        1.00
       8000        1.00        1.00        1.00        1.00        1.00
      32000        1.00        1.00        0.80        1.00        1.00

How to read a real result: with 5 trials per cell, one miss moves a cell by 0.20, and a 95% interval around 4 of 5 runs from roughly 0.38 to 0.96, so a single 0.80 is within noise. Raise --trials to 20 or more before you conclude anything, and remember that current frontier models often score near 1.00 on this literal-match test even when they struggle on harder long-context tasks (next section).

Effective context versus advertised context

The advertised context is the largest input the API accepts. The effective context is the length up to which the model still performs your task well. They differ, often a lot.

  • RULER (Hsieh et al., NVIDIA, arXiv 2404.06654, 2024) tested 17 long-context models on 13 tasks that go beyond single-needle retrieval (multiple needles, tracing variables through the text, aggregating counts, question answering). Models scored nearly perfectly on the vanilla needle test, yet performance dropped sharply as length grew, and only about half of the models that claimed 32K tokens or more kept satisfactory performance at 32K.
  • NoLiMa (Modarressi et al., arXiv 2502.05167, 2025) removed the literal word overlap between question and needle, so the model has to make a one-step association instead of matching words. Of 13 models claiming at least 128K tokens, 11 fell below half of their short-context score at 32K; GPT-4o dropped from 99.3% to 69.7%.
  • Context Rot (Hong, Troynikov, and Huber, Chroma technical report, July 2025) tested 18 models including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3, and found performance became increasingly unreliable as input grew, even on simple tasks, with distractors and haystack structure changing the results.

These results are for the models they tested at the time; newer models improve. The durable lesson is methodological: advertised length says what fits, not what works. Measure effective context for your task at your lengths, and treat the smallest context that contains what the model needs as the default.

SituationUse thisWhy
Deciding where to put a critical instructionStart of the system prompt, restated near the endPrimacy and recency; the middle is weakest
Several retrieved articlesMost relevant last, next to the questionRecency helps the article that matters most
Choosing a model for long inputsYour own harness at your lengths, 20+ trials per cellAdvertised length and needle scores overstate usable length
Long input that is mostly irrelevantRetrieve or compact first (Part C)Distractors hurt even when everything fits