Part B: The context window
What fills it
The context window is the maximum number of tokens a model can attend to in one call, and it covers the input and the output together. For Brightlane's reply assistant, the input side has five parts, and the output needs its own reserve:
The budget equation is: system + tools + history + retrieved + user + output reserve + safety margin must not exceed the window. The output reserve is the max_tokens you pass; if you do not reserve it, a long input leaves no room for the answer. The safety margin covers the error of counting with a tokenizer that is not the model's (up to 11% in Part A) and template overhead.
# examples/m02_budget.py
"""Module 2: a context budget for one Brightlane assistant request, checked against several windows."""
import json
from dataclasses import dataclass, field
from m02_conversation import SYSTEM, conversation
from supportdesk.data import get_article
from supportdesk.kb_search import KBSearch
from supportdesk.tokens import count_messages, count_tokens
TOOLS = [
{"type": "function", "function": {
"name": "lookup_invoice", "description": "Fetch one invoice by id, with amount, date, and payment status.",
"parameters": {"type": "object", "properties": {"invoice_id": {"type": "string", "pattern": "^INV-\\d{4}-\\d{6}$"}},
"required": ["invoice_id"]}}},
{"type": "function", "function": {
"name": "create_refund_request", "description": "Open a refund request for a human billing agent to approve.",
"parameters": {"type": "object", "properties": {"invoice_id": {"type": "string"}, "reason": {"type": "string"},
"amount_usd": {"type": "number"}},
"required": ["invoice_id", "reason"]}}},
]
@dataclass
class ContextBudget:
"""window >= system + tools + history + retrieved + user + output_reserve + safety margin."""
window: int
output_reserve: int
safety: float = 0.05 # our tokenizer is not the model's tokenizer: keep 5% slack
parts: dict[str, int] = field(default_factory=dict)
def add(self, name: str, tokens: int) -> None:
self.parts[name] = self.parts.get(name, 0) + tokens
@property
def usable(self) -> int:
return int(self.window * (1 - self.safety)) - self.output_reserve
@property
def used(self) -> int:
return sum(self.parts.values())
def report(self, label: str) -> str:
status = "fits" if self.used <= self.usable else f"OVER by {self.used - self.usable}"
return f"{label:28} window {self.window:>9,} usable input {self.usable:>9,} used {self.used:>6,} {status}"
messages = conversation()
history, question = messages[1:-1], messages[-1]["content"]
hits = KBSearch().search("duplicate charge refund invoice", k=3)
retrieved = "\n\n".join(f"[{h.article_id}]\n{get_article(h.article_id).body}" for h in hits)
parts = {
"system": count_tokens(SYSTEM) + 4,
"tools": count_tokens(json.dumps(TOOLS)),
"history": count_messages(history) - 3,
"retrieved": count_tokens(retrieved) + 4,
"user": count_tokens(question) + 4 + 3,
}
print("Request parts (o200k_base estimates):")
for name, n in parts.items():
print(f" {name:10} {n:5}")
print(f" {'total':10} {sum(parts.values()):5} retrieved articles: {[h.article_id for h in hits]}\n")
for label, window, reserve in [
("gpt-oss-120b on Groq", 131_072, 4_096), # reasoning model: reserve room to think
("gemini-3.5-flash", 1_048_576, 4_096),
("qwen3:8b on Ollama, 4k", 4_096, 1_024), # Ollama default below 24 GiB VRAM
("a 2k-token small model", 2_048, 1_024),
("TinyLM", 128, 40),
]:
budget = ContextBudget(window=window, output_reserve=reserve)
for name, n in parts.items():
budget.add(name, n)
print(budget.report(label))
per_message = parts["history"] / len(history)
fixed = sum(parts.values()) - parts["history"]
for label, usable in [("qwen3:8b on Ollama, 4k", ContextBudget(4_096, 1_024).usable),
("gpt-oss-120b on Groq", ContextBudget(131_072, 4_096).usable)]:
print(f"{label}: room for about {int((usable - fixed) / per_message):,} history messages "
f"at {per_message:.0f} tokens each")Code explained
- In simple words: a calculator that adds up every part of one request and checks it against several models' windows.
- What happens:
TOOLSare two tool definitions in the Chat Completions format (Module 10 uses tools for real; here we count their JSON, which is roughly how providers bill them).ContextBudgetholds the window, the output reserve, a 5% safety margin, and named parts;usableis the input room left. We measure each part withcount_tokensandcount_messages, retrieve help-center articles withKBSearch(Module 7), and check five targets: gpt-oss-120b on Groq (131,072 tokens), gemini-3.5-flash (1,048,576), qwen3:8b on Ollama at its default 4,096 on machines with under 24 GiB of GPU memory, a 2,048-token small model, and TinyLM (128). The last lines estimate how many more history messages fit. - Comes out:
Request parts (o200k_base estimates):
system 32
tools 174
history 505
retrieved 218
user 21
total 950 retrieved articles: ['billing-refunds', 'billing-invoices']
gpt-oss-120b on Groq window 131,072 usable input 120,422 used 950 fits
gemini-3.5-flash window 1,048,576 usable input 992,051 used 950 fits
qwen3:8b on Ollama, 4k window 4,096 usable input 2,867 used 950 fits
a 2k-token small model window 2,048 usable input 921 used 950 OVER by 29
TinyLM window 128 usable input 81 used 950 OVER by 869
qwen3:8b on Ollama, 4k: room for about 105 history messages at 23 tokens each
gpt-oss-120b on Groq: room for about 5,226 history messages at 23 tokens eachThe history is already the biggest part (505 of 950 tokens) after 23 short turns, and it is the only part that grows without limit. On the hosted models this request is tiny. On Ollama's default 4,096 window, the chat can grow to about 105 messages before it overflows, which a real support chat can reach in a day. The 2,048-token model is already over by 29 tokens. Window sizes: Groq's model page for gpt-oss-120b lists 131,072 tokens of context and 65,536 max output; Google's model page lists 1,048,576 input tokens for gemini-3.5-flash; Ollama's documentation gives its default context as 4k below 24 GiB of VRAM, 32k from 24 to 48 GiB, and 256k above (all checked 21 Sep 2026).
| Situation | Use this | Why |
|---|---|---|
Deciding max_tokens for a reasoning model | Reserve thinking plus answer (thousands) | Reasoning tokens come out of the same output budget |
| Counting for a model whose tokenizer you lack | Your count times the calibrated ratio, plus 5 to 10% | Up to 11% disagreement between tokenizers here |
| Local model on Ollama | Set the context length explicitly and budget to it | The default may be 4k even if the model supports far more |
| Hosted model with a huge window | Budget for cost and quality, not only fit | Fitting is not the same as being used well (next sections) |
At and past the limit: TinyLM
TinyLM's window is 128 tokens, so we can watch overflow happen for real. Read the first line of generate() in supportdesk/tinylm.py: ids = tokenizer.encode(prompt).ids[-(model.cfg.context - params.max_new_tokens):]. It keeps only the newest prompt tokens, leaving room for the output.
# examples/m02_tinylm_limit.py
"""Module 2: what TinyLM does with a prompt longer than its 128-token context window."""
import torch
from supportdesk.tinylm import SamplingParams, generate, load
torch.manual_seed(0)
torch.set_num_threads(2)
model, tok = load()
window = model.cfg.context
prompt = (
"Customer (Ana): Hi, I was charged twice for invoice INV-2026-004512 on the Business plan. Please refund it.\n"
"Agent (Omar): Sorry about that. A billing agent will review the duplicate charge.\n"
"Customer (Ana): Also, Slack notifications stopped arriving in our channel.\n"
"Agent (Omar): Choose the channel again in the Slack integration settings.\n"
"Customer (Ana): Thanks, that worked. How do I export a board to CSV?\n"
"Agent (Omar): Open the board and choose More, then Export, then CSV.\n"
"Customer (Ana): And where are we on my refund for the duplicate charge?\n"
"Agent (Omar):"
)
ids = tok.encode(prompt).ids
params = SamplingParams(max_new_tokens=40, temperature=0)
kept = ids[-(window - params.max_new_tokens):] # the same slice generate() applies
print(f"prompt tokens: {len(ids)} window: {window} reserved for output: {params.max_new_tokens}")
print(f"kept: last {len(kept)} tokens, dropped: first {len(ids) - len(kept)} tokens")
print(f"dropped text: {tok.decode(ids[:len(ids) - len(kept)])!r}")
print(f"kept text starts: {tok.decode(kept)[:60]!r}")
print(f"invoice id still in view: {'INV-2026-004512' in tok.decode(kept)}")
out = generate(model, tok, prompt, params)
print(f"\nTinyLM continues ({out.stop_reason}, {len(out.token_ids)} tokens): {out.text.strip()[:160]!r}")
print("\nAsking for the whole window as output:")
try:
generate(model, tok, prompt, SamplingParams(max_new_tokens=window, temperature=0))
except IndexError as err:
print(f" IndexError: {err}")Code explained
- In simple words: we hand TinyLM a conversation that is too long and look at exactly which part it never sees.
- What happens: the prompt is a support chat in TinyLM's training format. We apply the same slice
generate()uses, print what was dropped and kept, then let TinyLM continue with greedy decoding (temperature=0, taught in Module 3). Finally we ask formax_new_tokens=128, the whole window, and catch the result. - Comes out:
prompt tokens: 154 window: 128 reserved for output: 40
kept: last 88 tokens, dropped: first 66 tokens
dropped text: 'Customer (Ana): Hi, I was charged twice for invoice INV-2026-004512 on the Business plan. Please refund it.\nAgent (Omar): Sorry about that. A billing agent will review the duplicate charge.\nCustomer (Ana): Also, Slack notifications stopped arriving in our channel.\n'
kept text starts: 'Agent (Omar): Choose the channel again in the Slack integrat'
invoice id still in view: False
TinyLM continues (length, 40 tokens): 'Sorry about the sign-in page to 10 business days to the original payment method.\nTwo-factor authentication (2FA) recovery: Team costs 12 USD per user per user p'
Asking for the whole window as output:
IndexError: index out of range in selfThe prompt is 154 tokens; with 40 reserved for output, only the last 88 are kept, so the first 66 tokens (the invoice id, the plan, and the refund request) are silently gone. Nothing in the returned Generation says truncation happened. TinyLM then produces fluent support-desk text stitched from memorized replies ("to 10 business days to the original payment method"), which sounds relevant but is not grounded in the dropped facts. This is the worst kind of failure: no error, a plausible answer, missing information.
The last line is a real bug in the canonical generate(): when max_new_tokens is 128 or more, the slice ids[-(128 - 128):] becomes ids[-0:], which in Python means the whole list, and the position embedding lookup fails with IndexError: index out of range in self. The workaround is to keep max_new_tokens below model.cfg.context (the module's test pins this behavior); the file is canonical, so we do not edit it here.
At and past the limit: providers
Hosted providers behave in one of three ways at the limit, and your code needs a plan for each:
| Behavior | Where you see it | What to do |
|---|---|---|
| Reject the request with HTTP 400 | Hosted APIs such as Groq, Gemini, OpenAI when input plus max_tokens exceeds the window | Trim before sending; catch the error, trim harder, retry once |
| Silently drop the oldest tokens | Local servers such as Ollama when the prompt exceeds the configured context | Budget before sending; you get no error, only worse answers |
Stop the output early with finish_reason == "length" | Any provider when the reply reaches max_tokens | Check finish_reason; raise max_tokens or ask for less |
The wording of overflow errors differs by provider, so detect them by code first and message second, and log the raw error body the first time you meet a new provider.
# examples/m02_overflow.py
"""Module 2: stay inside the context window before sending, and recover if the provider still says no."""
import httpx
import openai
from m02_context import drop_oldest
from m02_conversation import conversation
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages
OVERFLOW_HINTS = ("context_length_exceeded", "context length", "maximum context", "too many tokens",
"exceeds the maximum", "reduce the length")
def is_context_overflow(err: Exception) -> bool:
"""Providers word this differently; check the code first, then the message."""
if not isinstance(err, openai.BadRequestError):
return False
text = f"{getattr(err, 'code', '')} {err.message} {err.body}".lower()
return any(hint in text for hint in OVERFLOW_HINTS)
def send_within_window(messages, llm, window: int, max_tokens: int, safety: float = 0.05):
"""Trim to fit before sending; on an overflow error, trim 20% harder and retry once."""
budget = int(window * (1 - safety)) - max_tokens
for attempt in (1, 2):
trimmed = drop_oldest(messages, budget)
print(f" attempt {attempt}: sending {count_messages(trimmed)} estimated tokens "
f"({len(messages) - len(trimmed)} messages dropped)")
try:
result = llm(trimmed, max_tokens=max_tokens)
except openai.BadRequestError as err:
if not is_context_overflow(err) or attempt == 2:
raise
print(f" provider refused: {err.body['message']}")
budget = int(budget * 0.8)
continue
if result.finish_reason == "length":
print(" warning: the reply hit max_tokens and is cut off")
return result
def fake_provider(limit: int, tokenizer_ratio: float):
"""A ScriptedLLM that behaves like a provider whose tokenizer counts more tokens than ours (plumbing only)."""
def respond(messages, kwargs):
real = int(count_messages(messages) * tokenizer_ratio)
if real + kwargs["max_tokens"] > limit:
request = httpx.Request("POST", "https://example.invalid/v1/chat/completions")
raise openai.BadRequestError(
"Error code: 400", response=httpx.Response(400, request=request),
body={"message": f"This model's maximum context length is {limit} tokens. "
f"Your request used {real + kwargs['max_tokens']} tokens.",
"type": "invalid_request_error", "code": "context_length_exceeded"})
return "(scripted reply: request accepted)"
return ScriptedLLM(responder=respond)
chat_log = conversation()
print(f"Conversation: {count_messages(chat_log)} estimated tokens")
print("Provider A (tokenizer counts like ours):")
reply = send_within_window(chat_log, fake_provider(limit=512, tokenizer_ratio=1.0), window=512, max_tokens=128)
print(f" reply: {reply.text!r}")
print("Provider B (tokenizer counts 30% more than ours):")
reply = send_within_window(chat_log, fake_provider(limit=512, tokenizer_ratio=1.3), window=512, max_tokens=128)
print(f" reply: {reply.text!r}")
# With a real provider the call is the same; only the window changes, for example:
# from supportdesk.llm import chat
# send_within_window(chat_log, chat, window=131_072, max_tokens=1_024) # gpt-oss-120b on GroqCode explained
- In simple words: a guard that trims the chat to fit before sending and, if the provider still refuses, trims harder and tries once more.
- What happens:
is_context_overflow()accepts onlyopenai.BadRequestError(HTTP 400 from any OpenAI-compatible endpoint) and looks for the standardcontext_length_exceededcode or common phrases.send_within_window()computes the input budget (window minus safety margin minusmax_tokens), trims withdrop_oldest()(Part C), calls the model, and on an overflow error retries once with a budget 20% smaller. It also warns whenfinish_reasonislength.fake_provider()is aScriptedLLMthat raises a realopenai.BadRequestErrorwhen its own count exceeds the limit; withtokenizer_ratio=1.3it simulates a model whose tokenizer counts 30% more than ours. It tests plumbing, not a model. - Comes out:
Conversation: 558 estimated tokens
Provider A (tokenizer counts like ours):
attempt 1: sending 336 estimated tokens (9 messages dropped)
reply: '(scripted reply: request accepted)'
Provider B (tokenizer counts 30% more than ours):
attempt 1: sending 336 estimated tokens (9 messages dropped)
provider refused: This model's maximum context length is 512 tokens. Your request used 564 tokens.
attempt 2: sending 272 estimated tokens (12 messages dropped)
reply: '(scripted reply: request accepted)'Provider A accepts the trimmed request. Provider B rejects it (564 tokens by its count against 512), and the retry with a tighter budget succeeds. Notice what the retry cost: 12 of 23 history messages dropped instead of 9. A calibrated safety margin avoids the round trip. With a real key, the same function works with supportdesk.llm.chat, as the comment at the end of the file shows. Illustrative example of a provider's overflow error body (not captured in this build; wording varies by provider): {"message": "Please reduce the length of the messages or completion.", "type": "invalid_request_error", "param": "messages", "code": "context_length_exceeded"}.
Position effects: primacy, recency, and lost in the middle
Fitting in the window does not mean the model uses every token equally. Three effects are well documented:
- Primacy: information at the start of the context is used relatively well.
- Recency: information at the end, near the question, is used best of all.
- Lost in the middle: information in the middle of a long context is used worst.
The standard reference is Liu et al., "Lost in the Middle: How Language Models Use Long Contexts" (arXiv 2307.03172, 2023; published in TACL in 2024). In multi-document question answering, they moved the one document containing the answer among up to 30 distractors. Accuracy traced a U shape over position: highest when the answer document was first or last, lowest in the middle. For GPT-3.5-Turbo the drop was more than 20 points, and in the worst case, performance with 20 or 30 documents fell below its closed-book accuracy of 56.1%, meaning the model did better with no documents at all than with the answer buried in the middle. The effect also appeared in models built for long contexts. Newer models have reduced it, but no one guarantees it is gone for your model and your data, so measure.
Illustration: A line chart with the position of the relevant document (first to last) on the x-axis and answer accuracy on the y-axis, showing a U-shaped curve that is high at both ends and dips in the middle, with a horizontal dashed line for closed-book accuracy crossing above the dip. | Alt text: U-shaped accuracy curve over document position | File: m02-lost-in-the-middle.png
For Brightlane, the practical rules are: put the instructions that matter most at the start (system prompt) and restate the specific task right before the output (the last user message); put the retrieved article most likely to answer the question closest to the question; and do not assume a fact in the middle of a 40-message history will be used just because it is present.
A needle-in-a-haystack harness
A needle-in-a-haystack test hides one fact (the needle) at a controlled depth inside filler text (the haystack) and asks for it. It is a narrow test, but it is cheap and it catches gross position failures in your own stack. This harness builds haystacks from Brightlane help-center paragraphs, hides a random approval code at five depths and three lengths, and scores exact matches.
# examples/m02_needle.py
"""Module 2: a needle-in-a-haystack harness for position effects.
Run with a stand-in (default) to test the plumbing, or with a real model:
PYTHONPATH=. python examples/m02_needle.py --real
"""
import argparse
import random
import re
from supportdesk.data import load_articles
from supportdesk.llm import Usage, chat, resolve
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_tokens
DEPTHS = (0.0, 0.25, 0.5, 0.75, 1.0)
QUESTION = "What is the refund approval code for workspace {ws}? Reply with the code only."
NEEDLE = "Internal billing note: the refund approval code for workspace {ws} is {code}."
def build_haystack(target_tokens: int, rng: random.Random) -> list[str]:
"""Help-center paragraphs, shuffled and repeated until the text reaches about target_tokens."""
paragraphs = [p for a in load_articles() for p in a.body.split("\n") if p.strip()]
out, total = [], 0
while total < target_tokens:
p = rng.choice(paragraphs)
out.append(p)
total += count_tokens(p) + 1
return out
def make_case(target_tokens: int, depth: float, rng: random.Random) -> tuple[list[dict], str]:
ws, code = f"ACME-{rng.randint(10, 99)}", f"PLUM-{rng.randint(1000, 9999)}"
paragraphs = build_haystack(target_tokens, rng)
paragraphs.insert(round(depth * len(paragraphs)), NEEDLE.format(ws=ws, code=code))
context = "\n".join(paragraphs)
messages = [
{"role": "system", "content": "Answer using only the reference text."},
{"role": "user", "content": f"Reference text:\n{context}\n\n{QUESTION.format(ws=ws)}"},
]
return messages, code
def run(llm, lengths, trials: int, seed: int = 0) -> dict[tuple[int, float], float]:
rng = random.Random(seed)
scores = {}
for length in lengths:
for depth in DEPTHS:
hits = 0
for _ in range(trials):
messages, code = make_case(length, depth, rng)
# Reasoning models spend max_tokens on thinking first, so leave room beyond the short answer.
hits += code in llm(messages, max_tokens=512, temperature=0).text
scores[(length, depth)] = hits / trials
return scores
def perfect_reader(messages, kwargs):
"""Stand-in that finds the code with a regex: it can only prove the plumbing works."""
found = re.search(r"code for workspace \S+ is (PLUM-\d+)", messages[-1]["content"])
return found.group(1) if found else "not found"
def head_and_tail_reader(messages, kwargs, keep: float = 0.3):
"""Stand-in with a planted blind spot: it only reads the first and last 30% of the text."""
text = messages[-1]["content"]
cut = int(len(text) * keep)
return perfect_reader([{"role": "user", "content": text[:cut] + text[-cut:]}], kwargs)
def show(title: str, scores: dict, trials: int) -> None:
lengths = sorted({k[0] for k in scores})
print(f"{title}\n share answered correctly, n = {trials} per cell")
print(f" {'tokens':>7}" + "".join(f"{f'depth {d:.2f}':>12}" for d in DEPTHS))
for length in lengths:
print(f" {length:>7}" + "".join(f"{scores[(length, d)]:>12.2f}" for d in DEPTHS))
if __name__ == "__main__":
parser = argparse.ArgumentParser()
parser.add_argument("--real", action="store_true", help="call supportdesk.llm.chat (needs a provider key)")
parser.add_argument("--trials", type=int, default=5)
args = parser.parse_args()
lengths = (2_000, 8_000, 32_000)
if args.real:
calls = len(lengths) * len(DEPTHS) * args.trials
input_tokens = sum(lengths) * len(DEPTHS) * args.trials
_, model = resolve()
estimate = cost_usd(Usage(input_tokens=input_tokens, output_tokens=calls * 512), model) if model in PRICES else float("nan")
print(f"{calls} calls to {model}, about {input_tokens:,} input tokens, at most {estimate:.2f} USD")
show("Real model (supportdesk.llm.chat):", run(chat, lengths, args.trials), args.trials)
else:
show("Stand-in: perfect regex reader (plumbing check, not a model)",
run(ScriptedLLM(responder=perfect_reader), lengths, args.trials), args.trials)
show("Stand-in: planted head-and-tail blind spot (does the harness detect it?)",
run(ScriptedLLM(responder=head_and_tail_reader), lengths, args.trials), args.trials)Code explained
- In simple words: hide a code at the start, quarter, middle, three quarters, and end of a long reference text, ask for it, and tabulate how often the answer is right.
- What happens:
build_haystack()draws shuffled help-center paragraphs until the text reaches the target token count.make_case()generates a random workspace and code (so no answer can be memorized), inserts the needle at the chosen depth, and builds the messages.run()loops over lengths, depths, and trials and records the share of correct answers.max_tokens=512leaves room for a reasoning model's hidden thinking; with a tinymax_tokens, a reasoning model can spend the whole budget thinking and return an empty answer. Two stand-ins validate the harness:perfect_readerfinds the code with a regex, andhead_and_tail_readerhas a planted blind spot (it reads only the first and last 30% of the text). With--real, the same harness callsllm.chatand first prints the number of calls, input tokens, and a cost ceiling. - Comes out (stand-ins, captured):
Stand-in: perfect regex reader (plumbing check, not a model)
share answered correctly, n = 5 per cell
tokens depth 0.00 depth 0.25 depth 0.50 depth 0.75 depth 1.00
2000 1.00 1.00 1.00 1.00 1.00
8000 1.00 1.00 1.00 1.00 1.00
32000 1.00 1.00 1.00 1.00 1.00
Stand-in: planted head-and-tail blind spot (does the harness detect it?)
share answered correctly, n = 5 per cell
tokens depth 0.00 depth 0.25 depth 0.50 depth 0.75 depth 1.00
2000 1.00 1.00 0.00 1.00 1.00
8000 1.00 1.00 0.00 1.00 1.00
32000 1.00 1.00 0.00 1.00 1.00The perfect reader scores 1.00 everywhere, so the plumbing works (needles are inserted, codes are checked). The planted blind spot shows up exactly where it should, at depth 0.50, at all three lengths. That second check matters: a harness you have never seen fail cannot be trusted to detect failure. Both tables are stand-in output, not model behavior.
With --real and the default Groq model, the preamble reads 75 calls to openai/gpt-oss-120b, about 1,050,000 input tokens, at most 0.18 USD (computed, not sent). Illustrative sample run (not captured in this build; produced for teaching). Your output will differ:
tokens depth 0.00 depth 0.25 depth 0.50 depth 0.75 depth 1.00
2000 1.00 1.00 1.00 1.00 1.00
8000 1.00 1.00 1.00 1.00 1.00
32000 1.00 1.00 0.80 1.00 1.00How to read a real result: with 5 trials per cell, one miss moves a cell by 0.20, and a 95% interval around 4 of 5 runs from roughly 0.38 to 0.96, so a single 0.80 is within noise. Raise --trials to 20 or more before you conclude anything, and remember that current frontier models often score near 1.00 on this literal-match test even when they struggle on harder long-context tasks (next section).
Effective context versus advertised context
The advertised context is the largest input the API accepts. The effective context is the length up to which the model still performs your task well. They differ, often a lot.
- RULER (Hsieh et al., NVIDIA, arXiv 2404.06654, 2024) tested 17 long-context models on 13 tasks that go beyond single-needle retrieval (multiple needles, tracing variables through the text, aggregating counts, question answering). Models scored nearly perfectly on the vanilla needle test, yet performance dropped sharply as length grew, and only about half of the models that claimed 32K tokens or more kept satisfactory performance at 32K.
- NoLiMa (Modarressi et al., arXiv 2502.05167, 2025) removed the literal word overlap between question and needle, so the model has to make a one-step association instead of matching words. Of 13 models claiming at least 128K tokens, 11 fell below half of their short-context score at 32K; GPT-4o dropped from 99.3% to 69.7%.
- Context Rot (Hong, Troynikov, and Huber, Chroma technical report, July 2025) tested 18 models including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3, and found performance became increasingly unreliable as input grew, even on simple tasks, with distractors and haystack structure changing the results.
These results are for the models they tested at the time; newer models improve. The durable lesson is methodological: advertised length says what fits, not what works. Measure effective context for your task at your lengths, and treat the smallest context that contains what the model needs as the default.
| Situation | Use this | Why |
|---|---|---|
| Deciding where to put a critical instruction | Start of the system prompt, restated near the end | Primacy and recency; the middle is weakest |
| Several retrieved articles | Most relevant last, next to the question | Recency helps the article that matters most |
| Choosing a model for long inputs | Your own harness at your lengths, 20+ trials per cell | Advertised length and needle scores overstate usable length |
| Long input that is mostly irrelevant | Retrieve or compact first (Part C) | Distractors hurt even when everything fits |