CourseLarge Language Models · Module 7: Context Engineering · part 31 of 80
Part 31 · Module 7: Context Engineering

Part B: Retrieved knowledge

15 min read·22 Sept 2026

The single most valuable thing in a support draft's context is the help-center article that answers the ticket. Everything else is support. So before we assemble anything, we measure how often the retriever finds that article.

B.1 The retriever: supportdesk/kb_search.py

This module introduces the canonical retriever. It is a BM25 keyword search: a classic ranking formula from information retrieval that scores an article higher when it contains the query's words, especially rare words, several times, and when the article is short. It runs offline in microseconds and has no model inside, which makes its behavior easy to measure and its failures easy to explain.

supportdesk/kb_search.py

python
"""A small BM25 keyword search over the help-center articles (no external services)."""
from __future__ import annotations

import math
import re
from collections import Counter
from dataclasses import dataclass

from supportdesk.data import Article, load_articles

WORD = re.compile(r"[a-z0-9]+")
STOPWORDS = frozenset("a an and are as at be but by can do does for from has have how i if in is it its me my no not of on or our so that the their this to was we what when where which who why will with you your".split())


def tokenize(text: str) -> list[str]:
    return [w for w in WORD.findall(text.lower()) if w not in STOPWORDS]


@dataclass(frozen=True)
class Hit:
    article_id: str
    title: str
    score: float


class KBSearch:
    """BM25 ranking: rewards rare query words that appear often in a short article."""

    def __init__(self, articles: list[Article] | None = None, k1: float = 1.2, b: float = 0.75) -> None:
        self.articles = articles if articles is not None else load_articles()
        self.k1, self.b = k1, b
        self.docs = [tokenize(a.title + " " + a.title + " " + a.body) for a in self.articles]
        self.avg_len = sum(len(d) for d in self.docs) / len(self.docs)
        df = Counter(w for d in self.docs for w in set(d))
        n = len(self.docs)
        self.idf = {w: math.log(1 + (n - c + 0.5) / (c + 0.5)) for w, c in df.items()}
        self.tf = [Counter(d) for d in self.docs]

    def search(self, query: str, k: int = 3) -> list[Hit]:
        terms = tokenize(query)
        scored = []
        for article, tf, doc in zip(self.articles, self.tf, self.docs):
            score = 0.0
            for t in terms:
                if t in tf:
                    f = tf[t]
                    score += self.idf[t] * f * (self.k1 + 1) / (f + self.k1 * (1 - self.b + self.b * len(doc) / self.avg_len))
            if score > 0:
                scored.append(Hit(article.id, article.title, round(score, 3)))
        scored.sort(key=lambda h: (-h.score, h.article_id))
        return scored[:k]

Code explained

  • In simple words: a card catalog for 12 articles that ranks them by how many of your query's distinctive words they contain.
  • What happens:
    • WORD and STOPWORDS: the regular expression keeps runs of lowercase ASCII letters and digits; the stopword list removes words such as "the" and "how" that appear everywhere and say nothing about the topic.
    • tokenize(text): lowercases, extracts [a-z0-9]+ runs, drops stopwords. Note what this means for other scripts: Japanese and Hindi characters match nothing and vanish, and an accented word such as "contraseña" splits into "contrase" and "a". We will see the consequence in B.2.
    • Hit: one search result, an article id, its title, and a score rounded to 3 decimals. It is frozen (immutable), so results can be safely cached and compared.
    • KBSearch.__init__: tokenizes each article with its title included twice (titles are short and on-topic, so doubling them is a cheap boost), records the average article length, computes each word's inverse document frequency (idf, high for words found in few articles), and stores term counts per article. k1 controls how quickly repeated words stop adding score; b controls how much long articles are penalized. 1.2 and 0.75 are the textbook defaults.
    • KBSearch.search(query, k): for every article, adds up each query word's BM25 contribution (idf times a saturating function of how often the word appears, adjusted for length), keeps articles with a positive score, sorts by score and then by id so ties are deterministic, and returns the top k.
  • Comes out: a list such as [Hit(article_id='billing-refunds', title='Refunds and cancellations', score=6.21), ...]. Scores are only comparable within one query.

B.2 Measuring it: hit rates by language

Two numbers describe a retriever for our purpose. Top-1 hit rate is the share of answerable tickets whose gold article comes first. Top-3 hit rate is the share where it appears anywhere in the first three (this is also called recall@3). We measure on all 62 answerable tickets, broken down by language, with a 95 percent Wilson interval: a confidence interval for a proportion that behaves sensibly when n is small, so you can see how little two tickets tell you.

examples/m07_retrieval.py

python
"""Module 7: measure KB retrieval by language, then try one improvement.

Run from the repo root:  PYTHONPATH=. python examples/m07_retrieval.py
"""
from __future__ import annotations

import math
from collections import defaultdict

from supportdesk.data import Ticket, get_article, load_tickets
from supportdesk.kb_search import Hit, KBSearch
from supportdesk.tokens import count_tokens

# A small glossary built from the help-center vocabulary (not from the tickets):
# each English KB term with the words a Spanish, German, Japanese, or Hindi
# customer would use for it. Keys are matched as lowercase substrings, because
# the canonical tokenizer keeps only [a-z0-9] and silently drops Japanese and
# Hindi characters (and splits words at accented letters).
GLOSSARY: dict[str, tuple[str, ...]] = {
    "refund refunded": ("reembolso", "rueckerstattung", "rückerstattung", "erstatt", "返金", "रिफंड", "धनवापसी"),
    "duplicate charge charged twice": ("duplicado", "dos veces", "doppelt", "zweimal", "二重", "दो बार"),
    "cancel cancellation": ("cancelar", "kündig", "kuendig", "解約", "キャンセル", "रद्द"),
    "password reset link": ("contraseña", "contrasena", "passwort", "パスワード", "पासवर्ड"),
    "locked attempts": ("bloquead", "gesperrt", "ロック", "लॉक"),
    "reset": ("restablecer", "zuruecksetzen", "zurücksetzen", "リセット", "रीसेट"),
    "price cost plans pricing": ("precio", "cuesta", "preis", "kosten", "料金", "価格", "कीमत", "मूल्य"),
    "per user": ("por usuario", "pro benutzer", "ユーザーあたり", "प्रति उपयोगकर्ता"),
    "invoice invoices": ("factura", "rechnung", "請求書", "चालान", "इनवॉइस"),
    "annual renewal": ("anual", "jährlich", "jaehrlich", "年間", "वार्षिक"),
    "export csv": ("exportar", "exportier", "エクスポート", "निर्यात"),
    "notifications slack channel": ("notificaciones", "benachrichtigung", "通知", "सूचना"),
    "sign in login": ("iniciar sesión", "anmelden", "ログイン", "लॉगिन"),
    "subscription plan": ("suscripción", "abonnement", "サブスクリプション", "プラン", "सदस्यता", "प्लान"),
    "delete deletion data": ("eliminar", "borrar", "löschen", "loeschen", "削除", "हटा"),
}


def expand_query(ticket: Ticket, subject_weight: int = 1, glossary: bool = True) -> str:
    """Build the search query: the subject repeated, the body, and English glossary terms."""
    text = f"{ticket.subject} " * subject_weight + ticket.body
    extra = []
    if glossary:
        lowered = text.lower()
        extra = [english for english, foreign in GLOSSARY.items() if any(f in lowered for f in foreign)]
    return " ".join([text, *extra])


def answerable(split: str | None = None) -> list[Ticket]:
    return [t for t in load_tickets(split) if t.gold["answerable"]]


def evaluate(search, tickets: list[Ticket]) -> dict[str, dict[str, int]]:
    """Top-1 and top-3 hits per language. `search(ticket)` returns a ranked list of Hits."""
    table: dict[str, dict[str, int]] = defaultdict(lambda: {"n": 0, "top1": 0, "top3": 0})
    for t in tickets:
        ids = [h.article_id for h in search(t)]
        for key in (t.language, "all"):
            row = table[key]
            row["n"] += 1
            row["top1"] += ids[:1] == [t.gold["kb_article"]]
            row["top3"] += t.gold["kb_article"] in ids[:3]
    return dict(table)


def wilson(hits: int, n: int, z: float = 1.96) -> tuple[float, float]:
    """95 percent Wilson interval for a proportion: honest error bars for small n."""
    if n == 0:
        return (0.0, 0.0)
    p = hits / n
    centre = (p + z * z / (2 * n)) / (1 + z * z / n)
    half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
    return (centre - half, centre + half)


def print_table(title: str, table: dict[str, dict[str, int]]) -> None:
    print(title)
    print(f"  {'lang':5} {'n':>3} {'top1':>6} {'top3':>6}   top1 95% CI")
    for lang in ("en", "es", "de", "ja", "hi", "all"):
        if lang in table:
            r = table[lang]
            lo, hi = wilson(r["top1"], r["n"])
            print(f"  {lang:5} {r['n']:>3} {r['top1']:>3}/{r['n']:<2} {r['top3']:>3}/{r['n']:<2}   {lo:.2f} to {hi:.2f}")


def format_articles(hits: list[Hit]) -> str:
    """How retrieved articles are pasted into a prompt (same format the ContextBuilder uses)."""
    return "\n\n".join(f'<article id="{h.article_id}">\n{get_article(h.article_id).body}\n</article>' for h in hits)


if __name__ == "__main__":
    kb = KBSearch()
    tickets = answerable()
    base = evaluate(lambda t: kb.search(t.text, k=3), tickets)
    print_table("Baseline KBSearch on ticket.text (all 62 answerable tickets)", base)

    print("\nMisses at top-1 (baseline):")
    for t in tickets:
        ids = [h.article_id for h in kb.search(t.text, k=3)]
        if ids[:1] != [t.gold["kb_article"]]:
            print(f"  {t.id} {t.language} {t.split:4} gold={t.gold['kb_article']:18} got={ids}")

    print("\nWhat the tokenizer keeps from a Japanese ticket:")
    ja = next(t for t in tickets if t.language == "ja")
    from supportdesk.kb_search import tokenize
    print(f"  {ja.text!r}\n  -> {tokenize(ja.text)}")

    variants = {
        "baseline": lambda t: kb.search(t.text, k=3),
        "subject x2": lambda t: kb.search(expand_query(t, 2, glossary=False), k=3),
        "glossary": lambda t: kb.search(expand_query(t, 1, glossary=True), k=3),
        "subject x2 + glossary": lambda t: kb.search(expand_query(t, 2, glossary=True), k=3),
    }
    print("\nAblation, top-1 hits (dev n=42, test n=20, non-English n=8):")
    for name, fn in variants.items():
        dev = evaluate(fn, answerable("dev"))["all"]
        test = evaluate(fn, answerable("test"))["all"]
        non_en = evaluate(fn, [t for t in tickets if t.language != "en"])["all"]
        print(f"  {name:22} dev {dev['top1']:>2}/{dev['n']}  test {test['top1']:>2}/{test['n']}"
              f"  non-en {non_en['top1']}/{non_en['n']}  (top-3 all: "
              f"{evaluate(fn, tickets)['all']['top3']}/{len(tickets)})")

    best = variants["glossary"]
    print("\nPer-ticket flips at top-1, baseline vs glossary:")
    for t in tickets:
        before = [h.article_id for h in kb.search(t.text, k=1)] == [t.gold["kb_article"]]
        after = [h.article_id for h in best(t)][:1] == [t.gold["kb_article"]]
        if before != after:
            print(f"  {t.id} {t.language} {t.split:4} {'gained' if after else 'LOST'}")
    print_table("\nGlossary search (all 62 answerable tickets)", evaluate(best, tickets))

    print("\nRecall@k vs tokens added (glossary search, 62 tickets, o200k_base):")
    print(f"  {'k':>2} {'recall@k':>9} {'mean tokens':>12} {'extra tokens per extra hit':>27}")
    prev_hits, prev_tokens = 0, 0.0
    for k in range(1, 6):
        hits_found, tokens = 0, 0
        for t in tickets:
            hits = kb.search(expand_query(t), k=k)
            hits_found += t.gold["kb_article"] in [h.article_id for h in hits]
            tokens += count_tokens(format_articles(hits))
        mean_tokens = tokens / len(tickets)
        gained = hits_found - prev_hits
        per_hit = f"{(mean_tokens - prev_tokens) * len(tickets) / gained:.0f}" if gained and k > 1 else "-"
        print(f"  {k:>2} {hits_found:>3}/{len(tickets)} ({hits_found / len(tickets):.2f}) {mean_tokens:>9.0f}"
              f" {per_hit:>27}")
        prev_hits, prev_tokens = hits_found, mean_tokens

Code explained

  • In simple words: run every answerable ticket through the retriever, count how often the right article comes back, then try one fix and count again.
  • What happens:
    • GLOSSARY maps English help-center terms to the words Spanish, German, Japanese, and Hindi customers use for them. It is matched by substring, not by the tokenizer, because the tokenizer throws Japanese and Hindi away.
    • expand_query(ticket, subject_weight, glossary) builds the query: the subject repeated subject_weight times, the body, and the English term for every glossary entry whose foreign word appears in the ticket.
    • evaluate(search, tickets) counts top-1 and top-3 hits per language and overall. wilson computes the interval; print_table prints it.
    • format_articles renders hits exactly the way the prompt will contain them, so token counts in the recall@k table are the real cost of adding them.
    • The main block prints the baseline table, every top-1 miss, what the tokenizer keeps of a Japanese ticket, an ablation of four query variants on the dev and test splits, the per-ticket flips, the improved table, and recall@k against tokens for k = 1 to 5.
  • Comes out: (deterministic; this runs in under a second)

B.3 One improvement, measured honestly

We tried two cheap ideas, alone and together:

  • Subject weighting: repeat the ticket subject so its words count double. The subject is usually a compact statement of the topic.
  • A tiny multilingual glossary: for each key help-center term, the words a customer would write in each supported language. When one appears, append the English term to the query.

The ablation rows say subject weighting did nothing useful (dev 34 of 42 vs 35 of 42 without it, a one-ticket difference in the wrong direction), so we drop it. The glossary alone moved top-1 from 52 to 56 of 62 and top-3 from 57 to 59. The per-ticket flips show where: four gains (both Japanese tickets, the Hindi ticket, one Spanish ticket) and no losses.

Now the honest part. Is this real or noise?

  • Paired test. Both variants ran on the same 62 tickets, so the right comparison counts only tickets where they disagree: 4 gained, 0 lost. An exact McNemar test (a sign test on those disagreements) gives p = 2 x 0.5^4 = 0.125. Four flips in one direction is suggestive, not conclusive at the usual 0.05 level.
  • Held-out split. On the test split the gain is one ticket (17 to 18 of 20). The glossary was written after seeing the dev misses, so the dev gain is partly by construction; the test split is the fairer check, and it has only two non-English tickets.
  • Intervals. Overall top-1 went from 0.84 (0.73 to 0.91) to 0.90 (0.80 to 0.95). The intervals overlap heavily.
  • What did not move. The English misses are untouched. T-1001 ("Charged twice... invoice INV-2026-004512") still retrieves billing-invoices first, because the word "invoice" outweighs "charged twice". Spanish T-1031 now ranks the right article second, just behind invoices, for the same reason.

The conclusion we can defend: the glossary fixes a mechanical failure (non-Latin scripts are invisible to the tokenizer), it cost nothing in English, and the before-and-after difference is consistent but within noise on this sample. For production you would reach for a multilingual embedding model or a translation step (Module 14 covers retrieval patterns more broadly), and you would grow the glossary from real traffic rather than from our eight tickets.

SituationUse thisWhy
Keyword retriever and customers write in several scriptsA glossary or a translate-then-search step in front of BM25The tokenizer silently drops non-Latin text; no ranking tweak fixes an empty query
Retrieval quality matters across many languages at scaleA multilingual embedding model, evaluated with the same harnessKeyword maps do not scale to every phrasing; measure it before switching, not after
You changed the retriever and want to know if it helpedPaired comparison on the same tickets, count wins and losses, check a held-out splitOverall percentages hide flips, and small splits hide overfitting

B.4 How much to retrieve: recall@k against tokens

Every article you add costs tokens on every call and, per Part A, spends attention. The last table in the output answers "how many articles?" with numbers. Recall@k is the share of tickets whose gold article is among the first k; mean tokens is what those k articles add to the prompt:

krecall@k (glossary search)tokens added per tickettokens spent per extra ticket answered
156/62 (0.90)109
258/62 (0.94)213about 3,200
359/62 (0.95)308about 5,900
460/62 (0.97)391about 5,100
561/62 (0.98)463about 4,400

The second article buys two more tickets for about 100 tokens each call; after that, each additional ticket rescued costs thousands of tokens spread across all 62 calls, and brings in articles that are wrong for most tickets. Our help-center articles are short (69 to 115 tokens each), so k = 3 is still cheap here. With 2,000-token articles the same curve would say k = 1 or 2 plus a fallback. The general method: plot recall@k against tokens on your own data and stop where the curve flattens.

SituationUse thisWhy
Short articles, recall still climbing at k = 3k = 3Recall 0.95 for about 300 tokens
Long articles or a tight windowk = 1 or 2, plus a "search again" tool (Part F.2)Each extra article costs more than it rescues
The top score is far above the restSend only the top hitA confident retriever rarely needs backups; saves tokens and attention

B.5 In what order

Once you know which articles to send, their order still matters because of the position effects in Part A.3. Two reasonable layouts:

  • Rank order: best article first. Easy to read, and puts the best article near the start of the volatile section.
  • Best last: best article immediately before the ticket, where recency helps it most.

The ContextBuilder in Part C supports both. For Japanese ticket T-1062 the glossary search returns account-login first:

python
# Where does the best article go? PYTHONPATH=.:examples python examples/m07_order.py
from m07_context import ContextBuilder
from m07_retrieval import expand_query
from supportdesk.data import load_tickets
from supportdesk.kb_search import KBSearch

kb = KBSearch()
t = next(x for x in load_tickets() if x.id == "T-1062")
hits = kb.search(expand_query(t), k=3)
for best_last in (False, True):
    b = ContextBuilder().knowledge(hits, best_last=best_last)
    print(f"best_last={best_last}: order sent = {b.reports['knowledge'].note}")

Code explained

  • In simple words: show the same three articles in both orders.
  • What happens: ContextBuilder().knowledge(hits, best_last=...) fits whole articles into the knowledge budget and records the order it used in the report's note.
  • Comes out: