CourseLarge Language Models · Module 7: Context Engineering · part 33 of 80
Part 33 · Module 7: Context Engineering

Part D: Memory

15 min read·22 Sept 2026

D.1 Session memory vs long-term memory

Session memory is what the assistant knows within one ticket: the history and the state scaffold. It is born when the ticket opens and discarded when it closes. Long-term memory is what survives across tickets: "this customer is the billing admin", "prefers replies in English", "has an open refund request". It lives in a store outside the model, and a retrieval step decides which memories go into the next context.

The model itself remembers nothing between calls. "Memory" in an LLM system is always your code choosing text to put back into the window, which means it has all the problems of retrieval (relevance, size, ordering) plus new ones: facts go stale, sources disagree, and customers have a right to see and delete what you keep.

SituationUse thisWhy
Facts needed only while this ticket is open (invoice number, what was already tried)Session memory: history plus state scaffoldDiscarded at close, no privacy footprint
Durable preferences and roles (language, billing admin, plan quirks)Long-term memory store, pinned into the profile sectionSaves the customer repeating themselves on every ticket
Facts your systems of record already hold (plan, seats, invoices)Look them up with a tool or from the CRM, do not "remember" themThe CRM is authoritative and current; a memory is a copy that goes stale
Anything sensitive (card numbers, passwords, 2FA codes, health)Never store; redact before the extractor sees itMemory multiplies exposure: every later prompt could leak it

D.2 The memory store

The store below does six jobs: extract facts from a conversation with a model call, validate and redact them, store them with a timestamp and source, retrieve the relevant ones, resolve conflicts, and give the customer control (forget, opt out, export).

examples/m07_memory.py

python
"""Module 7: session memory and a long-term memory store for the support assistant.

Long-term memories are facts about a customer that outlive one ticket
("prefers English", "is the billing admin"). This file extracts them from a
conversation with a model call, validates and redacts them, stores them with a
timestamp and a source, retrieves the relevant ones for a new ticket, resolves
conflicts by source trust and recency, forgets on request, and honors opt-out.
Run:  PYTHONPATH=. python examples/m07_memory.py
"""
from __future__ import annotations

import json
import re
from collections.abc import Callable
from dataclasses import asdict, dataclass, field
from datetime import datetime, timedelta
from pathlib import Path
from typing import Any, Literal

from pydantic import BaseModel, Field, ValidationError

from supportdesk.kb_search import tokenize
from supportdesk.stand_in import ScriptedLLM

Source = Literal["crm", "customer_said", "agent_note", "model_inferred"]
TRUST: dict[str, int] = {"crm": 3, "customer_said": 2, "agent_note": 2, "model_inferred": 1}
PINNED_KEYS = {"reply_language", "role"}          # always injected when present
TTL = {"open_issue": timedelta(days=30)}          # keys that go stale; others never expire on their own

EXTRACT_PROMPT = """Extract durable facts about this customer from the conversation below.
Return a JSON list. Each item: {"key": snake_case slot, "value": short value, "text": one sentence,
"source": "customer_said" if the customer stated it, "model_inferred" if you are guessing}.
Only facts useful for future tickets: preferences, role, plan details, open issues.
Never include passwords, card numbers, 2FA codes, or health, religion, or other sensitive data.
Return [] if nothing qualifies.

Conversation:
"""

REDACT = [
    (re.compile(r"\b\d(?:[ -]?\d){12,18}\b"), "[card]"),
    (re.compile(r"(?i)\b(password|passwort|contraseña)\b\s*(is|:)?\s*\S+"), "[secret]"),
    (re.compile(r"(?<![-\w])\d{6}\b"), "[code]"),  # 2FA codes, but not the tail of INV-2026-004977
]


class ExtractedFact(BaseModel):
    """What we accept from the extractor model. Anything else is rejected."""
    key: str = Field(pattern=r"^[a-z][a-z0-9_]{1,40}$")
    value: str = Field(min_length=1, max_length=120)
    text: str = Field(min_length=3, max_length=240)
    source: Literal["customer_said", "model_inferred"]


@dataclass
class Memory:
    customer_id: str
    key: str
    value: str
    text: str
    source: str
    created_at: datetime
    ticket_id: str
    active: bool = True

    def line(self) -> str:
        return f"{self.text} ({self.source}, {self.created_at:%Y-%m-%d}, {self.ticket_id})"


@dataclass
class SessionMemory:
    """Scratch state for one ticket. Cleared when the ticket closes; never shared across customers."""
    ticket_id: str
    notes: list[str] = field(default_factory=list)
    confirmed: dict[str, str] = field(default_factory=dict)

    def as_state(self) -> dict[str, Any]:
        return {"ticket": self.ticket_id, "confirmed": self.confirmed, "notes": self.notes[-5:]}


def redact(text: str) -> tuple[str, int]:
    count = 0
    for pattern, replacement in REDACT:
        text, n = pattern.subn(replacement, text)
        count += n
    return text, count


class MemoryStore:
    def __init__(self, path: Path | None = None) -> None:
        self.path = path
        self.memories: list[Memory] = []
        self.opted_out: set[str] = set()
        self.audit: list[dict[str, str]] = []  # what happened, never the forgotten content

    # --- writing ----------------------------------------------------------------------------------
    def add(self, memory: Memory) -> Memory:
        """Store one memory, then re-resolve the active winner for its key."""
        if memory.customer_id in self.opted_out:
            self._log("skipped_opt_out", memory.customer_id, memory.key)
            return memory
        memory.text, n = redact(memory.text)
        memory.value, m = redact(memory.value)
        if n + m:
            self._log("redacted", memory.customer_id, memory.key)
        self.memories.append(memory)
        self._resolve(memory.customer_id, memory.key)
        return memory

    def extract(self, conversation: list[dict[str, str]], customer_id: str, ticket_id: str, when: datetime,
                chat: Callable[..., Any]) -> tuple[list[Memory], list[str]]:
        """Ask a model for facts, validate each one, store the valid ones. Returns (stored, rejected reasons)."""
        if customer_id in self.opted_out:
            return [], ["customer opted out of memory"]
        transcript = "\n".join(f"{m['role']}: {redact(m['content'])[0]}" for m in conversation)
        reply = chat([{"role": "user", "content": EXTRACT_PROMPT + transcript}], temperature=0.0, max_tokens=400)
        try:
            items = json.loads(reply.text)
            if not isinstance(items, list):
                raise ValueError("not a list")
        except (json.JSONDecodeError, ValueError) as err:
            return [], [f"unparseable extractor output: {err}"]
        stored, rejected = [], []
        for item in items:
            try:
                fact = ExtractedFact.model_validate(item)
            except ValidationError as err:
                rejected.append(f"{str(item)[:60]}: {err.errors()[0]['msg']}")
                continue
            if redact(fact.value)[1]:
                rejected.append(f"{fact.key}: value is sensitive, not stored")
                self._log("refused_sensitive", customer_id, fact.key)
                continue
            stored.append(self.add(Memory(customer_id, fact.key, fact.value, fact.text, fact.source, when, ticket_id)))
        return stored, rejected

    def _resolve(self, customer_id: str, key: str) -> None:
        """For one slot, the most trusted source wins; among equals, the most recent wins."""
        same = [m for m in self.memories if m.customer_id == customer_id and m.key == key]
        winner = max(same, key=lambda m: (TRUST[m.source], m.created_at))
        for m in same:
            m.active = m is winner

    # --- reading ----------------------------------------------------------------------------------
    def retrieve(self, customer_id: str, query: str, now: datetime, k: int = 3) -> list[Memory]:
        """Pinned slots always, plus the k active memories that share the most words with the query."""
        if customer_id in self.opted_out:
            return []
        live = [m for m in self.memories if m.customer_id == customer_id and m.active
                and not (m.key in TTL and now - m.created_at > TTL[m.key])]
        pinned = [m for m in live if m.key in PINNED_KEYS]
        q = set(tokenize(query))

        def score(m: Memory) -> tuple[float, datetime]:
            words = set(tokenize(m.text + " " + m.key.replace("_", " ")))
            return (len(q & words) / (len(words) ** 0.5 or 1), m.created_at)

        relevant = sorted((m for m in live if m.key not in PINNED_KEYS), key=score, reverse=True)
        relevant = [m for m in relevant if score(m)[0] > 0][:k]
        return pinned + relevant

    # --- user control -------------------------------------------------------------------------------
    def forget(self, customer_id: str, key: str | None = None) -> int:
        """Delete one slot or everything for a customer. Deleted means gone, including superseded versions."""
        before = len(self.memories)
        self.memories = [m for m in self.memories
                         if not (m.customer_id == customer_id and (key is None or m.key == key))]
        removed = before - len(self.memories)
        self._log("forgot", customer_id, key or "*", str(removed))
        return removed

    def opt_out(self, customer_id: str) -> int:
        self.opted_out.add(customer_id)
        return self.forget(customer_id)

    def export(self, customer_id: str) -> list[dict[str, Any]]:
        """Everything stored about one customer, for an access request."""
        return [{**asdict(m), "created_at": m.created_at.isoformat()} for m in self.memories if m.customer_id == customer_id]

    def _log(self, action: str, customer_id: str, key: str, detail: str = "") -> None:
        self.audit.append({"action": action, "customer": customer_id, "key": key, "detail": detail})

    def save(self) -> None:
        if self.path:
            data = {"memories": [{**asdict(m), "created_at": m.created_at.isoformat()} for m in self.memories],
                    "opted_out": sorted(self.opted_out), "audit": self.audit}
            self.path.write_text(json.dumps(data, ensure_ascii=False, indent=1))


# Scripted extractor replies. These are hand-written stand-ins for what a model might return,
# including two bad items, so the validation and redaction paths run. NOT model output.
SCRIPTED_EXTRACTIONS = [
    json.dumps([
        {"key": "role", "value": "billing admin", "text": "The customer is the workspace billing admin.", "source": "customer_said"},
        {"key": "reply_language", "value": "en", "text": "Prefers replies in English.", "source": "customer_said"},
        {"key": "open_issue", "value": "annual renewal refund INV-2026-004977", "text": "Asked for a refund of the 16 September annual renewal, invoice INV-2026-004977.", "source": "customer_said"},
        {"key": "finance_contact_language", "value": "es", "text": "Their finance lead Carla prefers Spanish.", "source": "model_inferred"},
    ]),
    json.dumps([
        {"key": "reply_language", "value": "es", "text": "Seems to prefer Spanish.", "source": "model_inferred"},
        {"key": "payment_card", "value": "4111 1111 1111 1111", "text": "Pays with card 4111 1111 1111 1111.", "source": "customer_said"},
        {"key": "Mood", "value": "annoyed", "text": "Customer is annoyed.", "source": "model_inferred"},
        {"key": "slack_channel", "value": "#finance-ops", "text": "Posts Brightlane notifications to Slack channel #finance-ops.", "source": "customer_said"},
    ]),
    json.dumps([
        {"key": "reply_language", "value": "es", "text": "Asked to switch replies to Spanish from now on.", "source": "customer_said"},
    ]),
    "Sure! Here are the facts I found: the customer likes English.",
]

CONVERSATIONS = [
    ("T-1004", datetime(2026, 9, 18, 10, 0), [
        {"role": "user", "content": "We renewed our annual Business plan by mistake, invoice INV-2026-004977, renewed 16 September. I'm the billing admin. Please write to me in English even though our finance lead Carla prefers Spanish."}]),
    ("T-2210", datetime(2026, 9, 19, 9, 30), [
        {"role": "user", "content": "Hola, the Slack notifications to #finance-ops stopped. My card 4111 1111 1111 1111 was charged too, whatever."}]),
    ("T-2231", datetime(2026, 9, 20, 16, 5), [
        {"role": "user", "content": "From now on please answer me in Spanish, my colleague will read the replies."}]),
    ("T-2240", datetime(2026, 9, 21, 8, 0), [
        {"role": "user", "content": "Another question about invoices."}]),
]


if __name__ == "__main__":
    store = MemoryStore()
    extractor = ScriptedLLM(replies=list(SCRIPTED_EXTRACTIONS))  # plumbing test, not a model
    customer = "cust-acme"
    for ticket_id, when, convo in CONVERSATIONS:
        stored, rejected = store.extract(convo, customer, ticket_id, when, extractor)
        print(f"{ticket_id} {when:%Y-%m-%d}: stored {len(stored)}, rejected {len(rejected)}")
        for m in stored:
            print(f"  + {m.key} = {m.value!r} [{m.source}]{'' if m.active else ' (inactive: outranked)'}")
        for r in rejected:
            print(f"  - rejected: {r}")

    print("\nActive vs superseded for reply_language:")
    for m in store.memories:
        if m.key == "reply_language":
            print(f"  {'ACTIVE ' if m.active else 'old    '} {m.value} {m.source:14} {m.created_at:%Y-%m-%d %H:%M} {m.ticket_id}")

    for now, query in [(datetime(2026, 9, 21, 9, 0), "Where do I download last month's invoice?"),
                       (datetime(2026, 9, 21, 9, 0), "Slack stopped posting notifications"),
                       (datetime(2026, 11, 2, 9, 0), "Did my refund go through?")]:
        print(f"\nretrieve at {now:%Y-%m-%d} for {query!r}:")
        for m in store.retrieve(customer, query, now):
            print(f"  {m.key:15} {m.line()}")

    print("\nWhat the extractor model was shown for T-2210 (redacted before the call):")
    print(" ", extractor.calls[1]["messages"][0]["content"].split("Conversation:\n")[1])

    print("\nCustomer asks: 'forget my Slack setup'")
    print("  removed:", store.forget(customer, "slack_channel"))
    print("Customer opts out of memory entirely")
    print("  removed:", store.opt_out(customer))
    stored, rejected = store.extract(CONVERSATIONS[0][2], customer, "T-2250", datetime(2026, 9, 21, 12, 0),
                                     ScriptedLLM(replies=[SCRIPTED_EXTRACTIONS[0]]))
    print("  extract after opt-out:", len(stored), rejected)
    print("  export:", store.export(customer))
    print("\nAudit log (no forgotten content is kept):")
    for entry in store.audit:
        print(" ", entry)

Code explained

  • In simple words: a notebook about each customer, where every entry says who said it and when, the most trustworthy recent entry wins, and the customer can tear out pages.
  • What happens:
    • TRUST, PINNED_KEYS, TTL: the policy. The CRM (system of record) outranks what the customer said, which outranks what the model guessed. reply_language and role are always injected. open_issue memories expire after 30 days.
    • EXTRACT_PROMPT: the instruction for the extractor model. It asks for a JSON list of {key, value, text, source} and forbids sensitive data. The prompt is a request; the code below enforces it.
    • REDACT and redact(text): regular expressions for card numbers, "password: ..." phrases, and six-digit codes, replaced with placeholders. Redaction runs on the transcript before the extractor sees it and again on everything stored.
    • ExtractedFact: a pydantic model (Module 6) that every extracted item must pass: a snake_case key, bounded lengths, and a source the model is allowed to claim. The model cannot claim crm or agent_note; only your code can.
    • Memory and Memory.line(): one stored fact with customer, key, value, text, source, timestamp, the ticket it came from, and whether it is the active winner for its key. line() renders it with provenance, so the model sees "(customer_said, 2026-09-20, T-2231)" and so can you when debugging.
    • SessionMemory: the per-ticket counterpart, shown for contrast; its as_state() feeds the state scaffold.
    • MemoryStore.add: skips opted-out customers, redacts, appends, then re-resolves the key.
    • MemoryStore.extract: redacts the transcript, calls the extractor, parses JSON, validates each item, refuses items whose value is sensitive, and stores the rest. It returns what it stored and why it rejected the rest, so failures are visible rather than silent.
    • MemoryStore._resolve: for one customer and key, the winner is the highest trust, then the most recent. Older versions stay stored but inactive, which keeps an audit trail and lets you explain why an answer used what it used.
    • MemoryStore.retrieve: live memories only (active and not expired), pinned keys always, then up to k others ranked by word overlap with the new ticket, using the same tokenize as the retriever.
    • MemoryStore.forget, opt_out, export: delete one key or everything (including inactive versions), stop storing anything for a customer, and return everything stored for an access request. _log records actions, never the forgotten content. save writes JSON if you gave it a path.
    • SCRIPTED_EXTRACTIONS and CONVERSATIONS: four hand-written extractor replies for four conversations. They are a ScriptedLLM script, not model output. They include deliberately bad items (a card number, an invalid key, a non-JSON reply) so every validation path runs.
  • Comes out: run python examples/m07_memory.py. The extractor is a ScriptedLLM replaying hand-written replies, so this output tests the store's logic, not extraction quality:

D.3 Reading the run, and a bug it caught

Read the output top to bottom:

  • T-1004 stores four facts, including the model's guess about Carla's language (marked model_inferred).
  • T-2210 stores the model's guess that the customer prefers Spanish, but it is immediately inactive: a model_inferred guess cannot outrank what the customer said on T-1004. The card number is refused as sensitive, and Mood fails the key pattern.
  • T-2231 is the customer explicitly asking for Spanish. Same trust level as the English request, more recent, so it wins. The "Active vs superseded" table shows all three versions with their sources and times.
  • T-2240: the extractor returned chatty prose instead of JSON; the store rejects it and stores nothing. A memory system that "does its best" with unparseable output is how junk gets in.
  • Retrieval always returns the pinned role and reply_language, then the relevant extra: the open refund issue for an invoice question, the Slack channel for a Slack question. On 2 November the open issue is past its 30-day TTL and no longer appears.
  • Redaction replaced the card number before the extractor ever saw the text.
  • Forget and opt-out remove 1 and then 6 memories (including superseded versions), later extraction refuses to run, and export returns an empty list. The audit log records the actions and counts, not the content.

The first version of this file had a real bug, found by reading this output. Its six-digit-code pattern was \b\d{6}\b, and the first run printed:

text
  + open_issue = 'annual renewal refund INV-2026-[code]' [customer_said]

Code explained

  • In simple words: the redactor mistook the tail of an invoice number for a 2FA code.
  • What happens: \b treats the hyphen in INV-2026-004977 as a word boundary, so 004977 matched "six digits standing alone". The fix is the lookbehind (?<![-\w]) in the current file: six digits that do not follow a hyphen or a word character. A second, subtler bug in the card pattern swallowed the space after the number ("[card]was charged"); the current pattern starts and ends on a digit.
  • Comes out: after both fixes the invoice number survives intact and the card number is still redacted. tests/test_m07_context.py::test_memory_redaction_keeps_invoice_numbers pins both behaviors so they cannot regress.

Over-redaction is not harmless: a memory that says "refund for INV-2026-[code]" is useless to the next agent, and the model may invent a number to fill the gap. Test redaction on real ticket text, in both directions.

D.4 What to remember, and how to inject it

Remember what saves the customer effort on the next ticket and cannot be looked up: preferences, roles, ongoing issues, and commitments made. Do not remember what a system of record holds (plan, seat count, invoices), and never remember secrets.

Injection strategies:

SituationUse thisWhy
A few facts that apply to every ticket (language, role)Pin them: always injectCheap, and forgetting them is visible to the customer
Many facts, only some relevantRetrieve by relevance to the new ticket, cap at kKeeps the profile section small; irrelevant memories are distractors
Facts that go stale (open issues, temporary workarounds)Attach a TTL per key and filter at retrievalStale memories become poisoning (Part E)
The model needs to know how sure a memory isInject provenance: source and dateLets the model and the human reviewer weigh "said last week" against "guessed last year"

D.5 Conflicting and stale memories

Conflicts are normal: people change their minds, and extractors guess. The store's rule, "highest trust, then most recent", encodes two decisions you should make explicitly:

  • A model's inference never overrides something a person said. Without this, one bad extraction ("seems to prefer Spanish") silently flips every future reply.
  • Among equally trusted sources, newer wins. Without timestamps you cannot apply this rule at all, which is why every memory carries one.

Stale memories are a quieter problem. The open_issue TTL removes resolved-looking issues after 30 days; a better signal, when you have it, is the ticket system's own status ("closed"), which should delete the memory directly.

D.6 Privacy and user control

Long-term memory is personal data. The store gives customers four controls, each visible in the output: forget one thing, forget everything and opt out, export what is stored (an access request), and an audit log that proves deletions happened without keeping the deleted content. Ticket T-1024 in the dataset is a GDPR Article 17 erasure request; the help center says such requests go to privacy@brightlane.example, and the memory store must be in scope when they are processed. Module 11 covers retention and logging policy; the rule here is simpler: if you cannot delete it, do not store it.