CourseLarge Language Models · Module 8: Agents · part 37 of 80
Part 37 · Module 8: Agents

Part 2: The core loop

24 min read·22 Sept 2026

Perceive, decide, act, observe

Every agent, however it is marketed, is this loop:

  1. Perceive: build the context: system prompt, the ticket, every earlier step, tool definitions.
  2. Decide: call the model. It returns either tool calls or a final answer.
  3. Act: run the tool calls, behind validation, approval gates, and retries.
  4. Observe: append each result to the conversation as a tool message.
  5. Repeat, until the model answers without a tool call or a guard stops the run.

You built one turn of this in Module 6 (call, execute, return result, continue). An agent is that turn in a while loop, and nearly all the engineering is in what surrounds the loop: when to stop, what to spend, what to record, what to do when a tool fails, and which actions need a person.

.

The agent file

Here is the whole agent. It is long because it contains the fake back office (a CRM and a billing API with a JSON file standing in for the payment system), the seven tools, the approval gate, the tracer, the budget math, the loop, and the scripted policy that plays the model's part. Read the loop (run_agent) first; everything else supports it.

examples/m08_agent.py

python
"""Module 8: a support agent loop for Brightlane, with guards, budgets, tracing, and checkpoints.

The loop talks to any function with the same call signature as `supportdesk.llm.chat`.
Pass the real `chat` when you have a key, or a `ScriptedLLM` to drive the plumbing
without one. A ScriptedLLM is NOT a model: runs driven by it test the loop, the
guards, and the bookkeeping, never the quality of a model's decisions.
"""
from __future__ import annotations

import hashlib
import json
import random
import re
import time
from collections.abc import Callable
from dataclasses import asdict, dataclass, field
from pathlib import Path
from typing import Any

from supportdesk.data import get_article
from supportdesk.kb_search import KBSearch
from supportdesk.llm import ChatResult, ToolCall, Usage
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.tokens import count_messages, count_tokens

ChatFn = Callable[..., ChatResult]

# ---------------------------------------------------------------------------
# Fake back-office systems (a CRM and a billing API). Billing state lives in a
# JSON file so that it survives a simulated crash, like a real external system.
# ---------------------------------------------------------------------------

ACCOUNTS = {
    "dana.k@northwind.example": {"account_id": "ACC-311", "name": "Northwind Studio", "plan": "team",
                                 "seats": 24, "billing": "monthly", "role": "owner"},
    "li.wei@kestrel.example": {"account_id": "ACC-502", "name": "Kestrel Labs", "plan": "business",
                               "seats": 40, "billing": "annual", "role": "member"},
}

INVOICES = {
    "INV-2026-004512": {"invoice_id": "INV-2026-004512", "account_id": "ACC-311", "amount_due": 288.0,
                        "currency": "USD", "period": "2026-09",
                        "charges": [{"charge_id": "ch_7Hq1", "amount": 288.0, "date": "2026-09-03"},
                                    {"charge_id": "ch_7Hq2", "amount": 288.0, "date": "2026-09-03"}]},
    "INV-2026-004601": {"invoice_id": "INV-2026-004601", "account_id": "ACC-502", "amount_due": 11520.0,
                        "currency": "USD", "period": "2026-2027",
                        "charges": [{"charge_id": "ch_8Kc4", "amount": 11520.0, "date": "2026-09-01"}]},
    "INV-2026-004377": {"invoice_id": "INV-2026-004377", "account_id": "ACC-311", "amount_due": 288.0,
                        "currency": "USD", "period": "2026-08",
                        "charges": [{"charge_id": "ch_6Fa9", "amount": 288.0, "date": "2026-08-03"}]},
}


class TransientToolError(Exception):
    """A failure worth retrying: timeouts, 503s, rate limits."""


class BillingStore:
    """Refunds issued so far, persisted to disk. Refunds are keyed by charge id (idempotent)."""

    def __init__(self, path: str | Path) -> None:
        self.path = Path(path)
        self.refunds: dict[str, dict] = json.loads(self.path.read_text()) if self.path.exists() else {}

    def refund(self, invoice_id: str, charge_id: str, amount: float, reason: str) -> dict:
        if charge_id in self.refunds:  # idempotency: a repeated request returns the first result
            return {"already_refunded": True, **self.refunds[charge_id]}
        record = {"refund_id": f"RF-{len(self.refunds) + 1:05d}", "invoice_id": invoice_id,
                  "charge_id": charge_id, "amount": amount, "reason": reason}
        self.refunds[charge_id] = record
        self.path.write_text(json.dumps(self.refunds, indent=2))
        return record


class Flaky:
    """Wrap a tool so its first `failures` calls raise TransientToolError (a simulated outage)."""

    def __init__(self, fn: Callable[..., Any], failures: int = 2, message: str = "503 billing API timeout") -> None:
        self.fn, self.failures, self.message, self.calls = fn, failures, message, 0

    def __call__(self, **kwargs: Any) -> Any:
        self.calls += 1
        if self.calls <= self.failures:
            raise TransientToolError(f"{self.message} (call {self.calls})")
        return self.fn(**kwargs)


# ---------------------------------------------------------------------------
# Tools
# ---------------------------------------------------------------------------

@dataclass
class Tool:
    name: str
    description: str
    parameters: dict
    fn: Callable[..., Any]
    irreversible: bool = False

    def spec(self) -> dict:
        """The tool definition in Chat Completions format (Module 6)."""
        return {"type": "function", "function": {"name": self.name, "description": self.description,
                                                 "parameters": self.parameters}}


def _obj(props: dict, required: list[str]) -> dict:
    return {"type": "object", "properties": props, "required": required, "additionalProperties": False}


def make_tools(store: BillingStore, flaky_failures: int = 2, notes: list | None = None,
               escalations: list | None = None) -> dict[str, Tool]:
    """The seven support tools. `get_invoice` fails `flaky_failures` times before it works."""
    kb = KBSearch()
    notes = notes if notes is not None else []
    escalations = escalations if escalations is not None else []

    def search_kb(query: str) -> list[dict]:
        return [{"article_id": h.article_id, "title": h.title, "score": h.score} for h in kb.search(query, k=3)]

    def read_article(article_id: str) -> dict:
        a = get_article(article_id)  # raises KeyError for unknown ids, which becomes an error observation
        return {"article_id": a.id, "title": a.title, "body": a.body}

    def get_account(email: str) -> dict:
        if email not in ACCOUNTS:
            raise ValueError(f"No account for {email!r}")
        return ACCOUNTS[email]

    def get_invoice(invoice_id: str) -> dict:
        if invoice_id not in INVOICES:
            raise ValueError(f"No invoice {invoice_id!r}")
        inv = INVOICES[invoice_id]
        paid = sum(c["amount"] for c in inv["charges"])
        return {**inv, "amount_paid": paid, "overpaid": round(paid - inv["amount_due"], 2)}

    def issue_refund(invoice_id: str, charge_id: str, amount_usd: float, reason: str) -> dict:
        inv = INVOICES.get(invoice_id)
        if inv is None:
            raise ValueError(f"No invoice {invoice_id!r}")
        charge = next((c for c in inv["charges"] if c["charge_id"] == charge_id), None)
        if charge is None:
            raise ValueError(f"Charge {charge_id!r} is not on {invoice_id}")
        if not 0 < amount_usd <= charge["amount"]:
            raise ValueError(f"Refund must be between 0 and {charge['amount']} USD")
        return store.refund(invoice_id, charge_id, amount_usd, reason)

    def add_internal_note(ticket_id: str, note: str) -> dict:
        notes.append({"ticket_id": ticket_id, "note": note})
        return {"saved": True, "note_count": len(notes)}

    def escalate_to_human(ticket_id: str, reason: str) -> dict:
        escalations.append({"ticket_id": ticket_id, "reason": reason})
        return {"queued_for": "billing-team", "position": len(escalations)}

    s = {"type": "string"}
    tools = [
        Tool("search_kb", "Search Brightlane help-center articles. Returns up to 3 {article_id, title, score}; empty list if nothing matches, so rephrase with policy words.",
             _obj({"query": {**s, "description": "Keywords, in English."}}, ["query"]), search_kb),
        Tool("read_article", "Read one help-center article by id (from search_kb).",
             _obj({"article_id": s}, ["article_id"]), read_article),
        Tool("get_account", "Look up the customer's account by the requester email on the ticket.",
             _obj({"email": s}, ["email"]), get_account),
        Tool("get_invoice", "Fetch an invoice with all card charges, amount_paid and overpaid.",
             _obj({"invoice_id": {**s, "pattern": r"^INV-\d{4}-\d{6}$"}}, ["invoice_id"]),
             Flaky(get_invoice, failures=flaky_failures) if flaky_failures else get_invoice),
        Tool("issue_refund", "Refund one card charge. IRREVERSIBLE: needs human approval. Only refund charges you saw in get_invoice.",
             _obj({"invoice_id": s, "charge_id": s, "amount_usd": {"type": "number"}, "reason": s},
                  ["invoice_id", "charge_id", "amount_usd", "reason"]), issue_refund, irreversible=True),
        Tool("add_internal_note", "Add a note to the ticket that only staff can see.",
             _obj({"ticket_id": s, "note": s}, ["ticket_id", "note"]), add_internal_note),
        Tool("escalate_to_human", "Hand the ticket to the billing team with a reason.",
             _obj({"ticket_id": s, "reason": s}, ["ticket_id", "reason"]), escalate_to_human),
    ]
    return {t.name: t for t in tools}


# ---------------------------------------------------------------------------
# Approval gate, tracing, budgets
# ---------------------------------------------------------------------------

@dataclass
class Approval:
    approved: bool
    approver: str
    note: str = ""


Approver = Callable[[str, dict], Approval]


def deny_all(tool: str, args: dict) -> Approval:
    """The safe default: nothing irreversible happens without a person."""
    return Approval(False, "default-deny", "no approver configured")


class ScriptedApprover:
    """Stands in for a human clicking Approve or Reject in a review queue."""

    def __init__(self, decisions: list[bool], name: str = "maya") -> None:
        self.decisions, self.name, self.asked = list(decisions), name, []

    def __call__(self, tool: str, args: dict) -> Approval:
        self.asked.append((tool, args))
        ok = self.decisions.pop(0) if self.decisions else False
        return Approval(ok, f"human:{self.name}", "approved in review queue" if ok else "rejected")


def console_approver(tool: str, args: dict) -> Approval:
    """Ask on the terminal. Use this when you run the agent by hand with a real model."""
    answer = input(f"Approve {tool} {json.dumps(args)}? [y/N] ").strip().lower()
    return Approval(answer == "y", "human:console")


class Tracer:
    """Append one JSON object per event to a JSONL file (and keep them in memory)."""

    def __init__(self, path: str | Path | None, run_id: str) -> None:
        self.path = Path(path) if path else None
        self.run_id = run_id
        self.events: list[dict] = []

    def event(self, kind: str, step: int, **fields: Any) -> None:
        record = {"ts": round(time.time(), 3), "run_id": self.run_id, "step": step, "type": kind, **fields}
        self.events.append(record)
        if self.path:
            with self.path.open("a", encoding="utf-8") as f:
                f.write(json.dumps(record, default=str) + "\n")


def prompt_tokens(messages: list[dict], tools: list[dict] | None = None) -> int:
    """Estimate prompt tokens: message text, tool-call arguments, and tool definitions.

    count_messages (Module 2) counts message content only. Providers also bill the
    tool definitions and the arguments of earlier tool calls, so we add them.
    """
    total = count_messages(messages)
    for m in messages:
        for call in m.get("tool_calls") or []:
            total += count_tokens(call["function"]["name"] + call["function"]["arguments"])
    if tools:
        total += count_tokens(json.dumps(tools))
    return total


@dataclass
class AgentConfig:
    max_steps: int = 12
    max_tokens: int = 60_000        # input plus output, summed over the whole task
    max_usd: float = 0.05
    price_model: str = "openai/gpt-oss-120b"   # used for cost when the reply's model has no price (stand-ins)
    output_reserve: int = 400       # output tokens assumed when checking the budget before a call
    no_progress_limit: int = 3      # consecutive steps with no new observation
    repeat_limit: int = 2           # the same tool call with the same arguments, at most this many times
    max_attempts: int = 3           # tool retries on TransientToolError
    backoff_base_s: float = 0.05    # 0.05, 0.1, 0.2 ... plus jitter
    model: str | None = None        # passed to chat_fn; None lets llm.resolve pick


@dataclass
class Plan:
    steps: list[str] = field(default_factory=list)
    version: int = 0
    needs_replan: bool = False


@dataclass
class AgentState:
    task_id: str
    messages: list[dict]
    step: int = 0
    input_tokens: int = 0
    output_tokens: int = 0
    cost_usd: float = 0.0
    llm_calls: int = 0
    plan: Plan = field(default_factory=Plan)
    seen_calls: dict[str, int] = field(default_factory=dict)
    fingerprints: list[str] = field(default_factory=list)
    stale_steps: int = 0
    status: str = "running"
    final: str = ""

    def to_json(self) -> str:
        return json.dumps(asdict(self), indent=1)

    @classmethod
    def from_json(cls, text: str) -> "AgentState":
        raw = json.loads(text)
        raw["plan"] = Plan(**raw["plan"])
        return cls(**raw)


class SimulatedCrash(RuntimeError):
    """Raised on purpose to show that a checkpointed run resumes where it stopped."""


SYSTEM_PROMPT = """You are Brightlane's support agent. Work the ticket with the tools.
Before any refund: find the policy with search_kb, look up the account and the invoice, and refund
only a charge that get_invoice shows as a duplicate. Refunds need human approval; if approval is
refused or a system is down, escalate_to_human. Write 'Thought:' before each action and keep a
'PLAN:' line listing the remaining steps. When done, reply to the customer in plain text with no tool call."""


def new_state(task_id: str, ticket_text: str, system_prompt: str = SYSTEM_PROMPT) -> AgentState:
    return AgentState(task_id, [{"role": "system", "content": system_prompt},
                                {"role": "user", "content": ticket_text}])


PLAN_LINE = re.compile(r"^PLAN:\s*(.+)$", re.MULTILINE)


def _update_plan(state: AgentState, text: str) -> bool:
    match = PLAN_LINE.search(text or "")
    if not match:
        return False
    state.plan = Plan([s.strip() for s in re.split(r"\s*;\s*", match.group(1)) if s.strip()],
                      state.plan.version + 1, False)
    return True


# ---------------------------------------------------------------------------
# Tool execution: approval gate, retries with backoff, error observations
# ---------------------------------------------------------------------------

def execute_tool(tool: Tool, args: dict, approver: Approver, tracer: Tracer, step: int,
                 config: AgentConfig, rng: random.Random, sleep: Callable[[float], None] = time.sleep) -> dict:
    """Run one tool call and return an observation dict. Never raises for tool failures."""
    if "_unparseable" in args:
        return {"ok": False, "error": "Arguments were not valid JSON. Send a JSON object."}
    if tool.irreversible:
        decision = approver(tool.name, args)
        tracer.event("approval", step, tool=tool.name, args=args, approved=decision.approved,
                     approver=decision.approver, note=decision.note)
        if not decision.approved:
            return {"ok": False, "error": f"Not approved by {decision.approver}: {decision.note}", "denied": True}
    for attempt in range(1, config.max_attempts + 1):
        started = time.perf_counter()
        try:
            result = tool.fn(**args)
            tracer.event("tool_result", step, tool=tool.name, attempt=attempt, ok=True,
                         ms=round((time.perf_counter() - started) * 1000, 2),
                         preview=json.dumps(result, default=str)[:70])
            return {"ok": True, "result": result}
        except TransientToolError as exc:
            delay = config.backoff_base_s * 2 ** (attempt - 1) * (1 + rng.random() * 0.25)
            tracer.event("tool_retry", step, tool=tool.name, attempt=attempt, error=str(exc),
                         backoff_s=round(delay, 3) if attempt < config.max_attempts else None)
            if attempt < config.max_attempts:
                sleep(delay)
        except (ValueError, KeyError, TypeError) as exc:  # bad arguments: tell the model, do not retry
            tracer.event("tool_result", step, tool=tool.name, attempt=attempt, ok=False, error=str(exc))
            return {"ok": False, "error": str(exc)}
    return {"ok": False, "error": f"{tool.name} failed {config.max_attempts} times (transient)", "retryable": True}


def _fingerprint(name: str, observation: dict) -> str:
    return hashlib.sha256((name + json.dumps(observation, sort_keys=True)).encode()).hexdigest()[:12]


# ---------------------------------------------------------------------------
# The loop
# ---------------------------------------------------------------------------

def run_agent(ticket_text: str, task_id: str, tools: dict[str, Tool], chat_fn: ChatFn, *,
              config: AgentConfig | None = None, approver: Approver = deny_all,
              trace_path: str | Path | None = None, checkpoint_path: str | Path | None = None,
              crash_at_step: int | None = None, seed: int = 0, system_prompt: str = SYSTEM_PROMPT,
              sleep: Callable[[float], None] = time.sleep) -> AgentState:
    """Perceive, decide, act, observe, until the model answers or a guard stops the run."""
    config = config or AgentConfig()
    ckpt = Path(checkpoint_path) if checkpoint_path else None
    tracer = Tracer(trace_path, task_id)
    rng = random.Random(seed)
    specs = [t.spec() for t in tools.values()]
    if ckpt and ckpt.exists():
        state = AgentState.from_json(ckpt.read_text())
        tracer.event("resume", state.step, from_checkpoint=str(ckpt.name), llm_calls=state.llm_calls)
    else:
        state = new_state(task_id, ticket_text, system_prompt)
        tracer.event("start", 0, task_id=task_id)

    def stop(status: str, **why: Any) -> AgentState:
        state.status = status
        tracer.event("stop", state.step, status=status, **why)
        if ckpt:
            ckpt.write_text(state.to_json())
        return state

    while True:
        if state.step >= config.max_steps:
            return stop("max_steps", limit=config.max_steps)
        # Budget check BEFORE spending: projected tokens and dollars of the next call.
        est_in = prompt_tokens(state.messages, specs)
        projected = cost_usd(Usage(est_in, config.output_reserve), config.price_model)
        if state.input_tokens + state.output_tokens + est_in + config.output_reserve > config.max_tokens:
            return stop("budget_tokens", used=state.input_tokens + state.output_tokens, next_call=est_in)
        if state.cost_usd + projected > config.max_usd:
            return stop("budget_usd", used=round(state.cost_usd, 6), next_call=round(projected, 6))

        # Decide
        kwargs: dict[str, Any] = {"tools": specs}
        if config.model:
            kwargs["model"] = config.model
        result = chat_fn(state.messages, **kwargs)
        price_model = result.model if result.model in PRICES else config.price_model
        call_cost = cost_usd(result.usage, price_model)
        state.llm_calls += 1
        state.input_tokens += result.usage.input_tokens
        state.output_tokens += result.usage.output_tokens
        state.cost_usd += call_cost
        state.messages.append(result.as_message())
        replanned = _update_plan(state, result.text)
        tracer.event("llm_call", state.step, thought=result.text, plan=state.plan.steps if replanned else None,
                     plan_version=state.plan.version, tool_calls=[{"name": c.name, "args": c.arguments} for c in result.tool_calls],
                     input_tokens=result.usage.input_tokens, output_tokens=result.usage.output_tokens,
                     cost_usd=round(call_cost, 7), latency_ms=result.latency_ms)

        if not result.tool_calls:  # the model answered: done
            state.final = result.text
            state.step += 1
            return stop("done")

        # Act and observe
        progress = False
        for call in result.tool_calls:
            key = call.name + json.dumps(call.arguments, sort_keys=True)
            state.seen_calls[key] = state.seen_calls.get(key, 0) + 1
            if state.seen_calls[key] > config.repeat_limit:
                return stop("repeated_action", tool=call.name, args=call.arguments, times=state.seen_calls[key])
            tracer.event("tool_call", state.step, tool=call.name, args=call.arguments, call_id=call.id)
            if call.name not in tools:
                observation = {"ok": False, "error": f"Unknown tool {call.name!r}. Use one of {sorted(tools)}."}
            else:
                observation = execute_tool(tools[call.name], call.arguments, approver, tracer, state.step,
                                           config, rng, sleep)
            if observation.get("retryable"):
                state.plan.needs_replan = True
            fp = _fingerprint(call.name, observation)
            if fp not in state.fingerprints:
                state.fingerprints.append(fp)
                progress = True
            state.messages.append({"role": "tool", "tool_call_id": call.id,
                                   "content": json.dumps(observation, default=str)})
        if state.plan.needs_replan:
            state.messages.append({"role": "user", "content": "REPLAN: a tool failed after retries. "
                                   "Write a new PLAN: line that works around it."})
            tracer.event("replan_requested", state.step, old_plan=state.plan.steps)
        state.stale_steps = 0 if progress else state.stale_steps + 1
        if crash_at_step is not None and state.step == crash_at_step:
            # The worst moment: tools ran (side effects happened) but the checkpoint is not written yet.
            tracer.event("crash", state.step, note="simulated crash before checkpoint")
            raise SimulatedCrash(f"crashed during step {state.step}")
        state.step += 1
        if ckpt:
            ckpt.write_text(state.to_json())
            tracer.event("checkpoint", state.step, path=ckpt.name)
        if state.stale_steps >= config.no_progress_limit:
            return stop("no_progress", stale_steps=state.stale_steps)


# ---------------------------------------------------------------------------
# A scripted policy that plays the model's part (NOT a model)
# ---------------------------------------------------------------------------

INVOICE_RE = re.compile(r"INV-\d{4}-\d{6}")
EMAIL_RE = re.compile(r"[\w.+-]+@[\w-]+(?:\.[\w-]+)+")


def history(messages: list[dict]) -> list[tuple[str, dict, dict]]:
    """(tool name, arguments, observation) for every tool call so far, in order."""
    results = {m["tool_call_id"]: json.loads(m["content"]) for m in messages if m["role"] == "tool"}
    out = []
    for m in messages:
        for c in m.get("tool_calls") or []:
            out.append((c["function"]["name"], json.loads(c["function"]["arguments"]), results.get(c["id"], {})))
    return out


class SupportPolicy:
    """Plays a careful agent for duplicate-charge tickets, reading the conversation like a model would.

    `noise` > 0 adds seeded detours (extra searches, re-reads) to imitate run-to-run variation.
    """

    def __init__(self, noise: float = 0.0, seed: int = 0) -> None:
        self.noise, self.rng = noise, random.Random(seed)

    def reply(self, messages: list[dict], thought: str, name: str | None = None, args: dict | None = None,
              tools: list[dict] | None = None) -> ChatResult:
        calls = []
        if name:  # ids continue from the history, so a resumed run never reuses one
            n = len(history(messages)) + 1
            calls = [ToolCall(f"call_{n}", name, args or {}, json.dumps(args or {}))]
        out_tokens = count_tokens(thought) + sum(count_tokens(c.raw_arguments) + 5 for c in calls)
        return ChatResult(text=thought, tool_calls=calls, usage=Usage(prompt_tokens(messages, tools), out_tokens))

    def __call__(self, messages: list[dict], kwargs: dict) -> ChatResult:
        tools = kwargs.get("tools")
        ticket = next(m["content"] for m in messages if m["role"] == "user")
        tid = re.search(r"Ticket (T-\d+)", ticket).group(1)
        email = EMAIL_RE.search(ticket).group(0)
        done = history(messages)
        names = [d[0] for d in done]
        last_user = [m for m in messages if m["role"] == "user"][-1]["content"]
        r = lambda *a: self.reply(messages, *a, tools=tools)  # noqa: E731

        if last_user.startswith("REPLAN") and "escalate_to_human" not in names:
            return r("Thought: billing API is down after retries, so I cannot verify the charge. "
                     "Hand it to a person.\nPLAN: escalate_to_human; reply with expectations",
                     "escalate_to_human", {"ticket_id": tid, "reason": "get_invoice unavailable; duplicate charge unverified"})
        if "escalate_to_human" in names:
            return r("Thanks for flagging this. A member of our billing team is checking the duplicate "
                     "charge and will reply within one business day.")
        hits = [d for d in done if d[0] == "search_kb" and d[2].get("result")]
        if not hits:
            if "search_kb" not in names:
                return r("Thought: customer reports a double charge. Find the refund policy first.\n"
                         "PLAN: search_kb; read_article; get_account; get_invoice; issue_refund; add_internal_note; reply",
                         "search_kb", {"query": "charged twice"})
            return r("Thought: no results for the customer's words. Rephrase with policy words.",
                     "search_kb", {"query": "duplicate charge refund"})
        if self.noise and self.rng.random() < self.noise and names.count("search_kb") < 4:
            return r("Thought: check for a more specific article.", "search_kb",
                     {"query": self.rng.choice(["refund card payment", "billing charge policy", "invoice refund time"])})
        if "read_article" not in names:
            return r("Thought: read the refunds article to confirm duplicates are refundable.",
                     "read_article", {"article_id": hits[0][2]["result"][0]["article_id"]})
        if "get_account" not in names:
            return r("Thought: policy says duplicate charges are refunded in full. Confirm the requester's account.\n"
                     "PLAN: get_account; get_invoice; issue_refund; add_internal_note; reply",
                     "get_account", {"email": email})
        account = next(d[2] for d in done if d[0] == "get_account")
        if account.get("ok") and account["result"]["role"] not in ("owner", "billing_admin") \
                and "escalate_to_human" not in names:
            return r("Thought: policy says only owners and billing admins can request refunds; this requester is a "
                     f"{account['result']['role']}. Escalate.", "escalate_to_human",
                     {"ticket_id": tid, "reason": "refund requested by a non-owner"})
        if self.noise and self.rng.random() < self.noise / 2 and names.count("read_article") < 2:
            return r("Thought: re-read the policy wording on timing.", "read_article", {"article_id": "billing-refunds"})
        inv = [d for d in done if d[0] == "get_invoice" and d[2].get("ok")]
        if not inv:
            return r("Thought: fetch the invoice named in the ticket to see the charges.",
                     "get_invoice", {"invoice_id": INVOICE_RE.search(ticket).group(0)})
        invoice = inv[-1][2]["result"]
        refunds = [d for d in done if d[0] == "issue_refund"]
        if invoice["overpaid"] <= 0 and not refunds:
            return r("Thought: no duplicate on this invoice. A person should look.",
                     "escalate_to_human", {"ticket_id": tid, "reason": "customer reports duplicate, invoice shows none"})
        if not refunds:
            dup = invoice["charges"][-1]
            return r(f"Thought: two charges of {dup['amount']} on {dup['date']}; overpaid {invoice['overpaid']}. "
                     f"Refund the second one.\nPLAN: issue_refund; add_internal_note; reply",
                     "issue_refund", {"invoice_id": invoice["invoice_id"], "charge_id": dup["charge_id"],
                                      "amount_usd": invoice["overpaid"], "reason": "duplicate charge"})
        refund = refunds[-1][2]
        if refund.get("denied"):
            return r("Thought: refund was not approved. Escalate instead of retrying.",
                     "escalate_to_human", {"ticket_id": tid, "reason": "refund rejected by approver"})
        if "add_internal_note" not in names:
            rf = refund["result"]
            return r("Thought: record what I did for the team.", "add_internal_note",
                     {"ticket_id": tid, "note": f"Refunded duplicate {rf['charge_id']} ({rf['amount']} USD) as {rf['refund_id']}."})
        rf = refund["result"]
        return r(f"Hi, sorry about the double charge. We refunded the duplicate payment of {rf['amount']:.2f} USD "
                 f"on invoice {rf['invoice_id']} (refund {rf['refund_id']}). It goes back to your original card "
                 f"within 5 to 10 business days.")


T1001 = ("Ticket T-1001 from dana.k@northwind.example\n"
         "Subject: Charged twice this month\n\n"
         "Hi, my card was charged 288 USD twice on 3 September for the Team plan "
         "(invoice INV-2026-004512). Please refund the duplicate.")


def print_trace(path: str | Path) -> None:
    """A ReAct-style view of a JSONL trace: Thought, Action, Observation."""
    for line in Path(path).read_text().splitlines():
        e = json.loads(line)
        if e["type"] == "llm_call":
            first = (e["thought"] or "").splitlines()[0] if e["thought"] else ""
            print(f"[{e['step']}] THINK  {first[:96]}")
            for c in e["tool_calls"]:
                print(f"[{e['step']}] ACT    {c['name']}({json.dumps(c['args'])[:80]})")
        elif e["type"] == "tool_result":
            status = e["preview"] if e["ok"] else f"error: {e['error'][:60]}"
            print(f"[{e['step']}] OBS    {status}")
        elif e["type"] == "tool_retry":
            wait = f"backoff {e['backoff_s']} s" if e["backoff_s"] is not None else "giving up"
            print(f"[{e['step']}] RETRY  {e['tool']} attempt {e['attempt']}: {e['error']}; {wait}")
        elif e["type"] == "approval":
            print(f"[{e['step']}] GATE   {e['tool']} approved={e['approved']} by {e['approver']}")
        elif e["type"] in ("stop", "resume", "crash", "replan_requested"):
            extra = {k: v for k, v in e.items() if k not in ("ts", "run_id", "step", "type")}
            print(f"[{e['step']}] {e['type'].upper():6} {extra}")


if __name__ == "__main__":
    from supportdesk.stand_in import ScriptedLLM

    work = Path("runs/m08")
    work.mkdir(parents=True, exist_ok=True)
    for p in work.glob("agent_*"):
        p.unlink()
    store = BillingStore(work / "agent_billing.json")
    tools = make_tools(store, flaky_failures=2)
    llm = ScriptedLLM(responder=SupportPolicy())
    state = run_agent(T1001, "T-1001", tools, llm, approver=ScriptedApprover([True]),
                      trace_path=work / "agent_trace.jsonl")
    print_trace(work / "agent_trace.jsonl")
    print()
    print("status:", state.status, "| steps:", state.step, "| LLM calls:", state.llm_calls)
    print(f"tokens in/out: {state.input_tokens}/{state.output_tokens} | cost at {AgentConfig().price_model}: "
          f"{state.cost_usd:.6f} USD")
    print("final:", state.final)

Code explained

  • In simple words: a support agent in one file: tools that touch fake billing systems, a loop that lets a model pick tools until it answers, and the seatbelts around that loop.
  • What happens:
    • ACCOUNTS, INVOICES: the fake CRM and billing data. Invoice INV-2026-004512 has two identical 288 USD charges on 2026-09-03, which is the duplicate in T-1001. Dana is an owner; Li Wei is a plain member (that matters in the lab).
    • TransientToolError: the kind of failure worth retrying (timeouts, 503s). Validation errors are different: retrying get_invoice("INV-0") will never work, so those go straight back to the model.
    • BillingStore: refunds persisted to disk and keyed by charge id, so asking to refund the same charge twice returns the first refund with already_refunded: true. This is idempotency: repeating an action has the same effect as doing it once. It is what makes crash recovery safe (Part 5).
    • Flaky: wraps a tool so its first N calls raise TransientToolError. make_tools wraps get_invoice in it, so the billing API fails twice before it works.
    • Tool and make_tools: each tool has a name, a description written for the model, a JSON Schema for its arguments (Module 6), the Python function, and an irreversible flag. search_kb uses KBSearch from Module 7 and read_article uses get_article from Module 1. issue_refund validates that the charge is on the invoice and the amount is not above the charge.
    • Approval, deny_all, ScriptedApprover, console_approver: the approval gate for irreversible tools. The default denies, so forgetting to configure an approver is safe. ScriptedApprover stands in for Maya clicking Approve in a review queue; console_approver asks on the terminal when you run with a real model.
    • Tracer: appends one JSON object per event (LLM call, tool call, retry, approval, checkpoint, stop) to a JSONL file.
    • prompt_tokens: Module 2's count_messages counts message text only. Real providers also bill the tool definitions and the arguments of earlier tool calls, so this adds both. The loop uses it for the budget check before each call.
    • AgentConfig: every limit in one place. price_model is the model used for dollar math when the reply comes from a stand-in that has no price.
    • Plan, AgentState: the state that is checkpointed. The plan lives in state, not only in the model's head, so code can read it, trace it, and flag it for replanning.
    • execute_tool: the "act" step. Irreversible tools go through the approver first. Transient errors retry with exponential backoff plus jitter (0.05 s, then 0.1 s, each stretched by up to 25 percent at random so many agents do not retry in lockstep). Bad arguments become an error observation. It never raises, because a crashed loop teaches the model nothing (Module 6).
    • run_agent: the loop. Before each model call it checks the step limit and projects the next call's tokens and dollars against the budget, including output_reserve tokens for the reply. After the call it records usage and cost, updates the plan from any PLAN: line, and stops if there are no tool calls. For each tool call it checks the repeat counter, runs the tool, fingerprints the observation (a hash of tool name plus result) to detect progress, and appends the result. If a tool failed after all retries, it asks the model to replan. It writes a checkpoint after every step, and crash_at_step raises at the worst possible moment: after the tools ran but before the checkpoint.
    • SupportPolicy: plays a careful agent for duplicate-charge tickets by reading the conversation and choosing the next call. It writes Thought: lines and PLAN: lines the way we ask a real model to. With noise above zero it takes seeded detours (extra searches, re-reading) so Part 4 can show what run-to-run variation does to cost. It is not a model.
    • print_trace: renders a JSONL trace as THINK, ACT, OBS lines.
    • The __main__ block runs T-1001 with the flaky tool and an approver who says yes, then prints the trace and totals.
  • Comes out: see the next section, where we run it.

Running it: a ReAct-style trace

ReAct (Yao et al., "ReAct: Synergizing Reasoning and Acting in Language Models", 2022) is the pattern of interleaving a short reasoning step with each action: think, act, observe, think again. The reasoning is visible, which helps debugging, and each thought is grounded in the observation just before it. Our system prompt asks for a Thought: before each action, and the trace records it.

bash
python examples/m08_agent.py

Code explained

  • In simple words: run the agent on T-1001 and print what it thought, did, and saw at each step.
  • What happens: the loop runs with ScriptedLLM(responder=SupportPolicy()), the flaky get_invoice, and a ScriptedApprover([True]). Each step appends events to runs/m08/agent_trace.jsonl, and print_trace renders them. Tokens are counted with o200k_base on the exact prompts the loop sent; cost uses the openai/gpt-oss-120b price in pricing.py.
  • Comes out: (ScriptedLLM-driven: the path is the policy's, the counts are real)
text
  [0] THINK  Thought: customer reports a double charge. Find the refund policy first.
  [0] ACT    search_kb({"query": "charged twice"})
  [0] OBS    []
  [1] THINK  Thought: no results for the customer's words. Rephrase with policy words.
  [1] ACT    search_kb({"query": "duplicate charge refund"})
  [1] OBS    [{"article_id": "billing-refunds", "title": "Refunds and cancellations
  [2] THINK  Thought: read the refunds article to confirm duplicates are refundable.
  [2] ACT    read_article({"article_id": "billing-refunds"})
  [2] OBS    {"article_id": "billing-refunds", "title": "Refunds and cancellations"
  [3] THINK  Thought: policy says duplicate charges are refunded in full. Confirm the requester's account.
  [3] ACT    get_account({"email": "dana.k@northwind.example"})
  [3] OBS    {"account_id": "ACC-311", "name": "Northwind Studio", "plan": "team", 
  [4] THINK  Thought: fetch the invoice named in the ticket to see the charges.
  [4] ACT    get_invoice({"invoice_id": "INV-2026-004512"})
  [4] RETRY  get_invoice attempt 1: 503 billing API timeout (call 1); backoff 0.061 s
  [4] RETRY  get_invoice attempt 2: 503 billing API timeout (call 2); backoff 0.119 s
  [4] OBS    {"invoice_id": "INV-2026-004512", "account_id": "ACC-311", "amount_due
  [5] THINK  Thought: two charges of 288.0 on 2026-09-03; overpaid 288.0. Refund the second one.
  [5] ACT    issue_refund({"invoice_id": "INV-2026-004512", "charge_id": "ch_7Hq2", "amount_usd": 288.0, ")
  [5] GATE   issue_refund approved=True by human:maya
  [5] OBS    {"refund_id": "RF-00001", "invoice_id": "INV-2026-004512", "charge_id"
  [6] THINK  Thought: record what I did for the team.
  [6] ACT    add_internal_note({"ticket_id": "T-1001", "note": "Refunded duplicate ch_7Hq2 (288.0 USD) as RF-00)
  [6] OBS    {"saved": true, "note_count": 1}
  [7] THINK  Hi, sorry about the double charge. We refunded the duplicate payment of 288.00 USD on invoice IN
  [8] STOP   {'status': 'done'}

  status: done | steps: 8 | LLM calls: 8
  tokens in/out: 9362/377 | cost at openai/gpt-oss-120b: 0.001631 USD
  final: Hi, sorry about the double charge. We refunded the duplicate payment of 288.00 USD on invoice INV-2026-004512 (refund RF-00001). It goes back to your original card within 5 to 10 business days.

Read it top to bottom. Step 0 searches with the customer's words, "charged twice", and gets []: neither "charged" nor "twice" appears anywhere in the help center, so BM25 has nothing to match. That is a real miss from Module 7's retriever, not a staged one. Step 1 rephrases with policy words and finds billing-refunds. Step 4 shows two retries with growing backoff before the invoice arrives. Step 5 is the refund, and the GATE line shows it went through a human approval first. Eight model calls, 9,362 input tokens, 377 output tokens, about 0.0016 USD at gpt-oss-120b prices. Note how lopsided that is: almost everything is input, because each call resends the growing conversation.

.

Planning: upfront, incremental, and on failure

A plan is the agent's list of remaining steps. Three ways to handle it:

  • Upfront: write the whole plan first, then execute it. Cheap and easy to review, but brittle when step 2 reveals something step 1 did not expect.
  • Incremental: decide only the next step each time. Flexible, but the agent can wander, because nothing holds it to a goal.
  • Replanning: write a plan, execute it, and rewrite it when reality disagrees, for example when a tool fails for good.

Our loop keeps the plan in state (state.plan) and updates it whenever the model writes a PLAN: line. When a tool fails after all retries, the loop sets needs_replan and appends a REPLAN: message, so the next decision has to address the failure instead of retrying blindly. This short script reads the plan history out of two traces (run m08_guards.py from Part 3 first for the second one):

python
"""Module 8: read the plan's history out of two traces (run m08_agent.py and m08_guards.py first)."""
import json
from pathlib import Path

for trace in ["runs/m08/agent_trace.jsonl", "runs/m08/guards/tool_down_replan.jsonl"]:
    print(trace)
    for line in Path(trace).read_text().splitlines():
        e = json.loads(line)
        if e["type"] == "llm_call" and e["plan"]:
            print(f"  step {e['step']}: plan v{e['plan_version']}: {' ; '.join(e['plan'])}")
        elif e["type"] == "replan_requested":
            print(f"  step {e['step']}: REPLAN requested (tool failed after retries)")

Code explained

  • In simple words: pull every plan version and every replan request out of two JSONL traces.
  • What happens: it reads each event line, prints the plan attached to llm_call events that changed it, and marks replan_requested events. The first trace is the happy path; the second is a run where the billing API never recovers.

Comes out:

text
  runs/m08/agent_trace.jsonl
    step 0: plan v1: search_kb ; read_article ; get_account ; get_invoice ; issue_refund ; add_internal_note ; reply
    step 3: plan v2: get_account ; get_invoice ; issue_refund ; add_internal_note ; reply
    step 5: plan v3: issue_refund ; add_internal_note ; reply
  runs/m08/guards/tool_down_replan.jsonl
    step 0: plan v1: search_kb ; read_article ; get_account ; get_invoice ; issue_refund ; add_internal_note ; reply
    step 3: plan v2: get_account ; get_invoice ; issue_refund ; add_internal_note ; reply
    step 4: REPLAN requested (tool failed after retries)
    step 5: plan v3: escalate_to_human ; reply with expectations

In the happy path the plan shrinks as steps complete (versions 1 to 3). In the outage run, get_invoice fails three times at step 4, the loop requests a replan, and version 3 of the plan drops the refund entirely in favor of escalation. The plan changed because of an observation, which is the whole point of keeping it in state where code can see it.

SituationUse thisWhy
Short tasks with a known shape (most support tickets)Upfront plan, executed with checksThe plan is reviewable and the path is predictable
Exploration where each result decides the next move (debugging, research)Incremental steps with a step budgetYou cannot plan what you have not seen yet
Long tasks that touch flaky or changing systemsPlan in state plus replanning on failureA failed tool changes the plan explicitly, not by accident
Any irreversible stepA plan that names it, reviewed by a person before it runsThe approval request can show the plan, not just one call