CourseLarge Language Models · Module 8: Agents · part 40 of 80
Part 40 · Module 8: Agents

Part 5: Reliability in the loop

7 min read·22 Sept 2026

Parts 2 and 3 already ran the reliability machinery. This part slows down on each piece and on the design choices behind it.

Retries with backoff, and why idempotency comes first

Tools fail. Billing APIs time out, rate limits bite, a network blips. Inside an agent there are two very different places to handle that:

  • In the loop, around the tool (execute_tool): retry transient errors a few times with exponential backoff and jitter, and only then report failure to the model. The model never sees the two 503s in step 4 of our trace; it sees one clean invoice.
  • In the model's reasoning: after retries are exhausted, the model gets {"ok": false, "retryable": true} and a REPLAN: message, and chooses another path (escalation).

Retrying in the loop is cheap (no model call); retrying through the model costs a full call each time. So retry mechanically first, and involve the model only when the failure changes the plan.

There is a trap. A timeout does not tell you whether the action happened. If issue_refund times out after the payment processor accepted it, a blind retry refunds twice. The fix is to make the action idempotent: the billing store keys refunds by charge id, so the second request returns the first refund. Real payment APIs offer the same thing as an idempotency key you send with the request. Rule of thumb: never put an automatic retry around a non-idempotent side effect.

SituationUse thisWhy
Read-only tool, transient error (timeout, 503, 429)Retry in the loop with exponential backoff and jitter, 2 to 4 attemptsCheap, and invisible to the model
Bad arguments (unknown invoice, schema violation)No retry; return the error to the modelRetrying the same input cannot succeed; the model can fix the input
Side-effecting tool (refund, email, write)Idempotency key, then retries are safeA timeout may hide a success
Retries exhaustedFlag a replan; let the model choose another path or escalateThe plan is now wrong, not just delayed

Checkpointing long-running tasks

A checkpoint is the saved state of a run, written after each step, from which a new process can continue. Ours is AgentState serialized to JSON. Here is what is in the one the crash scenario left behind:

bash
python -c "
import json
d = json.load(open('runs/m08/guards/crash_resume_ckpt.json'))
for k, v in d.items():
    print(f'{k:13}', v if not isinstance(v, (list, dict)) or k == 'plan' else f'<{len(v)} items>')
"

Code explained

  • In simple words: open the checkpoint file and list what it holds, summarizing long lists.
  • What happens: the checkpoint is plain JSON, so any tool can read it. Lists and dicts are shown as item counts except the plan, which is small enough to print.

Comes out:

text
  task_id       T-1001
  messages      <17 items>
  step          8
  input_tokens  9376
  output_tokens 377
  cost_usd      0.0016326
  llm_calls     8
  plan          {'steps': ['issue_refund', 'add_internal_note', 'reply'], 'version': 3, 'needs_replan': False}
  seen_calls    <7 items>
  fingerprints  <7 items>
  stale_steps   0
  status        done
  final         Hi, sorry about the double charge. We refunded the duplicate payment of 288.00 USD on invoice INV-2026-004512 (refund RF-00001). It goes back to your original card within 5 to 10 business days.

The whole conversation (17 messages), the counters that the guards need (seen_calls, fingerprints, stale_steps), the budget spent so far, and the plan. If any of these were missing, a resumed run would reset its guards or its budget, which is how "resumable" agents quietly overspend.

Three things make resuming safe, and our run used all three: the checkpoint holds everything the guards and budget need; external side effects are idempotent, because the crash can land between the side effect and the checkpoint; and the external world (the billing file) is re-read on resume rather than trusted from memory. Anthropic describes the same approach in its multi-agent research system post: "retry logic and regular checkpoints" so agents can "resume from where the agent was when the errors occurred" instead of restarting.

One design gap is visible in the crash trace: Maya was asked to approve the same refund twice. A production gate should record approvals durably, keyed by the exact arguments, so a replayed step finds its existing approval.

Human-in-the-loop approval gates

An approval gate is code that pauses before an irreversible action and asks a person. The important word is code: a sentence in the system prompt ("always ask before refunding") is a request to the model, while if tool.irreversible: decision = approver(...) is a guarantee. The gate in execute_tool has these properties, each for a reason:

  • Default deny. deny_all is the default approver, so a misconfigured deployment refuses rather than refunds.
  • Keyed on the tool, not the model's intent. issue_refund is flagged irreversible=True; the model cannot talk its way around the flag.
  • Shows the exact arguments. The approver sees invoice_id, charge_id, and amount_usd, and approves exactly those.
  • Records who approved. The trace stores approver: "human:maya". Part 8's evaluator uses that to fail runs approved by an automatic policy.
  • Rejection is an observation. A denial returns {"ok": false, "denied": true} to the model, which should escalate, not retry.

Keep the number of gated actions small. If every step needs approval, reviewers click Approve without reading (Module 14 calls this reviewer fatigue), and the gate becomes decoration.

Approval queue entry for a refund requested by the agent
SituationUse thisWhy
Irreversible or money-moving action (refund, delete, send to customer)Code-level gate, default deny, exact arguments shownThe only guarantee that survives a wrong model decision
Reversible, low-impact action (internal note, tag)No gate; log itGating everything trains reviewers to rubber-stamp
High volume of one gated actionRules that auto-approve a narrow safe class (for example refunds under 50 USD with an exact duplicate), human for the restKeeps humans on the cases that need judgment
Approver does not respond in timeEscalate or park the task with its checkpointSilence is not consent

Observability: tracing every step

A trace is the step-by-step record of one run: every model call with its tokens and cost, every tool call with its arguments, every retry, approval, checkpoint, and stop reason. Ours is JSONL, one event per line, so it streams, greps, and loads into anything. These are three real events from runs/m08/agent_trace.jsonl:

json
{"ts": 1789975220.226, "run_id": "T-1001", "step": 0, "type": "llm_call", "thought": "Thought: customer reports a double charge. Find the refund policy first.\nPLAN: search_kb; read_article; get_account; get_invoice; issue_refund; add_internal_note; reply", "plan": ["search_kb", "read_article", "get_account", "get_invoice", "issue_refund", "add_internal_note", "reply"], "plan_version": 1, "tool_calls": [{"name": "search_kb", "args": {"query": "charged twice"}}], "input_tokens": 787, "output_tokens": 50, "cost_usd": 0.0001481, "latency_ms": 0.0}
{"ts": 1789975220.243, "run_id": "T-1001", "step": 4, "type": "tool_retry", "tool": "get_invoice", "attempt": 1, "error": "503 billing API timeout (call 1)", "backoff_s": 0.061}
{"ts": 1789975220.311, "run_id": "T-1001", "step": 4, "type": "tool_retry", "tool": "get_invoice", "attempt": 2, "error": "503 billing API timeout (call 2)", "backoff_s": 0.119}

Code explained

  • In simple words: each line is one thing that happened, with enough detail to replay the run in your head.
  • What happens: every event carries run_id (the task), step, and type. llm_call events carry the thought, the plan if it changed, the requested tool calls, tokens, cost, and latency (0.0 here because ScriptedLLM does not wait; a real chat fills it in). tool_retry events carry the error and the backoff chosen, so you can see flakiness per tool.
  • Comes out: this is data rather than a program. Your own timestamps will differ; the structure will not.

The trace is also your first diagnostic tool. When a run ends badly, read the trace before you touch the prompt. For example, the budget_usd scenario from Part 3 left a ticket with a refund and no reply. The last four events say why:

bash
tail -n 4 runs/m08/guards/budget_usd.jsonl | python -c "
import json, sys
for line in sys.stdin:
    e = json.loads(line)
    print(e['step'], e['type'], {k: v for k, v in e.items() if k in ('tool', 'status', 'used', 'next_call', 'cost_usd', 'preview')})
"

Code explained

  • In simple words: print the last four events of the run that stopped on budget, keeping only the fields that explain the stop.
  • What happens: tail takes the end of the JSONL file and the Python one-liner prints step, type, and a few fields.
  • Comes out:
text
  6 tool_call {'tool': 'add_internal_note'}
  6 tool_result {'tool': 'add_internal_note', 'preview': '{"saved": true, "note_count": 1}'}
  7 checkpoint {}
  7 stop {'status': 'budget_usd', 'used': 0.014571, 'next_call': 0.006006}

The internal note was written, then the guard refused call 8 because 0.014571 already spent plus 0.006006 projected exceeds 0.02. The diagnosis is "the output reserve is too conservative for this price point", not "the model is bad". Fix the budget (or route budget_usd stops to a person), not the prompt.

In production you would send the same events to a tracing system instead of a file. The common standard is OpenTelemetry; the MCP 2026-07-28 spec documents conventions for carrying OpenTelemetry trace context (traceparent, tracestate) in request _meta, so a trace can follow a call from your agent into an MCP server. Module 13 covers dashboards and retention; the rule for now is simple: if a step is not in the trace, you cannot debug it, bill it, or evaluate it.