CourseLarge Language Models · Module 8: Agents · part 38 of 80
Part 38 · Module 8: Agents

Part 3: Stopping, budgets, and recovery

8 min read·22 Sept 2026

Why a loop needs brakes

A model deciding when to stop is a model that can fail to stop. Real agents loop for mundane reasons: a search that never finds anything, a tool that keeps returning the same error, a model that "double checks" the same record forever, a context so long that the goal gets lost (Module 7's context rot). Each of those burns money on every turn. So the loop, not the model, owns termination:

  • Max steps: a hard ceiling on iterations. Simple, and always on.
  • No-progress detection: stop after N consecutive steps that produced no new observation. We fingerprint each observation (tool name plus result) and count steps that added no new fingerprint.
  • Repeated-action detection: stop when the model issues the identical tool call (same name, same arguments) more than repeat_limit times.
  • Budget ceilings: stop before a call that would push the task over its token or dollar limit. The check happens before spending, using prompt_tokens for the input and output_reserve for the reply.

When a guard fires, the run ends with a status (max_steps, no_progress, repeated_action, budget_tokens, budget_usd) that the caller can route on, typically to a human queue with the trace attached.

Triggering every guard on purpose

This script runs eight scenarios through the same loop. Two small policies misbehave on purpose: StuckSearcher keeps rephrasing a search that finds nothing, and Looper asks for the same account again and again.

python
"""Module 8: every termination guard, the budget ceiling, replanning, and crash recovery, one scenario each.

All runs are driven by ScriptedLLM policies (not a model), so the numbers are real
counts of the loop's bookkeeping, not measurements of model behavior.
"""
from __future__ import annotations

import json
from pathlib import Path

from m08_agent import (T1001, AgentConfig, BillingStore, ScriptedApprover, SimulatedCrash, SupportPolicy,
                       history, make_tools, print_trace, run_agent)
from supportdesk.stand_in import ScriptedLLM

WORK = Path("runs/m08/guards")
WORK.mkdir(parents=True, exist_ok=True)
for old in WORK.glob("*"):
    old.unlink()


class StuckSearcher(SupportPolicy):
    """Keeps rephrasing the customer's words; every query returns nothing new."""
    QUERIES = ["charged twice", "paid twice", "twice charged", "charged 2x", "twice paid"]

    def __call__(self, messages, kwargs):
        n = len(history(messages))
        return self.reply(messages, "Thought: try other words.", "search_kb",
                          {"query": self.QUERIES[n % len(self.QUERIES)]}, kwargs.get("tools"))


class Looper(SupportPolicy):
    """Calls get_account with the same arguments again and again."""

    def __call__(self, messages, kwargs):
        return self.reply(messages, "Thought: check the account.", "get_account",
                          {"email": "dana.k@northwind.example"}, kwargs.get("tools"))


def scenario(name: str, policy, *, config: AgentConfig | None = None, approve=(True,), flaky: int = 2,
             crash_at_step: int | None = None, resume: bool = False) -> dict:
    store = BillingStore(WORK / f"{name}_billing.json")
    tools = make_tools(store, flaky_failures=flaky)
    kwargs = dict(config=config, approver=ScriptedApprover(list(approve)),
                  trace_path=WORK / f"{name}.jsonl", checkpoint_path=WORK / f"{name}_ckpt.json",
                  sleep=lambda s: None)
    try:
        state = run_agent(T1001, "T-1001", tools, ScriptedLLM(responder=policy), crash_at_step=crash_at_step, **kwargs)
    except SimulatedCrash as exc:
        print(f"  {name}: {exc}; refunds on disk: {len(json.loads((WORK / f'{name}_billing.json').read_text()))}")
        if not resume:
            raise
        store = BillingStore(WORK / f"{name}_billing.json")        # a fresh process reloads everything
        tools = make_tools(store, flaky_failures=0)
        kwargs["approver"] = ScriptedApprover([True])                 # Maya approves the re-sent request
        state = run_agent(T1001, "T-1001", tools, ScriptedLLM(responder=SupportPolicy()), **kwargs)
    return {"scenario": name, "status": state.status, "steps": state.step, "llm_calls": state.llm_calls,
            "tokens": state.input_tokens + state.output_tokens, "cost_usd": round(state.cost_usd, 6),
            "refunds": len(store.refunds), "plan_version": state.plan.version}


rows = [
    scenario("happy_path", SupportPolicy()),
    scenario("max_steps", SupportPolicy(), config=AgentConfig(max_steps=4)),
    scenario("no_progress", StuckSearcher()),
    scenario("repeated_action", Looper()),
    scenario("budget_usd", SupportPolicy(), config=AgentConfig(price_model="gemini-3.5-flash", max_usd=0.02)),
    scenario("approval_denied", SupportPolicy(), approve=(False,)),
    scenario("tool_down_replan", SupportPolicy(), flaky=10),
    scenario("crash_resume", SupportPolicy(), crash_at_step=5, resume=True),
]
print()
print(f"{'scenario':18} {'status':16} {'steps':>5} {'calls':>5} {'tokens':>7} {'cost_usd':>9} {'refunds':>7} {'plan_v':>6}")
for r in rows:
    print(f"{r['scenario']:18} {r['status']:16} {r['steps']:>5} {r['llm_calls']:>5} {r['tokens']:>7} "
          f"{r['cost_usd']:>9.6f} {r['refunds']:>7} {r['plan_version']:>6}")

print("\nTrace of tool_down_replan (steps 4 onward):")
lines = [ln for ln in (WORK / "tool_down_replan.jsonl").read_text().splitlines() if json.loads(ln)["step"] >= 4]
(WORK / "replan_tail.jsonl").write_text("\n".join(lines) + "\n")
print_trace(WORK / "replan_tail.jsonl")
print("\nTrace of crash_resume (steps 5 onward):")
lines = [ln for ln in (WORK / "crash_resume.jsonl").read_text().splitlines() if json.loads(ln)["step"] >= 5]
(WORK / "crash_tail.jsonl").write_text("\n".join(lines) + "\n")
print_trace(WORK / "crash_tail.jsonl")

Code explained

  • In simple words: one run per guard, each set up so that exactly that guard (or recovery path) fires, then a summary table and two traces.
  • What happens:
    • scenario builds fresh tools and a fresh billing file, runs the agent with a trace and a checkpoint, and returns the totals. sleep=lambda s: None skips the real backoff waits so the script is fast; the backoff values still appear in the trace.
    • max_steps sets max_steps=4. no_progress uses StuckSearcher, whose queries all return []. repeated_action uses Looper. budget_usd prices the run as gemini-3.5-flash (1.50 USD input, 9.00 USD output per million tokens in pricing.py) with a 0.02 USD ceiling. approval_denied has Maya reject the refund. tool_down_replan makes get_invoice fail 10 times, more than the 3 attempts allowed.
    • crash_resume crashes at step 5, after the refund ran but before the checkpoint was written. The except branch then does what a restarted worker would: reload the billing store from disk, rebuild the tools, and call run_agent again with the same checkpoint path, which resumes from the last checkpoint.
  • Comes out: (ScriptedLLM-driven; counts are real)
text
    crash_resume: crashed during step 5; refunds on disk: 1

  scenario           status           steps calls  tokens  cost_usd refunds plan_v
  happy_path         done                 8     8    9739  0.001631       1      3
  max_steps          max_steps            4     4    3811  0.000641       0      2
  no_progress        no_progress          4     4    3428  0.000548       0      0
  repeated_action    repeated_action      2     3    2664  0.000431       0      0
  budget_usd         budget_usd           7     7    8084  0.014571       1      3
  approval_denied    done                 8     8    9628  0.001597       0      3
  tool_down_replan   done                 7     7    7791  0.001291       0      3
  crash_resume       done                 8     8    9753  0.001633       1      3

  Trace of tool_down_replan (steps 4 onward):
  [4] THINK  Thought: fetch the invoice named in the ticket to see the charges.
  [4] ACT    get_invoice({"invoice_id": "INV-2026-004512"})
  [4] RETRY  get_invoice attempt 1: 503 billing API timeout (call 1); backoff 0.061 s
  [4] RETRY  get_invoice attempt 2: 503 billing API timeout (call 2); backoff 0.119 s
  [4] RETRY  get_invoice attempt 3: 503 billing API timeout (call 3); giving up
  [4] REPLAN_REQUESTED {'old_plan': ['get_account', 'get_invoice', 'issue_refund', 'add_internal_note', 'reply']}
  [5] THINK  Thought: billing API is down after retries, so I cannot verify the charge. Hand it to a person.
  [5] ACT    escalate_to_human({"ticket_id": "T-1001", "reason": "get_invoice unavailable; duplicate charge unv)
  [5] OBS    {"queued_for": "billing-team", "position": 1}
  [6] THINK  Thanks for flagging this. A member of our billing team is checking the duplicate charge and will
  [7] STOP   {'status': 'done'}

  Trace of crash_resume (steps 5 onward):
  [5] THINK  Thought: two charges of 288.0 on 2026-09-03; overpaid 288.0. Refund the second one.
  [5] ACT    issue_refund({"invoice_id": "INV-2026-004512", "charge_id": "ch_7Hq2", "amount_usd": 288.0, ")
  [5] GATE   issue_refund approved=True by human:maya
  [5] OBS    {"refund_id": "RF-00001", "invoice_id": "INV-2026-004512", "charge_id"
  [5] CRASH  {'note': 'simulated crash before checkpoint'}
  [5] RESUME {'from_checkpoint': 'crash_resume_ckpt.json', 'llm_calls': 5}
  [5] THINK  Thought: two charges of 288.0 on 2026-09-03; overpaid 288.0. Refund the second one.
  [5] ACT    issue_refund({"invoice_id": "INV-2026-004512", "charge_id": "ch_7Hq2", "amount_usd": 288.0, ")
  [5] GATE   issue_refund approved=True by human:maya
  [5] OBS    {"already_refunded": true, "refund_id": "RF-00001", "invoice_id": "INV
  [6] THINK  Thought: record what I did for the team.
  [6] ACT    add_internal_note({"ticket_id": "T-1001", "note": "Refunded duplicate ch_7Hq2 (288.0 USD) as RF-00)
  [6] OBS    {"saved": true, "note_count": 1}
  [7] THINK  Hi, sorry about the double charge. We refunded the duplicate payment of 288.00 USD on invoice IN
  [8] STOP   {'status': 'done'}

How to read the table, row by row:

  • happy_path: 8 steps, one refund. The baseline.
  • max_steps: stopped after 4 model calls with no refund. A hard cap is crude but it bounds the worst case: the most this task can ever cost is 4 calls.
  • no_progress: the first empty search was new information (a new fingerprint); the next three returned the same [], so after three stale steps the loop stopped. Without this guard, StuckSearcher would spin until max_steps.
  • repeated_action: the third identical get_account call was refused before it ran. The model made 3 calls but only 2 steps completed.
  • budget_usd: this one is worth studying. The full run at Gemini prices would cost 0.0174 USD (Part 4), under the 0.02 ceiling, yet the guard stopped it after 7 calls at 0.0146 USD. Why? Before call 8 the loop projected about 1,600 input tokens plus a 400-token output reserve. At 9 USD per million output tokens, the reserve alone is 0.0036 USD, and 0.0146 plus the projection crosses 0.02. The actual final reply was far shorter than 400 tokens. A budget guard with a reserve is conservative by design; size the reserve from the 95th percentile of real reply lengths, and when it trips, hand off to a person rather than leaving the ticket half done. Here the refund went through but the customer reply was never written, which is exactly the state you do not want to leave silently.
  • approval_denied: Maya rejected, the policy escalated instead of retrying the refund, zero refunds.
  • tool_down_replan: three attempts with backoff, then a replan request, then escalation. The agent did not guess at the charges it could not see.
  • crash_resume: the process died right after the refund. On resume, the loop replayed step 5 from the checkpoint, the model asked for the same refund, Maya approved again, and the billing store answered already_refunded: true with the original RF-00001. One refund on disk, not two.

Notice one more thing in crash_resume: its token total (9,753) is barely above the happy path, but the model call made just before the crash is missing from the state's totals, because that state was never saved. The trace file still has it. When you reconcile spend, reconcile from traces or provider bills, not from checkpointed counters.

SituationUse thisWhy
Every agent, alwaysmax_steps plus a dollar ceilingBounds the worst case even when every other guard misses
Searches or reads that can come back empty or identicalNo-progress detection on observation fingerprintsCatches loops where arguments change but nothing new is learned
A model that re-requests the same callRepeated-action detectionCheap, exact, and catches the most common loop
A cost-sensitive queuePer-task token and dollar budgets checked before each call, plus a batch ceilingStops overspend before it happens, not after the bill
A guard fires mid-taskRoute to a human with the trace and the checkpointA half-done task needs a person more than a retry