CourseLarge Language Models · Module 8: Agents · part 36 of 80
Part 36 · Module 8: Agents

Part 1: Workflows and agents

6 min read·22 Sept 2026

By the end of this module, you'll have:

  • A guarded agent loop for the Brightlane support desk (examples/m08_agent.py) that works a duplicate-charge ticket with seven tools, asks a human before it refunds anything, retries a flaky billing API with backoff, writes every step to a JSONL trace, and resumes after a crash without refunding twice.
  • Termination guards and a per-task budget (max steps, no-progress detection, repeated-action detection, token and dollar ceilings) that you have watched fire, one scenario at a time.
  • A measured comparison of the same ticket handled as a fixed workflow and as an agent: 1 LLM call against 8, 313 tokens against 9,739, and what that means per 1,000 tickets.
  • A working picture of the capability layer: MCP messages on the wire (spec version 2026-07-28), a resource-limited code execution tool and what it does not protect against, a workspace-confined file tool, and subagents.
  • An orchestrator-worker version of the agent with its token bill itemized, plus the probability math of failure compounding.
  • A trajectory evaluator that passes a careful run and fails a lucky one and an unsafe one, even though a final-answer grader passes all three.

Prerequisites: Modules 1 to 7. You need the tool-calling loop and "errors go back to the model" from Module 6, token counting and pricing.py from Module 2, and context budgeting from Module 7. Working Python; no machine learning background.

Where we are: Module 7 made the assistant's context deliberate: what goes in, in what order, and what gets dropped. So far the assistant only reads and drafts. In this module it acts. An agent is a model in a loop, choosing tools and reading their results until the job is done, and everything from Module 7 now repeats on every turn of that loop.

How this module is organized

PartWhat it covers
Part 1: Workflows and agentsThe two shapes, what autonomy costs, and cheapest-thing-that-works as the default
Part 2: The core loopPerceive, decide, act, observe; the full agent file; a ReAct-style trace; planning upfront, incrementally, and on failure
Part 3: Stopping, budgets, and recoveryEvery termination guard, the budget ceiling, replanning, and crash recovery, each triggered on purpose
Part 4: Is the agent worth it?The duplicate-charge ticket as a workflow and as an agent: calls, tokens, cost, latency, predictability
Part 5: Reliability in the loopRetries with backoff, idempotent actions, checkpoints, approval gates, and traces you can read
Part 6: CapabilitiesMCP as the interop layer, code execution, computer and browser use, files, subagents
Part 7: Multi-agent systemsOrchestrator and workers, parallel exploration, communication cost, failure compounding, when it is worth it
Part 8: Evaluating trajectoriesGrading what the agent did, not only what it said

All examples run in order from the root of your supportdesk copy, with PYTHONPATH=., and write their files under runs/m08/. Here is the setup once:

bash
cd supportdesk
source .venv/bin/activate          # or however you activated the course venv in Module 1
export PYTHONPATH=.
mkdir -p runs/m08
python -c "import supportdesk.llm, supportdesk.pricing, supportdesk.kb_search; print('ok')"

Code explained

  • In simple words: get into the project, make its package importable, and create a folder for this module's traces and checkpoints.
  • What happens: PYTHONPATH=. lets every example import supportdesk.*. The example scripts live in examples/ and import each other (m08_guards.py imports m08_agent.py), which works because Python puts a script's own folder on the import path. The last line checks that the canonical modules this module relies on import cleanly.
  • Comes out:
text
  ok

A word on honesty before we start. No API key was available while this module was built. Every agent run you see here is driven by ScriptedLLM from Module 1, with a small rule-based policy (a function that reads the conversation and decides the next tool call, the way a model would). That makes the loop, guards, traces, retries, and budgets real and testable, and every token count is a real count of the real prompts the loop sends. It does not tell you how well a model makes decisions. Wherever that matters, the text says so and tells you what to run with a key.

Part 1: Workflows and agents

Two shapes for the same job

Maya's team gets tickets like T-1001: "my card was charged 288 USD twice on 3 September (invoice INV-2026-004512). Please refund the duplicate." Handling it takes a few lookups, one check, one irreversible action, and a reply.

There are two ways to build that with an LLM.

A workflow is a fixed path written in code. Code extracts the invoice number, code fetches the invoice, code checks for two identical charges, code asks a human to approve the refund, and the model is called once, at the end, to write the reply. Anthropic's "Building effective agents" (December 2024) defines workflows as "systems where LLMs and tools are orchestrated through predefined code paths."

An agent is a model directing its own path. The model gets the ticket and the tools, decides which tool to call, reads the result, and decides again, until it chooses to stop. The same post defines agents as systems where "LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks."

mermaidCopy

text
flowchart LR
  subgraph Workflow["Workflow: code decides the path"]
    W1[Extract invoice id] --> W2[get_account + get_invoice] --> W3{Duplicate?}
    W3 -- yes --> W4[Approval gate] --> W5[issue_refund] --> W6[LLM writes reply]
    W3 -- no --> W7[escalate_to_human]
  end
  subgraph Agent["Agent: the model decides the path"]
    A1[LLM decides] --> A2[Run chosen tool] --> A3[Append observation] --> A1
    A1 -- no tool call --> A4[Final reply]
  end

What autonomy costs

Autonomy is paid for three ways, and all three show up in Part 4 as numbers:

  • Latency. Each decision is a model call, and the calls are sequential: the model cannot decide step 5 before it has seen the result of step 4.
  • Spend. Every call resends the whole conversation so far, plus the tool definitions. Context grows each step, so tokens grow faster than steps.
  • Unpredictability. The path can differ between runs of the same ticket. That makes cost harder to forecast, bugs harder to reproduce, and behavior harder to certify.

What you buy with it is flexibility: the agent can handle tickets whose path you did not write down in advance. The Anthropic post puts the trade plainly: "agentic systems often trade latency and cost for better task performance", and its advice is to "find the simplest solution possible, and only increase complexity when needed."

Cheapest thing that works

The default in this course is the cheapest design that meets the quality bar, measured on your eval set (Module 10), and you move up a row only when the row below fails on real tickets.

SituationUse thisWhy
One input, one output, no lookups (classify a ticket, draft from given facts)A single LLM call, maybe with structured output (Module 6)One round trip; nothing to orchestrate
The steps are known in advance and the same for every ticket of this typeA workflow: code for lookups and checks, the LLM only where language is neededPredictable path, cost, and latency; easy to test and certify
A handful of known paths, chosen by the contentA router plus a few workflows (Module 5)Still predictable; the model only picks the branch
The steps depend on what earlier steps reveal, and you cannot enumerate the pathsA single agent loop with guardsFlexibility, paid for in calls, tokens, and variance
Broad, parallelizable exploration that overflows one contextOrchestrator plus workers (Part 7)More total tokens, but work runs in parallel with fresh contexts

For duplicate charges the steps are known, so the honest answer for T-1001 is "workflow". We still build the agent, because Brightlane's queue has plenty of tickets whose path is not known in advance ("my board automations stopped and I think it is related to our plan downgrade"), and because an agent is the only way to see what the guards, traces, and evaluations in this module protect against.