Part 4: Is the agent worth it?
The same ticket, two designs
Now the measurement promised in Part 1. run_workflow handles T-1001 on a fixed path: regex for the email and invoice id, get_account and get_invoice through the same execute_tool (so it gets the same retries and approval gate), a code check for a duplicate charge, the refund, and then exactly one model call to write the reply as a DraftReply (Module 6's schema). The agent is run_agent from Part 2. Both use the same flaky billing API and the same approver.
examples/m08_workflow_vs_agent.py
"""Module 8: the same duplicate-charge ticket as a fixed workflow and as an agent.
Token counts are real counts (o200k_base) of the real prompts each design sends.
The replies come from ScriptedLLM policies, so the counts show what each DESIGN
costs, not how well a model performs. Latency is modeled from stated assumptions.
"""
from __future__ import annotations
import json
import random
import statistics
import time
from pathlib import Path
from m08_agent import (EMAIL_RE, INVOICE_RE, T1001, AgentConfig, Approver, BillingStore, ScriptedApprover,
SupportPolicy, Tracer, execute_tool, make_tools, run_agent)
from supportdesk.data import get_article
from supportdesk.llm import Usage
from supportdesk.pricing import cost_usd
from supportdesk.schemas import DraftReply
from supportdesk.stand_in import ScriptedLLM
WORK = Path("runs/m08/compare")
WORK.mkdir(parents=True, exist_ok=True)
for old in WORK.glob("*"):
old.unlink()
DRAFT_SYSTEM = ("You write Brightlane support replies. Use only the facts and the policy given. "
"Return JSON with keys reply, cited_articles, confidence.")
def run_workflow(ticket_text: str, task_id: str, tools: dict, chat_fn, approver: Approver,
tracer: Tracer) -> dict:
"""Fixed path: code does every lookup and check; the model writes only the customer reply."""
config = AgentConfig(backoff_base_s=0.05)
rng = random.Random(0)
email, invoice_id = EMAIL_RE.search(ticket_text).group(0), INVOICE_RE.search(ticket_text).group(0)
account = execute_tool(tools["get_account"], {"email": email}, approver, tracer, 1, config, rng)
invoice = execute_tool(tools["get_invoice"], {"invoice_id": invoice_id}, approver, tracer, 2, config, rng)
if not (account["ok"] and invoice["ok"]) or invoice["result"]["account_id"] != account["result"]["account_id"]:
execute_tool(tools["escalate_to_human"], {"ticket_id": task_id, "reason": "lookup failed or mismatch"},
approver, tracer, 3, config, rng)
return {"route": "escalated", "llm_calls": 0, "usage": Usage()}
inv = invoice["result"]
charges = inv["charges"]
duplicate = [c for c in charges[1:] if (c["amount"], c["date"]) == (charges[0]["amount"], charges[0]["date"])]
if not duplicate or inv["overpaid"] <= 0:
execute_tool(tools["escalate_to_human"], {"ticket_id": task_id, "reason": "no duplicate found"},
approver, tracer, 3, config, rng)
return {"route": "escalated", "llm_calls": 0, "usage": Usage()}
refund = execute_tool(tools["issue_refund"], {"invoice_id": invoice_id, "charge_id": duplicate[0]["charge_id"],
"amount_usd": inv["overpaid"], "reason": "duplicate charge"},
approver, tracer, 3, config, rng)
if not refund["ok"]:
execute_tool(tools["escalate_to_human"], {"ticket_id": task_id, "reason": refund["error"]},
approver, tracer, 4, config, rng)
return {"route": "escalated", "llm_calls": 0, "usage": Usage()}
policy = get_article("billing-refunds")
facts = {"refund_id": refund["result"]["refund_id"], "amount_usd": inv["overpaid"], "invoice_id": invoice_id}
messages = [{"role": "system", "content": DRAFT_SYSTEM},
{"role": "user", "content": f"Policy ({policy.id}):\n{policy.body}\n\nFacts: {json.dumps(facts)}\n\nTicket:\n{ticket_text}"}]
result = chat_fn(messages, response_format={"type": "json_object"})
draft = DraftReply.model_validate_json(result.text)
execute_tool(tools["add_internal_note"], {"ticket_id": task_id, "note": f"Workflow refunded {facts}"},
approver, tracer, 5, config, rng)
tracer.event("final", 6, reply=draft.reply)
return {"route": "refunded", "llm_calls": 1, "usage": result.usage, "reply": draft.reply}
def draft_responder(messages, kwargs) -> str:
facts = json.loads(messages[-1]["content"].split("Facts: ")[1].split("\n")[0])
return json.dumps({"reply": f"Hi, sorry about the double charge. We refunded the duplicate payment of "
f"{facts['amount_usd']:.2f} USD on invoice {facts['invoice_id']} (refund {facts['refund_id']}). "
"It goes back to your original card within 5 to 10 business days.",
"cited_articles": ["billing-refunds"], "confidence": "high"})
# Latency model (ASSUMPTIONS, not measurements): replace with your own numbers from Module 3.
CALL_OVERHEAD_MS = 600 # network plus time to first token, per sequential LLM call
OUTPUT_TOK_PER_S = 150 # decode speed
PREFILL_TOK_PER_S = 5000 # input processing speed
def modeled_latency_ms(calls: int, input_tokens: int, output_tokens: int) -> float:
return calls * CALL_OVERHEAD_MS + input_tokens / PREFILL_TOK_PER_S * 1000 + output_tokens / OUTPUT_TOK_PER_S * 1000
def row(name: str, calls: int, usage: Usage, wall_ms: float) -> dict:
return {"design": name, "llm_calls": calls, "input": usage.input_tokens, "output": usage.output_tokens,
"usd_oss120b": cost_usd(usage, "openai/gpt-oss-120b"), "usd_flash": cost_usd(usage, "gemini-3.5-flash"),
"latency_model_s": modeled_latency_ms(calls, usage.input_tokens, usage.output_tokens) / 1000,
"wall_ms": wall_ms}
if __name__ == "__main__":
# 1) Fixed workflow
from supportdesk.tokens import count_tokens
count_tokens("warm up the tokenizer so its load time is not charged to either design")
tools = make_tools(BillingStore(WORK / "workflow_billing.json"), flaky_failures=2)
tracer = Tracer(WORK / "workflow_trace.jsonl", "T-1001")
started = time.perf_counter()
wf = run_workflow(T1001, "T-1001", tools, ScriptedLLM(responder=draft_responder),
ScriptedApprover([True]), tracer)
wf_ms = (time.perf_counter() - started) * 1000
print("workflow route:", wf["route"], "| reply:", wf["reply"][:70], "...")
# 2) Agent, same ticket, same flaky tool
tools = make_tools(BillingStore(WORK / "agent_billing.json"), flaky_failures=2)
started = time.perf_counter()
st = run_agent(T1001, "T-1001", tools, ScriptedLLM(responder=SupportPolicy()),
approver=ScriptedApprover([True]), trace_path=WORK / "agent_trace.jsonl")
ag_ms = (time.perf_counter() - started) * 1000
print("agent status:", st.status, "| reply:", st.final[:70], "...")
rows = [row("workflow", wf["llm_calls"], wf["usage"], wf_ms),
row("agent", st.llm_calls, Usage(st.input_tokens, st.output_tokens), ag_ms)]
print(f"\n{'design':9} {'calls':>5} {'input':>6} {'output':>6} {'usd@oss-120b':>13} {'usd@3.5-flash':>14} "
f"{'latency model s':>15} {'plumbing ms':>11}")
for r in rows:
print(f"{r['design']:9} {r['llm_calls']:>5} {r['input']:>6} {r['output']:>6} {r['usd_oss120b']:>13.6f} "
f"{r['usd_flash']:>14.6f} {r['latency_model_s']:>15.1f} {r['wall_ms']:>11.0f}")
ratio = (rows[1]["input"] + rows[1]["output"]) / (rows[0]["input"] + rows[0]["output"])
print(f"agent/workflow tokens: {ratio:.1f}x | per 1,000 tickets at gemini-3.5-flash: "
f"workflow {rows[0]['usd_flash'] * 1000:.2f} USD, agent {rows[1]['usd_flash'] * 1000:.2f} USD")
# 3) Predictability: 50 agent runs with seeded detours (simulated variation, not a model)
calls, tokens = [], []
for seed in range(50):
store = BillingStore(WORK / f"noisy_{seed}.json")
s = run_agent(T1001, "T-1001", make_tools(store, flaky_failures=0), ScriptedLLM(responder=SupportPolicy(noise=0.3, seed=seed)),
approver=ScriptedApprover([True]), sleep=lambda x: None)
assert s.status == "done", s.status
calls.append(s.llm_calls)
tokens.append(s.input_tokens + s.output_tokens)
q = statistics.quantiles(tokens, n=20)
print(f"\nagent over 50 noisy runs: LLM calls min {min(calls)} median {statistics.median(calls):.0f} max {max(calls)}; "
f"tokens median {statistics.median(tokens):.0f}, p95 {q[18]:.0f}, max {max(tokens)}; "
f"stdev {statistics.stdev(tokens):.0f}")
print("workflow over any number of runs: LLM calls always 1, tokens always", rows[0]["input"] + rows[0]["output"])Code explained
- In simple words: run T-1001 both ways, count what each costs, then run the agent 50 more times with seeded detours to see how much its cost moves.
- What happens:
run_workflowis a plain function: every decision is anifin code. The only model call is the reply, with the policy text and the facts (refund id, amount, invoice) in the prompt andresponse_format={"type": "json_object"}. Its reply is validated withDraftReply.model_validate_json. Any surprise (lookup failure, account mismatch, no duplicate, refund rejected) routes toescalate_to_human.draft_responderis the scripted stand-in for that one call.modeled_latency_msis a latency model, not a measurement: 600 ms of overhead per sequential call, 5,000 input tokens per second of prefill, 150 output tokens per second of decode. These are assumptions; replace them with the time to first token and decode speed you measured for your provider in Module 3.rowprices each design withcost_usdat two price points frompricing.py.- The tokenizer is warmed up before timing so its load time is not charged to whichever design runs first. "Plumbing ms" is real wall time of the Python around the model calls; most of it is the two real backoff sleeps (about 0.06 s and 0.12 s).
- The last block runs the agent 50 times with
SupportPolicy(noise=0.3, seed=...). The detours are simulated variation, a stand-in for a real model sometimes searching again or re-reading. It shows the shape of the effect, not its size for any real model.
- Comes out: (ScriptedLLM-driven; token counts real, latency modeled)
workflow route: refunded | reply: Hi, sorry about the double charge. We refunded the duplicate payment o ...
agent status: done | reply: Hi, sorry about the double charge. We refunded the duplicate payment o ...
design calls input output usd@oss-120b usd@3.5-flash latency model s plumbing ms
workflow 1 242 71 0.000079 0.001002 1.1 188
agent 8 9362 377 0.001631 0.017436 9.2 226
agent/workflow tokens: 31.1x | per 1,000 tickets at gemini-3.5-flash: workflow 1.00 USD, agent 17.44 USD
agent over 50 noisy runs: LLM calls min 8 median 10 max 11; tokens median 13425, p95 15581, max 15668; stdev 1632
workflow over any number of runs: LLM calls always 1, tokens always 313Reading the comparison
| Measure | Workflow | Agent | Where the number comes from |
|---|---|---|---|
| LLM calls | 1 | 8 | Counted by the loop |
| Tokens (input + output) | 313 | 9,739 | o200k_base counts of the prompts actually sent |
| Cost per ticket, gpt-oss-120b | 0.000079 USD | 0.001631 USD | pricing.py math |
| Cost per ticket, gemini-3.5-flash | 0.0010 USD | 0.0174 USD | pricing.py math |
| Latency | about 1.1 s | about 9.2 s | Modeled from stated assumptions |
| Predictability | Same path, same tokens every run | 8 to 11 calls; tokens stdev 1,632 in the simulated runs | 50 seeded runs |
The agent costs about 31 times the tokens for the same outcome. Two reasons, both visible in the data:
- Tool definitions ride along on every call. The seven tool schemas are 615 tokens. Over 8 calls that is 4,920 of the 9,362 input tokens, over half. Module 2's prompt caching would discount this stable prefix on providers that support it; it does not remove the round trips.
- Context grows every step. Call 1 sent 787 tokens; the last call sent the whole history. Total input grows roughly with the square of the number of steps.
Is the difference within noise? No: the workflow's number is exact and the agent's minimum across 50 runs (8 calls, about 9,700 tokens) is still about 30 times higher. What is uncertain is how a real model's path length varies, which is why you should rerun the 50-run block with chat and your key before you quote a p95 to anyone.
For duplicate charges, the workflow wins on every measure we have, and it is easier to certify: a compliance reviewer can read run_workflow in five minutes. The agent earns its cost only on tickets where you cannot write the path down, and even then the right move is often to turn the paths the agent discovers in production (read them from the traces) into workflows, keeping the agent for the long tail.
| Situation | Use this | Why |
|---|---|---|
| High-volume ticket type with a known procedure (duplicate charges, invoice copies) | Workflow | 31x fewer tokens here, fixed latency, auditable |
| Mixed queue where most tickets fit known procedures | Router to workflows, agent as fallback for the rest | Pays agent prices only on the long tail |
| Rare, open-ended investigations (why did automations stop?) | Agent with guards and a budget | The path is unknowable in advance; cap the cost instead |
| You are not sure which | Log agent trajectories for a few weeks, then promote common paths into workflows |