Part C: Model output as untrusted input
Everything a model writes, including tool arguments, is untrusted input, the same as a form a stranger filled in. The model might be wrong, confused by a long context, or steered by a prompt injection hidden in a ticket (Module 11 covers injection in depth). The rule of this part: nothing the model writes reaches a privileged operation without code checking it first. A privileged operation is one that changes money, access, or data, like issue_refund.
Validating arguments before execution
"""Model output is untrusted input: what stops a bad refund before money moves."""
from __future__ import annotations
import dataclasses
import time
from examples.m06_tools import SYSTEM, TICKET, TOOLS, WITH_REFUNDS, ToolContext, call, execute, run_tool_loop
from supportdesk.llm import ChatResult
from supportdesk.stand_in import ScriptedLLM
asked: list[str] = []
def approver(name, args) -> bool:
asked.append(f"{name} {args.amount_usd} USD")
return True
print("=== 1. Refund larger than the invoice (576 USD on a 288 USD invoice)")
ctx = ToolContext("cus_acme", WITH_REFUNDS, approver=approver)
outcome = execute(call("r1", "issue_refund", invoice_id="INV-2026-004513", amount_usd=576.0, reason="duplicate_charge"), ctx)
print(outcome.content)
print("approver was asked:", asked, "| refunded_usd:", ctx.db["INV-2026-004513"].refunded_usd)
print("\n=== 2. Amount outside the schema bound (20000 USD)")
print(execute(call("r2", "issue_refund", invoice_id="INV-2026-004513", amount_usd=20000, reason="duplicate_charge"), ctx).content)
print("\n=== 3. The error goes back to the model, which corrects itself (ScriptedLLM, plumbing only)")
llm = ScriptedLLM(replies=[
ChatResult("", tool_calls=[call("r3", "issue_refund", invoice_id="INV-2026-004513", amount_usd=576.0, reason="duplicate_charge")]),
ChatResult("", tool_calls=[call("r4", "issue_refund", invoice_id="INV-2026-004513", amount_usd=288.0, reason="duplicate_charge")]),
ChatResult("We refunded the duplicate charge of 288 USD. It arrives in 5 to 10 business days."),
])
messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": TICKET}]
result = run_tool_loop(llm, messages, ctx, verbose=True)
print("final:", result.text, "| approver asked:", asked)
print("\n=== 4. A human says no")
deny_ctx = ToolContext("cus_acme", WITH_REFUNDS, approver=lambda name, args: False)
print(execute(call("r5", "issue_refund", invoice_id="INV-2026-004513", amount_usd=288.0, reason="duplicate_charge"), deny_ctx).content)
print("refunded_usd:", deny_ctx.db["INV-2026-004513"].refunded_usd)
print("\n=== 5. Two parallel refunds of the same duplicate")
par_ctx = ToolContext("cus_acme", WITH_REFUNDS, approver=lambda name, args: True)
llm = ScriptedLLM(replies=[
ChatResult("", tool_calls=[call(f"p{i}", "issue_refund", invoice_id="INV-2026-004513", amount_usd=288.0,
reason="duplicate_charge") for i in (1, 2)]),
ChatResult("Refunded once."),
])
run_tool_loop(llm, messages, par_ctx, verbose=True)
print("refunded_usd:", par_ctx.db["INV-2026-004513"].refunded_usd)
print("\n=== 6. A tool that hangs is abandoned after its timeout")
def slow_invoice(args, ctx):
time.sleep(3) # imagine the billing API is down
return {"never": "returned"}
slow_registry = {**TOOLS, "get_invoice": dataclasses.replace(TOOLS["get_invoice"], run=slow_invoice, timeout_s=0.5)}
slow = execute(call("t1", "get_invoice", invoice_id="INV-2026-004512"), par_ctx, registry=slow_registry)
print(f"{slow.content} after {slow.ms:.0f} ms")
Code explained
- In simple words: try to move too much money in several ways, and watch each attempt stop at a different gate.
- What happens: case 1 asks for 576 USD on a 288 USD invoice, the classic "refund both charges" mistake. Case 2 exceeds the schema bound. Case 3 runs the loop with a
ScriptedLLMthat makes the case 1 mistake, reads the error, and corrects itself (scripted plumbing, not model behavior). Case 4 uses an approver that says no. Case 5 sends two identical refunds in one parallel turn. Case 6 replacesget_invoicewith a function that hangs for 3 s under a 0.5 s timeout. - Comes out:
=== 1. Refund larger than the invoice (576 USD on a 288 USD invoice)
{"error": "Refund of 576.00 USD exceeds the refundable amount of 288.00 USD on INV-2026-004513.", "error_type": "rejected"}
approver was asked: [] | refunded_usd: 0.0
=== 2. Amount outside the schema bound (20000 USD)
{"error": "Invalid arguments: amount_usd: Input should be less than or equal to 10000", "error_type": "bad_arguments"}
=== 3. The error goes back to the model, which corrects itself (ScriptedLLM, plumbing only)
turn 1 ERR issue_refund({"invoice_id": "INV-2026-004513", "amount_usd": 576.0, "reason": "duplicate_charge"})
-> {"error": "Refund of 576.00 USD exceeds the refundable amount of 288.00 USD on INV-2026-004513.", "error_type"
turn 2 ok issue_refund({"invoice_id": "INV-2026-004513", "amount_usd": 288.0, "reason": "duplicate_charge"})
-> {"status": "refunded", "invoice_id": "INV-2026-004513", "amount_usd": 288.0, "refundable_usd_now": 0.0, "arriv
final: We refunded the duplicate charge of 288 USD. It arrives in 5 to 10 business days. | approver asked: ['issue_refund 288.0 USD']
=== 4. A human says no
{"error": "This action needs human approval and was not approved now. Do not retry it; tell the customer a teammate will review the request.", "error_type": "not_approved"}
refunded_usd: 0.0
=== 5. Two parallel refunds of the same duplicate
turn 1 ok issue_refund({"invoice_id": "INV-2026-004513", "amount_usd": 288.0, "reason": "duplicate_charge"})
-> {"status": "refunded", "invoice_id": "INV-2026-004513", "amount_usd": 288.0, "refundable_usd_now": 0.0, "arriv
turn 1 ERR issue_refund({"invoice_id": "INV-2026-004513", "amount_usd": 288.0, "reason": "duplicate_charge"})
-> {"error": "Refund of 288.00 USD exceeds the refundable amount of 0.00 USD on INV-2026-004513.", "error_type":
refunded_usd: 288.0
=== 6. A tool that hangs is abandoned after its timeout
{"error": "get_invoice took longer than 0.5s and was abandoned. Try once more or continue without it.", "error_type": "timeout"} after 514 ms
The key line is under case 1: approver was asked: []. The oversized refund was rejected by the policy check before a human was even asked, so Maya never sees a request that should not exist, and refunded_usd stayed 0.0. In case 3 the approver was asked exactly once, for the corrected 288 USD. In case 5 exactly one refund succeeded; which of the two wins can change between runs, but the total never exceeds the invoice, thanks to the lock around check and update. In case 6 the call returned a timeout error after about 0.5 s (514 ms here; it varies).
Sandboxing and permission scoping
Permission scoping means each conversation gets only the capabilities its job needs. The executor enforces three scopes, all set by code:
- Which tools: a triage-only session gets
READ_ONLY; only the refund workflow getsWITH_REFUNDS. You saw a read-only session's refund attempt returnnot_available. - Which data:
get_invoiceonly sees the ticket's own customer, taken fromToolContext, never from arguments. - Which actions without a human: none that move money.
issue_refundalways needs approval.
A sandbox limits what running code can do even when it misbehaves. Tools that only call your own APIs, like these three, are sandboxed by the checks above. Tools that run code or touch files, such as a "run this spreadsheet formula" or "convert this attachment" tool, need process isolation, because a thread cannot be stopped from outside.
"""Sandboxing basics: run untrusted work in a separate process with hard limits.
A thread cannot be killed, but a process can. This runs a snippet of Python in
a child process with a wall-clock timeout, a CPU-time limit, a memory limit, an
empty environment (no API keys), and a throwaway working directory. It is a
first layer, not a complete sandbox: see the text for containers and gVisor.
"""
from __future__ import annotations
import resource
import subprocess
import sys
import tempfile
import time
def _limits() -> None: # runs in the child just before the snippet starts
resource.setrlimit(resource.RLIMIT_CPU, (1, 1)) # 1 s of CPU
resource.setrlimit(resource.RLIMIT_AS, (256 * 2**20, 256 * 2**20)) # 256 MB of memory
def run_sandboxed(code: str, timeout_s: float = 5.0) -> dict:
started = time.perf_counter()
with tempfile.TemporaryDirectory() as workdir:
try:
proc = subprocess.run([sys.executable, "-I", "-c", code], cwd=workdir, env={}, capture_output=True,
text=True, timeout=timeout_s, preexec_fn=_limits)
outcome = {"exit": proc.returncode, "stdout": proc.stdout.strip()[:80], "stderr": (proc.stderr.strip().splitlines() or [""])[-1][:80]}
except subprocess.TimeoutExpired:
outcome = {"exit": None, "error": f"killed after {timeout_s}s wall clock"}
outcome["ms"] = round((time.perf_counter() - started) * 1000)
return outcome
CASES = {
"normal": "print(sum([288, 288]) / 2)",
"reads env": "import os; print(sorted(os.environ))",
"cpu loop": "while True: pass",
"sleeps": "import time; time.sleep(60)",
"memory bomb": "x = bytearray(1024 * 2**20); print(len(x))",
}
for name, code in CASES.items():
print(f"{name:<12} {run_sandboxed(code)}")
Code explained
- In simple words: run untrusted snippets in a separate Python process with a time limit, CPU and memory limits, no environment variables, and a throwaway folder.
- What happens:
run_sandboxedstartspython -I -c code(isolated mode ignores user site packages andPYTHON*variables) withenv={}so no API keys are visible,cwdset to a temporary directory that is deleted afterwards, andpreexec_fn=_limits, which setsRLIMIT_CPUto 1 s of CPU time andRLIMIT_ASto 256 MB of address space in the child before it starts.subprocess.run(timeout=5.0)kills the child if the wall clock runs out. Five cases exercise each limit. - Comes out:
normal {'exit': 0, 'stdout': '288.0', 'stderr': '', 'ms': 138}
reads env {'exit': 0, 'stdout': "['LC_CTYPE']", 'stderr': '', 'ms': 104}
cpu loop {'exit': None, 'error': 'killed after 5.0s wall clock', 'ms': 5024}
sleeps {'exit': None, 'error': 'killed after 5.0s wall clock', 'ms': 5044}
memory bomb {'exit': 1, 'stdout': '', 'stderr': 'MemoryError', 'ms': 86}
The normal snippet works; the environment is empty apart from a locale variable Python sets itself; the memory bomb fails with MemoryError inside the child; the sleeper is killed at the 5 s wall-clock limit, which a CPU limit alone would never catch. The CPU loop is the honest surprise. RLIMIT_CPU counts CPU time, not wall time, and on this build machine (a VM) the child was killed by the wall-clock timeout in this run and by the CPU limit (exit code -9, after about 3.6 to 4.3 s of wall time) in two earlier runs. That is why you set both limits. This is a first layer, not a complete sandbox: it does not block network access or file reads outside the folder. For those, run tools in a container with no network and a read-only filesystem, or use a stronger isolation layer such as gVisor or a microVM.
Timeouts belong on every tool, not only risky ones. The executor's thread timeout in case 6 returns control to the loop after 0.5 s, but the stuck thread keeps running until its function returns (that demo script takes about 3 s to exit for this reason). For work that can hang forever, use a process, as above, or an HTTP client timeout inside the tool.
Never letting a model's output reach a privileged operation unchecked
Put the whole guard layer together and you get a short checklist. Every item is in execute and tested:
| Check | Where | What it stops |
|---|---|---|
| Tool offered in this context | ctx.allowed_tools | Invented tools, tools outside this workflow's scope |
| Arguments parse and validate | pydantic models with enums, patterns, bounds | Wrong formats, invented arguments, absurd amounts |
| Policy check against live data | precheck, then again inside run under a lock | Refunds above the invoice, other customers' invoices, double refunds |
| Human approval for privileged tools | ctx.approver | Any money movement no person has seen |
| Timeout | future.result(timeout=...) or a subprocess | A hung dependency freezing the conversation |
| Audit log | ctx.audit | Invisible actions; you can always answer "who refunded this?" |
The tests pin these behaviors so a later refactor cannot quietly remove one.
python -m pytest -q tests/test_m06_structured.py tests/test_m06_tools.py
Code explained
- In simple words: run the module's 32 tests, covering the parser, repair loop, salvage, strict schema, and every gate in the executor.
- What happens:
tests/test_m06_structured.pychecks that the corpus has at least 20 cases, that each parser stage never parses fewer cases than the one before (5 at the start, 21 at the end), each repair label, the apostrophe and brace edge cases, the strict schema rules, and the loop's ceiling.tests/test_m06_tools.pychecks the happy path, the oversized refund being rejected before approval, the schema bound, denial, a three-way parallel refund moving money once, read-only scope, cross-customer invisibility, errors instead of exceptions, a crashing tool not leaking its message, the timeout, call-ordertool_call_ids, the forced"none"last turn, and that calls on the last turn are not executed. - Comes out:
................................ [100%]
32 passed in 4.68s
Runtime varies (about 4 to 6 s here), mostly the deliberate sleep in the timeout test.
Module Lab
The lab runs the whole module end to end. Part 1 triages all 48 dev tickets through generate_structured with a ceiling of 3, then salvage. Part 2 sends the tickets that mention an invoice through the tool loop, where every refund waits in an approval queue for Maya. Offline, the triage "model" is a keyword classifier (a real classical baseline) whose JSON is deliberately wrapped in the hostile formats from Part A, one style per ticket, and the tool agent is scripted. That tests the pipeline, not model quality. With --live, llm.chat does both jobs.
-live, llm.chat does both jobs.
examples/m06_lab.py
"""Module 6 lab: structured triage for every dev ticket, then safe refund handling with tools.
Offline (default): the "model" is a keyword baseline (a real classical classifier)
whose JSON is deliberately wrapped in hostile formats, plus a scripted tool agent.
That tests the pipeline, not model quality. With --live, llm.chat does both jobs.
PYTHONPATH=. python examples/m06_lab.py
LLM_PROVIDER=groq GROQ_API_KEY=... PYTHONPATH=. python examples/m06_lab.py --live
"""
from __future__ import annotations
import json
import math
import re
import sys
from collections import Counter
from typing import Any
from examples.m06_structured import generate_structured, response_format_for, salvage
from examples.m06_tools import SYSTEM as AGENT_SYSTEM
from examples.m06_tools import WITH_REFUNDS, ToolContext, call, run_tool_loop
from supportdesk.data import Ticket, load_tickets
from supportdesk.llm import ChatResult, Usage
from supportdesk.schemas import Triage
from supportdesk.stand_in import ScriptedLLM
TRIAGE_SYSTEM = ("Triage this Brightlane support ticket. Reply with only a JSON object with keys category, "
"priority, language, summary, needs_human.")
LIVE = "--live" in sys.argv
# --- An offline triage "model": keyword rules, then hostile formatting -------------------------
RULES = [("feature_request", r"feature|would be great|please add|wish|could you add"),
("cancellation", r"cancel|refund for annual|stop now|money back|k\u00fcndig"),
("account_access", r"log ?in|password|sso|2fa|two-factor|locked|access|sign in|anmeld"),
("bug", r"error|crash|broken|not working|fails|down|bug|500"),
("billing", r"charge|invoice|bill|price|cost|vat|payment|refund|cobr|factura|plan")]
def keyword_triage(t: Ticket) -> dict[str, Any]:
text = t.text.lower()
category = next((c for c, pattern in RULES if re.search(pattern, text)), "how_to")
if re.search(r"all users|everyone|whole team|outage|security", text):
priority = "urgent"
elif re.search(r"charged|twice|locked|cannot|can't|blocked|refund", text):
priority = "high"
elif category == "feature_request":
priority = "low"
else:
priority = "normal"
if re.search(r"[\u3040-\u30ff\u4e00-\u9fff]", text):
language = "ja"
elif re.search(r"[\u0900-\u097f]", text):
language = "hi"
elif re.search(r"\b(hola|necesito|factura|cobraron)\b", text):
language = "es"
elif re.search(r"\b(ich|nicht|bitte|und)\b", text):
language = "de"
else:
language = "en"
return {"category": category, "priority": priority, "language": language,
"summary": ("Customer writes: " + t.subject)[:200],
"needs_human": category in ("billing", "cancellation", "account_access") and priority == "high"}
def hostile(data: dict[str, Any], style: int) -> str:
"""Format the baseline's answer the way real models misbehave (one style per ticket)."""
clean = json.dumps(data)
if style == 1:
return f"```json\n{clean}\n```"
if style == 2:
return f"Sure! Here is the triage:\n{clean}"
if style == 3:
return clean[:-1] + ",}"
if style == 4:
return str(data) # a Python dict: single quotes, True/False
if style == 6:
return json.dumps({**data, "priority": "critical" if data["priority"] == "urgent" else "medium"})
if style == 7:
return clean[: int(len(clean) * 0.6)] # cut off by max_tokens
return clean
def offline_triage_llm(tickets: dict[str, Ticket]) -> ScriptedLLM:
def respond(messages: list[dict], kwargs: dict) -> str:
ticket = tickets[messages[1]["content"]]
index = int(ticket.id.split("-")[1])
attempt = sum(m["role"] == "assistant" for m in messages)
data = keyword_triage(ticket)
if attempt == 0:
return hostile(data, index % 8)
if index % 16 == 7: # a stubborn case: stays broken, hits the ceiling
return hostile(data, 7)
return json.dumps(data) # after feedback, the stand-in sends clean JSON
return ScriptedLLM(responder=respond)
def wilson(hits: int, n: int, z: float = 1.96) -> tuple[float, float]:
p = hits / n
centre = (p + z * z / (2 * n)) / (1 + z * z / n)
half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
return centre - half, centre + half
def main() -> None:
# --- Part 1: triage every dev ticket ------------------------------------------------------------
tickets = load_tickets("dev")
if LIVE:
from supportdesk.llm import chat
triage_llm: Any = chat
extra = {"response_format": response_format_for(Triage), "max_tokens": 2000}
else:
triage_llm = offline_triage_llm({t.text: t for t in tickets})
extra = {}
status = Counter()
correct = 0
usage = Usage()
results: dict[str, dict[str, Any]] = {}
for t in tickets:
messages = [{"role": "system", "content": TRIAGE_SYSTEM}, {"role": "user", "content": t.text}]
out = generate_structured(triage_llm, messages, Triage, max_attempts=3, **extra)
usage.input_tokens += out.usage.input_tokens
usage.output_tokens += out.usage.output_tokens
if out.value is not None:
fields = out.value.model_dump()
status["valid, first try" if out.attempts == 1 else "valid after repair"] += 1
else:
fields, problems = salvage(out.last_data or {}, Triage)
if "category" in fields:
fields = {**fields, "priority": fields.get("priority", "high"), "needs_human": True}
status["salvaged"] += 1
else:
status["failed, sent to a human"] += 1
fields = {"category": None, "needs_human": True}
results[t.id] = fields
correct += fields.get("category") == t.gold["category"]
n = len(tickets)
print(f"Part 1: triage of {n} dev tickets ({'llm.chat' if LIVE else 'keyword baseline + injected hostile output'})")
for key in ("valid, first try", "valid after repair", "salvaged", "failed, sent to a human"):
print(f" {key:<26} {status[key]:>3}")
low, high = wilson(correct, n)
print(f" category accuracy vs gold: {correct}/{n} = {correct / n:.0%} (95% Wilson CI {low:.0%} to {high:.0%})")
print(f" tokens: {usage.input_tokens} in, {usage.output_tokens} out (estimates for the stand-in)")
# --- Part 2: tickets that mention an invoice go to the tool agent ---------------------------------
CUSTOMER_OF = {"T-1001": "cus_acme", "T-1031": "cus_sol"} # in production this comes from the ticket's account
pending: list[str] = []
def queue_for_maya(name: str, args: Any) -> bool:
pending.append(f"{name}({args.model_dump()})")
return False # nothing moves money until a person approves
def offline_agent(msgs: list[dict], kwargs: dict) -> ChatResult:
"""A scripted agent: look up, try to refund, then answer. Plumbing, not a model."""
tool_msgs = [m for m in msgs if m["role"] == "tool"]
ticket_text = msgs[1]["content"]
invoice_id = re.search(r"INV-\d{4}-\d{6}", ticket_text).group(0)
if not tool_msgs:
return ChatResult("", tool_calls=[call("a1", "search_kb", query="duplicate charge refund"),
call("a2", "get_invoice", invoice_id=invoice_id)])
if len(tool_msgs) == 2:
invoice = json.loads(tool_msgs[1]["content"])
return ChatResult("", tool_calls=[call("a3", "issue_refund", invoice_id=invoice_id,
amount_usd=invoice["refundable_usd"], reason="duplicate_charge")])
return ChatResult("Thanks for flagging the duplicate charge. A teammate is reviewing the refund now and will "
"confirm by email. [billing-refunds]")
print("\nPart 2: invoice tickets through the tool loop (refunds wait for human approval)")
for ticket_id, customer in CUSTOMER_OF.items():
t = next(x for x in tickets if x.id == ticket_id)
if results[ticket_id].get("category") not in ("billing", "cancellation"):
continue
ctx = ToolContext(customer, WITH_REFUNDS, approver=queue_for_maya)
agent = chat if LIVE else ScriptedLLM(responder=offline_agent)
messages = [{"role": "system", "content": AGENT_SYSTEM}, {"role": "user", "content": t.text}]
loop = run_tool_loop(agent, messages, ctx, max_turns=5)
print(f" {ticket_id} ({t.language}) {loop.stop_reason} in {loop.turns} calls; tools:",
[(o.name, "ok" if o.ok else json.loads(o.content)["error_type"]) for o in loop.outcomes])
print(f" draft for review: {loop.text}")
print(" approval queue for Maya:")
for item in pending:
print(" ", item)
if __name__ == "__main__":
main()
Code explained
- In simple words: a conveyor belt: messy answers in, typed triage out; invoice tickets then go through the guarded tool loop, and refunds stop at a human.
- What happens:
RULESandkeyword_triageare the offline classifier: the first matching regex picks the category, simple rules pick priority and language, andneeds_humanis true for high-priority billing, cancellation, and access tickets.hostile(data, style)formats that answer in one of eight ways: clean, fenced, with a preamble, with a trailing comma, as a Python dict, clean again, with an invalid priority, or truncated.offline_triage_llmwraps this in aScriptedLLMresponder that picks the style from the ticket number. After a correction message it answers with clean JSON, except for every sixteenth ticket, which stays truncated to exercise the ceiling and salvage.main()Part 1 runsgenerate_structuredfor each ticket and sorts results into valid first try, valid after repair, salvaged (category kept, priority defaulted, human review forced), or failed. It scores category accuracy against the gold labels with a Wilson interval and totals estimated tokens.main()Part 2 maps the two invoice tickets to their customers, builds aToolContextwithWITH_REFUNDSandqueue_for_mayaas the approver (it records the request and returnsFalse), and runsrun_tool_loop.offline_agentscripts a reasonable agent: look up the article and invoice, request a refund of the refundable amount, then write a reply.
- Comes out: nothing on import (
main()runs only as a script, som06_schema_eval.pycan importkeyword_triageandwilson). Run it with the next command.
python examples/m06_lab.py
Code explained
- In simple words: run the lab offline.
- What happens: the script prints the triage summary, then the tool-loop results and the approval queue.
- Comes out:
Part 1: triage of 48 dev tickets (keyword baseline + injected hostile output)
valid, first try 36
valid after repair 10
salvaged 2
failed, sent to a human 0
category accuracy vs gold: 24/48 = 50% (95% Wilson CI 36% to 64%)
tokens: 5092 in, 2371 out (estimates for the stand-in)
Part 2: invoice tickets through the tool loop (refunds wait for human approval)
T-1001 (en) answered in 3 calls; tools: [('search_kb', 'ok'), ('get_invoice', 'ok'), ('issue_refund', 'not_approved')]
draft for review: Thanks for flagging the duplicate charge. A teammate is reviewing the refund now and will confirm by email. [billing-refunds]
T-1031 (es) answered in 3 calls; tools: [('search_kb', 'ok'), ('get_invoice', 'ok'), ('issue_refund', 'not_approved')]
draft for review: Thanks for flagging the duplicate charge. A teammate is reviewing the refund now and will confirm by email. [billing-refunds]
approval queue for Maya:
issue_refund({'invoice_id': 'INV-2026-004512', 'amount_usd': 288.0, 'reason': 'duplicate_charge'})
issue_refund({'invoice_id': 'INV-2026-004871', 'amount_usd': 144.0, 'reason': 'duplicate_charge'})
Pipeline results: 36 tickets were valid on the first try (the styles the parser fixes), 10 needed one repair round (invalid priority, or truncation), 2 hit the ceiling and were salvaged into a human-reviewed route, and none crashed or was lost. The 50% category accuracy (95% CI 36% to 64%, n=48) belongs to the keyword baseline, not to any model, and is the bar a real model must clear on this data; the parser and repair loop cannot improve it, they only make sure it arrives intact. In Part 2 both refunds reached the approval queue with correct, validated arguments, nothing was refunded, and both drafts tell the customer a person is reviewing it. The scripted agent answers the Spanish ticket in English; a real model given Module 4's instructions should reply in the customer's language, and that is worth checking in your --live run.
To run the lab against a real model:
LLM_PROVIDER=groq GROQ_API_KEY=your-key python examples/m06_lab.py --live
Code explained
- In simple words: the same lab with
llm.chatdoing triage (with a strictresponse_format) and the tool loop. - What happens: Part 1 sends 48 triage requests plus any repairs; Part 2 runs the tool loop for the two invoice tickets, and refunds still go to the queue, never executed. Cost at the listed gpt-oss-120b prices is a fraction of a cent per ticket; check
pricing.pyfor your model. - Comes out: your own valid-first-try count, repair count, and category accuracy with its interval. Compare the accuracy to the keyword baseline's 50%, and remember that differences smaller than the interval width are noise at n=48.
Project Milestone
After this module, the Brightlane repository contains:
supportdesk/schemas.py(canonical, introduced here):Triage,DraftReply,Category,Priority,triage_json_schema().examples/m06_structured.py:strict_schema,response_format_for, the stagedparse_json,validate_output,describe_errors,generate_structuredwith a ceiling, andsalvage.examples/m06_hostile.py: the 24-case hostile corpus and the parser scoreboard (21% to 88% parsed).examples/m06_tools.py: the fake billing database,search_kb,get_invoice,issue_refund,ToolContextscopes, the never-raisingexecute, andrun_tool_loop.examples/m06_wire.py: offline capture of real request bodies fromllm.chat.- Demonstrations:
m06_freetext.py,m06_schema.py,m06_native.py,m06_native_wire.py,m06_repair.py,m06_salvage.py,m06_schema_cost.py,m06_schema_eval.py,m06_tool_request.py,m06_tool_errors.py,m06_parallel.py,m06_tool_choice.py,m06_many_tools.py,m06_untrusted.py,m06_sandbox.py, and the labm06_lab.py. tests/test_m06_structured.pyandtests/test_m06_tools.py: 32 passing tests.
The assistant can now produce triage objects that downstream code can trust, and it can look things up and propose refunds, with every privileged action checked by code and approved by a person.
Interview Questions
1. A teammate says "we use JSON mode, so we don't need validation." What do you tell them? JSON mode guarantees syntactically valid JSON, not your schema, and even strict schema modes guarantee shape, not truth. Some providers may not enforce keywords like pattern or maxLength (Gemini's documented keyword list omits them), best-effort modes can still return invalid output, and outputs can be truncated by max_tokens. Gemini's own docs say to always validate values in your application. Validate every response with the same typed model you generated the schema from, and treat a validation failure as a normal event with a repair or fallback path.
2. How would you design a parser for messy model output, and how do you know it works? In stages, each handling one failure class: strip code fences, extract the first balanced object while skipping braces inside strings, normalize near-JSON (single quotes, trailing commas, Python literals, smart quotes), and close truncated output. Record which repairs fired. Then measure on a corpus of hostile outputs with the stages turned on one at a time; in this module the parse rate went from 5/24 to 21/24. Also decide what not to fix: cases needing guesswork, like an apostrophe inside a single-quoted string, go to the repair loop instead.
3. What goes into a good repair loop? The specific validation errors in short, model-readable lines; the model's own previous answer so it can edit rather than restart; an instruction to return only the corrected object; and a hard attempt ceiling, usually 2 or 3. Track tokens, since each retry carries the whole failed exchange (47, 149, 251 estimated input tokens in the demo). Return a result object instead of raising, and keep the last parsed data for salvage.
4. When is it acceptable to use a partially valid object? When the valid fields are enough to take a safe action and the missing ones can be defaulted in the conservative direction. For triage, a valid category is enough to route; an invalid priority becomes "high" and the ticket is forced to human review. If the field that drives the action is invalid (no category), do not guess: escalate. Never treat a truncated-and-closed object as fully trusted.
5. Explain what the model actually sees when you pass tools. The provider renders the tool definitions into the prompt text using the model's chat template; for gpt-oss that is the harmony format, where tools appear as TypeScript-like type declarations with descriptions as comments. The model generates a call in a special format, and the provider parses it into tool_calls with an id, a name, and a JSON string of arguments. Constraints like pattern may not appear in the rendered text at all, so descriptions carry what the model must know, and code enforces the rest.
6. A tool call fails. Should the tool raise an exception? No. Catch it in the executor and return a short JSON error as the tool result, with a type code for your metrics: unknown tool, bad arguments, rejected by policy, not approved, timeout, internal. The model can then fix the argument, try something else, or explain the situation to the customer. Do not include stack traces or internal messages, which leak details and can be exploited.
7. How do you handle parallel tool calls correctly? Run independent calls concurrently, but send each result as a tool message tagged with the tool_call_id it answers, and answer every call before the next model request. Returning them in the original call order keeps logs deterministic. Tools with side effects need their own protection against concurrent calls: in the demo, two parallel refunds of the same invoice moved money only once because the check and the update happen under one lock.
8. When would you use each tool_choice setting? auto for normal turns. A forced function on the first turn when a lookup must happen, or when you use a tool's arguments as structured output. required when any tool must be used. none on the last allowed turn to guarantee an answer and bound the loop. And verify: some servers, Ollama's OpenAI-compatible endpoint per its docs, do not support tool_choice, so check the reply and keep a turn limit.
9. Your agent has 40 tools and picks the wrong one often. What do you do? Measure first: tool-selection accuracy on a labeled set of real requests, plus the token cost of definitions (30 tools cost 2,291 tokens per call in this module versus 533 for 3). Then reduce what each call sees: group tools by the triaged task, retrieve the top few by description with a fallback set, merge overlapping tools into one with an enum, and sharpen descriptions. Published work such as RAG-MCP reports large selection gains from retrieving tools instead of listing all of them, but re-measure on your own model and requests.
10. How do you stop a model from issuing a refund it should not? Layers, all in code: the refund tool is only offered in the refund workflow's context; the customer id comes from the session, not from arguments; arguments are validated (invoice pattern, amount bounds, reason enum); a policy check compares the amount with the invoice's refundable amount before anyone is asked; a human approves every refund; the execution repeats the check under a lock; and everything is audited. In the demo a 576 USD refund on a 288 USD invoice was rejected before the approver was called.
11. What is the difference between a timeout and a sandbox? A timeout bounds how long you wait. A sandbox bounds what the code can do while it runs: CPU, memory, files, network, secrets. A thread timeout returns control but leaves the work running; a subprocess with RLIMIT_CPU, RLIMIT_AS, an empty environment, a temporary directory, and a wall-clock timeout can actually be stopped. For untrusted code you want both, plus network and filesystem isolation from a container or a stronger layer such as gVisor.
12. Does making the schema bigger or more nested hurt accuracy? It certainly costs tokens (in this module, about 26 per described field and about 33 per nesting level), and providers may reject very large or deep schemas. Research such as "Let Me Speak Freely?" found format restrictions can hurt reasoning, and JSONSchemaBench found constrained-decoding support varies with schema complexity. But the effect on a given task is an empirical question: run the same tickets with a flat and a nested schema, compare validity and accuracy with confidence intervals, and keep schemas flat until the numbers say otherwise.
Other Tools and Providers
| Tool or provider | What it is | When to consider it instead of this module's approach |
|---|---|---|
| OpenAI API Structured Outputs | Native json_schema with strict mode, also for function parameters | If you use OpenAI models directly; same response_format shape as here |
| Anthropic tool use | Tools defined with input_schema; responses contain tool_use blocks, results go back as tool_result blocks | If you use Claude models; the loop and guard layer carry over unchanged, only the message format differs |
Gemini native API (google-genai) | response_schema for structured output, function-calling modes AUTO, ANY, NONE, VALIDATED | When you need Gemini features the OpenAI-compatible layer (beta) does not expose |
| Instructor | Library that wraps clients with pydantic response models, validation, and retries | If you want the repair loop as a library; you give up some visibility into each attempt |
| Outlines, XGrammar, llama.cpp grammars | Constrained decoding for models you run yourself | Self-hosted models where you control decoding; JSONSchemaBench compares several |
| vLLM and SGLang guided decoding | Structured output in open-source inference servers | Serving open-weight models at scale (Module 13) |
| json-repair (PyPI) | A general-purpose repairer for malformed JSON | If you would rather not maintain a parser; measure it on your own hostile corpus first, as we did |
| jsonschema (Python) | A JSON Schema validator | When the schema, not a pydantic model, is your source of truth |
| Model Context Protocol (MCP) | A standard for exposing tools from separate servers | When many apps share tools; the guard layer still belongs on your side |
| Containers, gVisor, Firecracker | Process and VM isolation | Tools that run code or touch files, beyond the subprocess limits shown here |
Coming Up in Module 7
Module 7, Context Engineering, treats everything you put in the window as a scarce budget. You just measured one of the biggest line items: tool definitions cost more than the ticket itself, and 30 tools cost over 2,000 tokens per call. Module 7 formally introduces kb_search.py, the BM25 search behind search_kb, and asks how much retrieved knowledge, history, and tool text a request should carry, in what order, and how to tell whether extra context actually improved the answer.