CourseLarge Language Models · Module 6: Structured Outputs and Tool Use · part 29 of 80
Part 29 · Module 6: Structured Outputs and Tool Use

Part C: Model output as untrusted input

24 min read·22 Sept 2026

Everything a model writes, including tool arguments, is untrusted input, the same as a form a stranger filled in. The model might be wrong, confused by a long context, or steered by a prompt injection hidden in a ticket (Module 11 covers injection in depth). The rule of this part: nothing the model writes reaches a privileged operation without code checking it first. A privileged operation is one that changes money, access, or data, like issue_refund.

.

Validating arguments before execution

python
"""Model output is untrusted input: what stops a bad refund before money moves."""
from __future__ import annotations

import dataclasses
import time

from examples.m06_tools import SYSTEM, TICKET, TOOLS, WITH_REFUNDS, ToolContext, call, execute, run_tool_loop
from supportdesk.llm import ChatResult
from supportdesk.stand_in import ScriptedLLM

asked: list[str] = []


def approver(name, args) -> bool:
    asked.append(f"{name} {args.amount_usd} USD")
    return True


print("=== 1. Refund larger than the invoice (576 USD on a 288 USD invoice)")
ctx = ToolContext("cus_acme", WITH_REFUNDS, approver=approver)
outcome = execute(call("r1", "issue_refund", invoice_id="INV-2026-004513", amount_usd=576.0, reason="duplicate_charge"), ctx)
print(outcome.content)
print("approver was asked:", asked, "| refunded_usd:", ctx.db["INV-2026-004513"].refunded_usd)

print("\n=== 2. Amount outside the schema bound (20000 USD)")
print(execute(call("r2", "issue_refund", invoice_id="INV-2026-004513", amount_usd=20000, reason="duplicate_charge"), ctx).content)

print("\n=== 3. The error goes back to the model, which corrects itself (ScriptedLLM, plumbing only)")
llm = ScriptedLLM(replies=[
    ChatResult("", tool_calls=[call("r3", "issue_refund", invoice_id="INV-2026-004513", amount_usd=576.0, reason="duplicate_charge")]),
    ChatResult("", tool_calls=[call("r4", "issue_refund", invoice_id="INV-2026-004513", amount_usd=288.0, reason="duplicate_charge")]),
    ChatResult("We refunded the duplicate charge of 288 USD. It arrives in 5 to 10 business days."),
])
messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": TICKET}]
result = run_tool_loop(llm, messages, ctx, verbose=True)
print("final:", result.text, "| approver asked:", asked)

print("\n=== 4. A human says no")
deny_ctx = ToolContext("cus_acme", WITH_REFUNDS, approver=lambda name, args: False)
print(execute(call("r5", "issue_refund", invoice_id="INV-2026-004513", amount_usd=288.0, reason="duplicate_charge"), deny_ctx).content)
print("refunded_usd:", deny_ctx.db["INV-2026-004513"].refunded_usd)

print("\n=== 5. Two parallel refunds of the same duplicate")
par_ctx = ToolContext("cus_acme", WITH_REFUNDS, approver=lambda name, args: True)
llm = ScriptedLLM(replies=[
    ChatResult("", tool_calls=[call(f"p{i}", "issue_refund", invoice_id="INV-2026-004513", amount_usd=288.0,
                                    reason="duplicate_charge") for i in (1, 2)]),
    ChatResult("Refunded once."),
])
run_tool_loop(llm, messages, par_ctx, verbose=True)
print("refunded_usd:", par_ctx.db["INV-2026-004513"].refunded_usd)

print("\n=== 6. A tool that hangs is abandoned after its timeout")


def slow_invoice(args, ctx):
    time.sleep(3)                                   # imagine the billing API is down
    return {"never": "returned"}


slow_registry = {**TOOLS, "get_invoice": dataclasses.replace(TOOLS["get_invoice"], run=slow_invoice, timeout_s=0.5)}
slow = execute(call("t1", "get_invoice", invoice_id="INV-2026-004512"), par_ctx, registry=slow_registry)
print(f"{slow.content} after {slow.ms:.0f} ms")

Code explained

  • In simple words: try to move too much money in several ways, and watch each attempt stop at a different gate.
  • What happens: case 1 asks for 576 USD on a 288 USD invoice, the classic "refund both charges" mistake. Case 2 exceeds the schema bound. Case 3 runs the loop with a ScriptedLLM that makes the case 1 mistake, reads the error, and corrects itself (scripted plumbing, not model behavior). Case 4 uses an approver that says no. Case 5 sends two identical refunds in one parallel turn. Case 6 replaces get_invoice with a function that hangs for 3 s under a 0.5 s timeout.
  • Comes out:
text
=== 1. Refund larger than the invoice (576 USD on a 288 USD invoice)
{"error": "Refund of 576.00 USD exceeds the refundable amount of 288.00 USD on INV-2026-004513.", "error_type": "rejected"}
approver was asked: [] | refunded_usd: 0.0

=== 2. Amount outside the schema bound (20000 USD)
{"error": "Invalid arguments: amount_usd: Input should be less than or equal to 10000", "error_type": "bad_arguments"}

=== 3. The error goes back to the model, which corrects itself (ScriptedLLM, plumbing only)
  turn 1 ERR issue_refund({"invoice_id": "INV-2026-004513", "amount_usd": 576.0, "reason": "duplicate_charge"})
           -> {"error": "Refund of 576.00 USD exceeds the refundable amount of 288.00 USD on INV-2026-004513.", "error_type"
  turn 2 ok  issue_refund({"invoice_id": "INV-2026-004513", "amount_usd": 288.0, "reason": "duplicate_charge"})
           -> {"status": "refunded", "invoice_id": "INV-2026-004513", "amount_usd": 288.0, "refundable_usd_now": 0.0, "arriv
final: We refunded the duplicate charge of 288 USD. It arrives in 5 to 10 business days. | approver asked: ['issue_refund 288.0 USD']

=== 4. A human says no
{"error": "This action needs human approval and was not approved now. Do not retry it; tell the customer a teammate will review the request.", "error_type": "not_approved"}
refunded_usd: 0.0

=== 5. Two parallel refunds of the same duplicate
  turn 1 ok  issue_refund({"invoice_id": "INV-2026-004513", "amount_usd": 288.0, "reason": "duplicate_charge"})
           -> {"status": "refunded", "invoice_id": "INV-2026-004513", "amount_usd": 288.0, "refundable_usd_now": 0.0, "arriv
  turn 1 ERR issue_refund({"invoice_id": "INV-2026-004513", "amount_usd": 288.0, "reason": "duplicate_charge"})
           -> {"error": "Refund of 288.00 USD exceeds the refundable amount of 0.00 USD on INV-2026-004513.", "error_type":
refunded_usd: 288.0

=== 6. A tool that hangs is abandoned after its timeout
{"error": "get_invoice took longer than 0.5s and was abandoned. Try once more or continue without it.", "error_type": "timeout"} after 514 ms

The key line is under case 1: approver was asked: []. The oversized refund was rejected by the policy check before a human was even asked, so Maya never sees a request that should not exist, and refunded_usd stayed 0.0. In case 3 the approver was asked exactly once, for the corrected 288 USD. In case 5 exactly one refund succeeded; which of the two wins can change between runs, but the total never exceeds the invoice, thanks to the lock around check and update. In case 6 the call returned a timeout error after about 0.5 s (514 ms here; it varies).

Sandboxing and permission scoping

Permission scoping means each conversation gets only the capabilities its job needs. The executor enforces three scopes, all set by code:

  • Which tools: a triage-only session gets READ_ONLY; only the refund workflow gets WITH_REFUNDS. You saw a read-only session's refund attempt return not_available.
  • Which data: get_invoice only sees the ticket's own customer, taken from ToolContext, never from arguments.
  • Which actions without a human: none that move money. issue_refund always needs approval.

A sandbox limits what running code can do even when it misbehaves. Tools that only call your own APIs, like these three, are sandboxed by the checks above. Tools that run code or touch files, such as a "run this spreadsheet formula" or "convert this attachment" tool, need process isolation, because a thread cannot be stopped from outside.

python
"""Sandboxing basics: run untrusted work in a separate process with hard limits.

A thread cannot be killed, but a process can. This runs a snippet of Python in
a child process with a wall-clock timeout, a CPU-time limit, a memory limit, an
empty environment (no API keys), and a throwaway working directory. It is a
first layer, not a complete sandbox: see the text for containers and gVisor.
"""
from __future__ import annotations

import resource
import subprocess
import sys
import tempfile
import time


def _limits() -> None:                      # runs in the child just before the snippet starts
    resource.setrlimit(resource.RLIMIT_CPU, (1, 1))                     # 1 s of CPU
    resource.setrlimit(resource.RLIMIT_AS, (256 * 2**20, 256 * 2**20))  # 256 MB of memory


def run_sandboxed(code: str, timeout_s: float = 5.0) -> dict:
    started = time.perf_counter()
    with tempfile.TemporaryDirectory() as workdir:
        try:
            proc = subprocess.run([sys.executable, "-I", "-c", code], cwd=workdir, env={}, capture_output=True,
                                  text=True, timeout=timeout_s, preexec_fn=_limits)
            outcome = {"exit": proc.returncode, "stdout": proc.stdout.strip()[:80], "stderr": (proc.stderr.strip().splitlines() or [""])[-1][:80]}
        except subprocess.TimeoutExpired:
            outcome = {"exit": None, "error": f"killed after {timeout_s}s wall clock"}
    outcome["ms"] = round((time.perf_counter() - started) * 1000)
    return outcome


CASES = {
    "normal": "print(sum([288, 288]) / 2)",
    "reads env": "import os; print(sorted(os.environ))",
    "cpu loop": "while True: pass",
    "sleeps": "import time; time.sleep(60)",
    "memory bomb": "x = bytearray(1024 * 2**20); print(len(x))",
}
for name, code in CASES.items():
    print(f"{name:<12} {run_sandboxed(code)}")

Code explained

  • In simple words: run untrusted snippets in a separate Python process with a time limit, CPU and memory limits, no environment variables, and a throwaway folder.
  • What happens: run_sandboxed starts python -I -c code (isolated mode ignores user site packages and PYTHON* variables) with env={} so no API keys are visible, cwd set to a temporary directory that is deleted afterwards, and preexec_fn=_limits, which sets RLIMIT_CPU to 1 s of CPU time and RLIMIT_AS to 256 MB of address space in the child before it starts. subprocess.run(timeout=5.0) kills the child if the wall clock runs out. Five cases exercise each limit.
  • Comes out:
text
normal       {'exit': 0, 'stdout': '288.0', 'stderr': '', 'ms': 138}
reads env    {'exit': 0, 'stdout': "['LC_CTYPE']", 'stderr': '', 'ms': 104}
cpu loop     {'exit': None, 'error': 'killed after 5.0s wall clock', 'ms': 5024}
sleeps       {'exit': None, 'error': 'killed after 5.0s wall clock', 'ms': 5044}
memory bomb  {'exit': 1, 'stdout': '', 'stderr': 'MemoryError', 'ms': 86}

The normal snippet works; the environment is empty apart from a locale variable Python sets itself; the memory bomb fails with MemoryError inside the child; the sleeper is killed at the 5 s wall-clock limit, which a CPU limit alone would never catch. The CPU loop is the honest surprise. RLIMIT_CPU counts CPU time, not wall time, and on this build machine (a VM) the child was killed by the wall-clock timeout in this run and by the CPU limit (exit code -9, after about 3.6 to 4.3 s of wall time) in two earlier runs. That is why you set both limits. This is a first layer, not a complete sandbox: it does not block network access or file reads outside the folder. For those, run tools in a container with no network and a read-only filesystem, or use a stronger isolation layer such as gVisor or a microVM.

Timeouts belong on every tool, not only risky ones. The executor's thread timeout in case 6 returns control to the loop after 0.5 s, but the stuck thread keeps running until its function returns (that demo script takes about 3 s to exit for this reason). For work that can hang forever, use a process, as above, or an HTTP client timeout inside the tool.

Never letting a model's output reach a privileged operation unchecked

Put the whole guard layer together and you get a short checklist. Every item is in execute and tested:

CheckWhereWhat it stops
Tool offered in this contextctx.allowed_toolsInvented tools, tools outside this workflow's scope
Arguments parse and validatepydantic models with enums, patterns, boundsWrong formats, invented arguments, absurd amounts
Policy check against live dataprecheck, then again inside run under a lockRefunds above the invoice, other customers' invoices, double refunds
Human approval for privileged toolsctx.approverAny money movement no person has seen
Timeoutfuture.result(timeout=...) or a subprocessA hung dependency freezing the conversation
Audit logctx.auditInvisible actions; you can always answer "who refunded this?"

The tests pin these behaviors so a later refactor cannot quietly remove one.

bash
python -m pytest -q tests/test_m06_structured.py tests/test_m06_tools.py

Code explained

  • In simple words: run the module's 32 tests, covering the parser, repair loop, salvage, strict schema, and every gate in the executor.
  • What happens: tests/test_m06_structured.py checks that the corpus has at least 20 cases, that each parser stage never parses fewer cases than the one before (5 at the start, 21 at the end), each repair label, the apostrophe and brace edge cases, the strict schema rules, and the loop's ceiling. tests/test_m06_tools.py checks the happy path, the oversized refund being rejected before approval, the schema bound, denial, a three-way parallel refund moving money once, read-only scope, cross-customer invisibility, errors instead of exceptions, a crashing tool not leaking its message, the timeout, call-order tool_call_ids, the forced "none" last turn, and that calls on the last turn are not executed.
  • Comes out:
text
................................                                         [100%]
32 passed in 4.68s

Runtime varies (about 4 to 6 s here), mostly the deliberate sleep in the timeout test.

Module Lab

The lab runs the whole module end to end. Part 1 triages all 48 dev tickets through generate_structured with a ceiling of 3, then salvage. Part 2 sends the tickets that mention an invoice through the tool loop, where every refund waits in an approval queue for Maya. Offline, the triage "model" is a keyword classifier (a real classical baseline) whose JSON is deliberately wrapped in the hostile formats from Part A, one style per ticket, and the tool agent is scripted. That tests the pipeline, not model quality. With --live, llm.chat does both jobs.

-live, llm.chat does both jobs.

examples/m06_lab.py

python
"""Module 6 lab: structured triage for every dev ticket, then safe refund handling with tools.

Offline (default): the "model" is a keyword baseline (a real classical classifier)
whose JSON is deliberately wrapped in hostile formats, plus a scripted tool agent.
That tests the pipeline, not model quality. With --live, llm.chat does both jobs.

  PYTHONPATH=. python examples/m06_lab.py
  LLM_PROVIDER=groq GROQ_API_KEY=... PYTHONPATH=. python examples/m06_lab.py --live
"""
from __future__ import annotations

import json
import math
import re
import sys
from collections import Counter
from typing import Any

from examples.m06_structured import generate_structured, response_format_for, salvage
from examples.m06_tools import SYSTEM as AGENT_SYSTEM
from examples.m06_tools import WITH_REFUNDS, ToolContext, call, run_tool_loop
from supportdesk.data import Ticket, load_tickets
from supportdesk.llm import ChatResult, Usage
from supportdesk.schemas import Triage
from supportdesk.stand_in import ScriptedLLM

TRIAGE_SYSTEM = ("Triage this Brightlane support ticket. Reply with only a JSON object with keys category, "
                 "priority, language, summary, needs_human.")
LIVE = "--live" in sys.argv


# --- An offline triage "model": keyword rules, then hostile formatting -------------------------

RULES = [("feature_request", r"feature|would be great|please add|wish|could you add"),
         ("cancellation", r"cancel|refund for annual|stop now|money back|k\u00fcndig"),
         ("account_access", r"log ?in|password|sso|2fa|two-factor|locked|access|sign in|anmeld"),
         ("bug", r"error|crash|broken|not working|fails|down|bug|500"),
         ("billing", r"charge|invoice|bill|price|cost|vat|payment|refund|cobr|factura|plan")]


def keyword_triage(t: Ticket) -> dict[str, Any]:
    text = t.text.lower()
    category = next((c for c, pattern in RULES if re.search(pattern, text)), "how_to")
    if re.search(r"all users|everyone|whole team|outage|security", text):
        priority = "urgent"
    elif re.search(r"charged|twice|locked|cannot|can't|blocked|refund", text):
        priority = "high"
    elif category == "feature_request":
        priority = "low"
    else:
        priority = "normal"
    if re.search(r"[\u3040-\u30ff\u4e00-\u9fff]", text):
        language = "ja"
    elif re.search(r"[\u0900-\u097f]", text):
        language = "hi"
    elif re.search(r"\b(hola|necesito|factura|cobraron)\b", text):
        language = "es"
    elif re.search(r"\b(ich|nicht|bitte|und)\b", text):
        language = "de"
    else:
        language = "en"
    return {"category": category, "priority": priority, "language": language,
            "summary": ("Customer writes: " + t.subject)[:200],
            "needs_human": category in ("billing", "cancellation", "account_access") and priority == "high"}


def hostile(data: dict[str, Any], style: int) -> str:
    """Format the baseline's answer the way real models misbehave (one style per ticket)."""
    clean = json.dumps(data)
    if style == 1:
        return f"```json\n{clean}\n```"
    if style == 2:
        return f"Sure! Here is the triage:\n{clean}"
    if style == 3:
        return clean[:-1] + ",}"
    if style == 4:
        return str(data)                                   # a Python dict: single quotes, True/False
    if style == 6:
        return json.dumps({**data, "priority": "critical" if data["priority"] == "urgent" else "medium"})
    if style == 7:
        return clean[: int(len(clean) * 0.6)]              # cut off by max_tokens
    return clean


def offline_triage_llm(tickets: dict[str, Ticket]) -> ScriptedLLM:
    def respond(messages: list[dict], kwargs: dict) -> str:
        ticket = tickets[messages[1]["content"]]
        index = int(ticket.id.split("-")[1])
        attempt = sum(m["role"] == "assistant" for m in messages)
        data = keyword_triage(ticket)
        if attempt == 0:
            return hostile(data, index % 8)
        if index % 16 == 7:                                 # a stubborn case: stays broken, hits the ceiling
            return hostile(data, 7)
        return json.dumps(data)                             # after feedback, the stand-in sends clean JSON
    return ScriptedLLM(responder=respond)


def wilson(hits: int, n: int, z: float = 1.96) -> tuple[float, float]:
    p = hits / n
    centre = (p + z * z / (2 * n)) / (1 + z * z / n)
    half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
    return centre - half, centre + half


def main() -> None:
    # --- Part 1: triage every dev ticket ------------------------------------------------------------

    tickets = load_tickets("dev")
    if LIVE:
        from supportdesk.llm import chat
        triage_llm: Any = chat
        extra = {"response_format": response_format_for(Triage), "max_tokens": 2000}
    else:
        triage_llm = offline_triage_llm({t.text: t for t in tickets})
        extra = {}

    status = Counter()
    correct = 0
    usage = Usage()
    results: dict[str, dict[str, Any]] = {}
    for t in tickets:
        messages = [{"role": "system", "content": TRIAGE_SYSTEM}, {"role": "user", "content": t.text}]
        out = generate_structured(triage_llm, messages, Triage, max_attempts=3, **extra)
        usage.input_tokens += out.usage.input_tokens
        usage.output_tokens += out.usage.output_tokens
        if out.value is not None:
            fields = out.value.model_dump()
            status["valid, first try" if out.attempts == 1 else "valid after repair"] += 1
        else:
            fields, problems = salvage(out.last_data or {}, Triage)
            if "category" in fields:
                fields = {**fields, "priority": fields.get("priority", "high"), "needs_human": True}
                status["salvaged"] += 1
            else:
                status["failed, sent to a human"] += 1
                fields = {"category": None, "needs_human": True}
        results[t.id] = fields
        correct += fields.get("category") == t.gold["category"]

    n = len(tickets)
    print(f"Part 1: triage of {n} dev tickets ({'llm.chat' if LIVE else 'keyword baseline + injected hostile output'})")
    for key in ("valid, first try", "valid after repair", "salvaged", "failed, sent to a human"):
        print(f"  {key:<26} {status[key]:>3}")
    low, high = wilson(correct, n)
    print(f"  category accuracy vs gold: {correct}/{n} = {correct / n:.0%} (95% Wilson CI {low:.0%} to {high:.0%})")
    print(f"  tokens: {usage.input_tokens} in, {usage.output_tokens} out (estimates for the stand-in)")

    # --- Part 2: tickets that mention an invoice go to the tool agent ---------------------------------

    CUSTOMER_OF = {"T-1001": "cus_acme", "T-1031": "cus_sol"}   # in production this comes from the ticket's account
    pending: list[str] = []


    def queue_for_maya(name: str, args: Any) -> bool:
        pending.append(f"{name}({args.model_dump()})")
        return False                                            # nothing moves money until a person approves


    def offline_agent(msgs: list[dict], kwargs: dict) -> ChatResult:
        """A scripted agent: look up, try to refund, then answer. Plumbing, not a model."""
        tool_msgs = [m for m in msgs if m["role"] == "tool"]
        ticket_text = msgs[1]["content"]
        invoice_id = re.search(r"INV-\d{4}-\d{6}", ticket_text).group(0)
        if not tool_msgs:
            return ChatResult("", tool_calls=[call("a1", "search_kb", query="duplicate charge refund"),
                                              call("a2", "get_invoice", invoice_id=invoice_id)])
        if len(tool_msgs) == 2:
            invoice = json.loads(tool_msgs[1]["content"])
            return ChatResult("", tool_calls=[call("a3", "issue_refund", invoice_id=invoice_id,
                                                   amount_usd=invoice["refundable_usd"], reason="duplicate_charge")])
        return ChatResult("Thanks for flagging the duplicate charge. A teammate is reviewing the refund now and will "
                          "confirm by email. [billing-refunds]")


    print("\nPart 2: invoice tickets through the tool loop (refunds wait for human approval)")
    for ticket_id, customer in CUSTOMER_OF.items():
        t = next(x for x in tickets if x.id == ticket_id)
        if results[ticket_id].get("category") not in ("billing", "cancellation"):
            continue
        ctx = ToolContext(customer, WITH_REFUNDS, approver=queue_for_maya)
        agent = chat if LIVE else ScriptedLLM(responder=offline_agent)
        messages = [{"role": "system", "content": AGENT_SYSTEM}, {"role": "user", "content": t.text}]
        loop = run_tool_loop(agent, messages, ctx, max_turns=5)
        print(f"  {ticket_id} ({t.language}) {loop.stop_reason} in {loop.turns} calls; tools:",
              [(o.name, "ok" if o.ok else json.loads(o.content)["error_type"]) for o in loop.outcomes])
        print(f"    draft for review: {loop.text}")
    print("  approval queue for Maya:")
    for item in pending:
        print("   ", item)


if __name__ == "__main__":
    main()

Code explained

  • In simple words: a conveyor belt: messy answers in, typed triage out; invoice tickets then go through the guarded tool loop, and refunds stop at a human.
  • What happens:
    • RULES and keyword_triage are the offline classifier: the first matching regex picks the category, simple rules pick priority and language, and needs_human is true for high-priority billing, cancellation, and access tickets.
    • hostile(data, style) formats that answer in one of eight ways: clean, fenced, with a preamble, with a trailing comma, as a Python dict, clean again, with an invalid priority, or truncated.
    • offline_triage_llm wraps this in a ScriptedLLM responder that picks the style from the ticket number. After a correction message it answers with clean JSON, except for every sixteenth ticket, which stays truncated to exercise the ceiling and salvage.
    • main() Part 1 runs generate_structured for each ticket and sorts results into valid first try, valid after repair, salvaged (category kept, priority defaulted, human review forced), or failed. It scores category accuracy against the gold labels with a Wilson interval and totals estimated tokens.
    • main() Part 2 maps the two invoice tickets to their customers, builds a ToolContext with WITH_REFUNDS and queue_for_maya as the approver (it records the request and returns False), and runs run_tool_loop. offline_agent scripts a reasonable agent: look up the article and invoice, request a refund of the refundable amount, then write a reply.
  • Comes out: nothing on import (main() runs only as a script, so m06_schema_eval.py can import keyword_triage and wilson). Run it with the next command.
bash
python examples/m06_lab.py

Code explained

  • In simple words: run the lab offline.
  • What happens: the script prints the triage summary, then the tool-loop results and the approval queue.
  • Comes out:
text
Part 1: triage of 48 dev tickets (keyword baseline + injected hostile output)
  valid, first try            36
  valid after repair          10
  salvaged                     2
  failed, sent to a human      0
  category accuracy vs gold: 24/48 = 50% (95% Wilson CI 36% to 64%)
  tokens: 5092 in, 2371 out (estimates for the stand-in)

Part 2: invoice tickets through the tool loop (refunds wait for human approval)
  T-1001 (en) answered in 3 calls; tools: [('search_kb', 'ok'), ('get_invoice', 'ok'), ('issue_refund', 'not_approved')]
    draft for review: Thanks for flagging the duplicate charge. A teammate is reviewing the refund now and will confirm by email. [billing-refunds]
  T-1031 (es) answered in 3 calls; tools: [('search_kb', 'ok'), ('get_invoice', 'ok'), ('issue_refund', 'not_approved')]
    draft for review: Thanks for flagging the duplicate charge. A teammate is reviewing the refund now and will confirm by email. [billing-refunds]
  approval queue for Maya:
    issue_refund({'invoice_id': 'INV-2026-004512', 'amount_usd': 288.0, 'reason': 'duplicate_charge'})
    issue_refund({'invoice_id': 'INV-2026-004871', 'amount_usd': 144.0, 'reason': 'duplicate_charge'})

Pipeline results: 36 tickets were valid on the first try (the styles the parser fixes), 10 needed one repair round (invalid priority, or truncation), 2 hit the ceiling and were salvaged into a human-reviewed route, and none crashed or was lost. The 50% category accuracy (95% CI 36% to 64%, n=48) belongs to the keyword baseline, not to any model, and is the bar a real model must clear on this data; the parser and repair loop cannot improve it, they only make sure it arrives intact. In Part 2 both refunds reached the approval queue with correct, validated arguments, nothing was refunded, and both drafts tell the customer a person is reviewing it. The scripted agent answers the Spanish ticket in English; a real model given Module 4's instructions should reply in the customer's language, and that is worth checking in your --live run.

To run the lab against a real model:

bash
LLM_PROVIDER=groq GROQ_API_KEY=your-key python examples/m06_lab.py --live

Code explained

  • In simple words: the same lab with llm.chat doing triage (with a strict response_format) and the tool loop.
  • What happens: Part 1 sends 48 triage requests plus any repairs; Part 2 runs the tool loop for the two invoice tickets, and refunds still go to the queue, never executed. Cost at the listed gpt-oss-120b prices is a fraction of a cent per ticket; check pricing.py for your model.
  • Comes out: your own valid-first-try count, repair count, and category accuracy with its interval. Compare the accuracy to the keyword baseline's 50%, and remember that differences smaller than the interval width are noise at n=48.

Project Milestone

After this module, the Brightlane repository contains:

  • supportdesk/schemas.py (canonical, introduced here): Triage, DraftReply, Category, Priority, triage_json_schema().
  • examples/m06_structured.py: strict_schema, response_format_for, the staged parse_json, validate_output, describe_errors, generate_structured with a ceiling, and salvage.
  • examples/m06_hostile.py: the 24-case hostile corpus and the parser scoreboard (21% to 88% parsed).
  • examples/m06_tools.py: the fake billing database, search_kb, get_invoice, issue_refund, ToolContext scopes, the never-raising execute, and run_tool_loop.
  • examples/m06_wire.py: offline capture of real request bodies from llm.chat.
  • Demonstrations: m06_freetext.py, m06_schema.py, m06_native.py, m06_native_wire.py, m06_repair.py, m06_salvage.py, m06_schema_cost.py, m06_schema_eval.py, m06_tool_request.py, m06_tool_errors.py, m06_parallel.py, m06_tool_choice.py, m06_many_tools.py, m06_untrusted.py, m06_sandbox.py, and the lab m06_lab.py.
  • tests/test_m06_structured.py and tests/test_m06_tools.py: 32 passing tests.

The assistant can now produce triage objects that downstream code can trust, and it can look things up and propose refunds, with every privileged action checked by code and approved by a person.

Interview Questions

1. A teammate says "we use JSON mode, so we don't need validation." What do you tell them? JSON mode guarantees syntactically valid JSON, not your schema, and even strict schema modes guarantee shape, not truth. Some providers may not enforce keywords like pattern or maxLength (Gemini's documented keyword list omits them), best-effort modes can still return invalid output, and outputs can be truncated by max_tokens. Gemini's own docs say to always validate values in your application. Validate every response with the same typed model you generated the schema from, and treat a validation failure as a normal event with a repair or fallback path.

2. How would you design a parser for messy model output, and how do you know it works? In stages, each handling one failure class: strip code fences, extract the first balanced object while skipping braces inside strings, normalize near-JSON (single quotes, trailing commas, Python literals, smart quotes), and close truncated output. Record which repairs fired. Then measure on a corpus of hostile outputs with the stages turned on one at a time; in this module the parse rate went from 5/24 to 21/24. Also decide what not to fix: cases needing guesswork, like an apostrophe inside a single-quoted string, go to the repair loop instead.

3. What goes into a good repair loop? The specific validation errors in short, model-readable lines; the model's own previous answer so it can edit rather than restart; an instruction to return only the corrected object; and a hard attempt ceiling, usually 2 or 3. Track tokens, since each retry carries the whole failed exchange (47, 149, 251 estimated input tokens in the demo). Return a result object instead of raising, and keep the last parsed data for salvage.

4. When is it acceptable to use a partially valid object? When the valid fields are enough to take a safe action and the missing ones can be defaulted in the conservative direction. For triage, a valid category is enough to route; an invalid priority becomes "high" and the ticket is forced to human review. If the field that drives the action is invalid (no category), do not guess: escalate. Never treat a truncated-and-closed object as fully trusted.

5. Explain what the model actually sees when you pass tools. The provider renders the tool definitions into the prompt text using the model's chat template; for gpt-oss that is the harmony format, where tools appear as TypeScript-like type declarations with descriptions as comments. The model generates a call in a special format, and the provider parses it into tool_calls with an id, a name, and a JSON string of arguments. Constraints like pattern may not appear in the rendered text at all, so descriptions carry what the model must know, and code enforces the rest.

6. A tool call fails. Should the tool raise an exception? No. Catch it in the executor and return a short JSON error as the tool result, with a type code for your metrics: unknown tool, bad arguments, rejected by policy, not approved, timeout, internal. The model can then fix the argument, try something else, or explain the situation to the customer. Do not include stack traces or internal messages, which leak details and can be exploited.

7. How do you handle parallel tool calls correctly? Run independent calls concurrently, but send each result as a tool message tagged with the tool_call_id it answers, and answer every call before the next model request. Returning them in the original call order keeps logs deterministic. Tools with side effects need their own protection against concurrent calls: in the demo, two parallel refunds of the same invoice moved money only once because the check and the update happen under one lock.

8. When would you use each tool_choice setting? auto for normal turns. A forced function on the first turn when a lookup must happen, or when you use a tool's arguments as structured output. required when any tool must be used. none on the last allowed turn to guarantee an answer and bound the loop. And verify: some servers, Ollama's OpenAI-compatible endpoint per its docs, do not support tool_choice, so check the reply and keep a turn limit.

9. Your agent has 40 tools and picks the wrong one often. What do you do? Measure first: tool-selection accuracy on a labeled set of real requests, plus the token cost of definitions (30 tools cost 2,291 tokens per call in this module versus 533 for 3). Then reduce what each call sees: group tools by the triaged task, retrieve the top few by description with a fallback set, merge overlapping tools into one with an enum, and sharpen descriptions. Published work such as RAG-MCP reports large selection gains from retrieving tools instead of listing all of them, but re-measure on your own model and requests.

10. How do you stop a model from issuing a refund it should not? Layers, all in code: the refund tool is only offered in the refund workflow's context; the customer id comes from the session, not from arguments; arguments are validated (invoice pattern, amount bounds, reason enum); a policy check compares the amount with the invoice's refundable amount before anyone is asked; a human approves every refund; the execution repeats the check under a lock; and everything is audited. In the demo a 576 USD refund on a 288 USD invoice was rejected before the approver was called.

11. What is the difference between a timeout and a sandbox? A timeout bounds how long you wait. A sandbox bounds what the code can do while it runs: CPU, memory, files, network, secrets. A thread timeout returns control but leaves the work running; a subprocess with RLIMIT_CPU, RLIMIT_AS, an empty environment, a temporary directory, and a wall-clock timeout can actually be stopped. For untrusted code you want both, plus network and filesystem isolation from a container or a stronger layer such as gVisor.

12. Does making the schema bigger or more nested hurt accuracy? It certainly costs tokens (in this module, about 26 per described field and about 33 per nesting level), and providers may reject very large or deep schemas. Research such as "Let Me Speak Freely?" found format restrictions can hurt reasoning, and JSONSchemaBench found constrained-decoding support varies with schema complexity. But the effect on a given task is an empirical question: run the same tickets with a flat and a nested schema, compare validity and accuracy with confidence intervals, and keep schemas flat until the numbers say otherwise.

Other Tools and Providers

Tool or providerWhat it isWhen to consider it instead of this module's approach
OpenAI API Structured OutputsNative json_schema with strict mode, also for function parametersIf you use OpenAI models directly; same response_format shape as here
Anthropic tool useTools defined with input_schema; responses contain tool_use blocks, results go back as tool_result blocksIf you use Claude models; the loop and guard layer carry over unchanged, only the message format differs
Gemini native API (google-genai)response_schema for structured output, function-calling modes AUTO, ANY, NONE, VALIDATEDWhen you need Gemini features the OpenAI-compatible layer (beta) does not expose
InstructorLibrary that wraps clients with pydantic response models, validation, and retriesIf you want the repair loop as a library; you give up some visibility into each attempt
Outlines, XGrammar, llama.cpp grammarsConstrained decoding for models you run yourselfSelf-hosted models where you control decoding; JSONSchemaBench compares several
vLLM and SGLang guided decodingStructured output in open-source inference serversServing open-weight models at scale (Module 13)
json-repair (PyPI)A general-purpose repairer for malformed JSONIf you would rather not maintain a parser; measure it on your own hostile corpus first, as we did
jsonschema (Python)A JSON Schema validatorWhen the schema, not a pydantic model, is your source of truth
Model Context Protocol (MCP)A standard for exposing tools from separate serversWhen many apps share tools; the guard layer still belongs on your side
Containers, gVisor, FirecrackerProcess and VM isolationTools that run code or touch files, beyond the subprocess limits shown here

Coming Up in Module 7

Module 7, Context Engineering, treats everything you put in the window as a scarce budget. You just measured one of the biggest line items: tool definitions cost more than the ticket itself, and 30 tools cost over 2,000 tokens per call. Module 7 formally introduces kb_search.py, the BM25 search behind search_kb, and asks how much retrieved knowledge, history, and tool text a request should carry, in what order, and how to tell whether extra context actually improved the answer.