Part A: Access model
Module 13: Deployment, Operations, and Economics
By the end of this module, you'll have:
- A small gateway in front of every model call, with fallback chains, retries that honor
Retry-After, per-user quotas, a kill switch, and an exact-match cache, tested against fake providers that fail on cue. - A sizing sheet for self-hosting: weights and KV-cache memory computed exactly for TinyLM and from the published configs of Llama 3.1 8B, Qwen3-8B, and gpt-oss-120b, plus measured quantization and batching tradeoffs on TinyLM.
- A total-cost-of-ownership calculator that compares a rented GPU with API prices for Brightlane's real ticket sizes and tells you the break-even volume.
- Three caching layers and a cost-aware router, each with measured hit rates, false hits, accuracy, and cost.
- Safe logs, JSONL traces, a latency and spend dashboard, and drift checks computed from those traces, plus an incident runbook.
- A change-management kit: a deprecation checker, a paired regression test for model upgrades, and feature flags with percentage rollout and automatic rollback.
Prerequisites: Modules 1 to 3 (the llm.chat helper, tokens and pricing, prefill and decode), Module 7 (cache-aware prompt layout), Module 8 (tracing agent steps), Module 10 (eval sets and noise), and Module 11 (PII handling and denial-of-wallet). Working Python and a terminal.
Where we are: Module 12 made the assistant read screenshots and documents. Everything so far has been about getting good answers. This module is about keeping them coming when a provider rate-limits you at 9 a.m., when finance asks what the assistant costs, and when a model you depend on is switched off.
How this module is organized
| Part | What it covers |
|---|---|
| Setup | Working copy, packages, and how the examples run |
| Part A: Access model | Hosted APIs vs self-hosting, cloud catalogs, privacy and residency, lock-in, and the gateway |
| Part B: Self-hosting | Memory arithmetic, serving engines, quantization, batching and utilization, total cost vs API pricing |
| Part C: Operating in production | Latency budgets and streaming, rate limits and backoff, fallback chains, caching, cost-aware routing, quotas and kill switches |
| Part D: Observability | Safe logging, tracing, dashboards, drift detection, incident response |
| Part E: Change management | Deprecations, regression-testing an upgrade, feature flags, rollback plans |
| Module Lab | Three simulated days of operating the assistant, end to end |
Setup
The examples run in order in one working copy of the course repository. Each is a complete script in examples/, and later scripts import earlier ones from the same folder (m13_gateway.py is used everywhere). No API key is needed. Where a real model would be called, a fake provider built on ScriptedLLM stands in: it replays text and fails on a script you control. That is exactly what you want for testing retries, fallbacks, and quotas, and it says nothing about model quality.
cp -r supportdesk ~/work/m13 && cd ~/work/m13
python3.11 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
pip install scikit-learn==1.9.1 scipy==1.17.1 matplotlib==3.11.2
export PYTHONPATH=.
python -m pytest -q tests/test_m13_gateway.py tests/test_m13_ops.py
Code explained
- In simple words: make a private copy of the project, install the pinned packages this module adds, and run its tests.
- What happens: the
cpgives you a sandbox so nothing you try here touches the canonical code. The three extra packages are pinned to the versions used to produce every output below: scikit-learn for TF-IDF and logistic regression, SciPy for statistical tests, and Matplotlib for one chart.PYTHONPATH=.lets scripts importsupportdesk.*; scripts inexamples/also find each other because Python puts a script's own folder on the import path. The tests exercise the gateway and the operations helpers. - Comes out: all tests pass. Timings differ per machine.
A note on the numbers in this module. Everything labeled "measured" was produced by the code shown, on a 2-core CPU that was shared with other jobs while this module was written. Timings moved by 20 to 50 percent between runs, so read their shape, not their digits. Token counts, memory arithmetic, cost math, accuracy on the ticket set, and the statistics are deterministic and will match your run exactly.
Part A: Access model
Hosted APIs vs open-weight self-hosting
There are two ways to get tokens out of a model. With a hosted API you send a request to a provider (Groq, Google, OpenAI, Anthropic, and others) and pay per token. With self-hosting you download open weights (the trained parameters, published under a license) and run them on hardware you rent or own, paying for the hardware whether or not anyone is using it.
The course's helper already speaks to both: groq and gemini are hosted APIs, and ollama is a self-hosted server on your own machine. The interesting question is not which is "better" but which fits a given workload.
| Situation | Use this | Why |
|---|---|---|
| Traffic is small or spiky, and you need the strongest models | Hosted API | You pay only for tokens used; the best proprietary models are only available this way |
| Steady, high volume on a task an open model handles well | Self-hosting (or a dedicated endpoint) | A busy GPU can beat per-token prices; Part B computes where |
| Data may not leave your network, or a regulator requires a specific location | Self-hosting, or a cloud catalog in your region | You control where prompts go and what is retained |
| You need to fine-tune weights and serve many adapters | Self-hosting an open model | You can only modify weights you have (Module 9) |
| Prototype, or a team with no one on call for GPUs | Hosted API | Operations effort is part of the cost; Part B prices it |
Cloud-provider model catalogs
Between "call a startup's API" and "run your own GPUs" sits a third option: the model catalog of the cloud you already use. You get many models (proprietary and open) behind one bill, one identity system, and your cloud's existing compliance agreements. The names change often, so here they are as of September 2026:
- Amazon Bedrock (AWS). Offers models from many vendors. Its cross-Region inference routes requests to other AWS Regions for capacity; geographic inference profiles keep that routing inside one geography, such as the US or the EU, which matters for residency (AWS docs: cross-Region inference).
- Gemini Enterprise Agent Platform (Google Cloud), the product formerly called Vertex AI. Google's product page now reads "Gemini Enterprise Agent Platform (formerly Vertex AI)"; its Model Garden lists Google, third-party, and open models (Google Cloud). Many tutorials and SDKs still say Vertex AI.
- Microsoft Foundry (Azure), renamed from Azure AI Foundry in late 2025 (Microsoft Foundry product page). It hosts OpenAI models plus a catalog of others.
Check the current name and the regional availability of each model before you design around it. A model being "in the catalog" does not mean it is available in your Region.
Privacy, residency, and compliance as deciding factors
For many teams these decide the access model before price or quality is discussed. Four terms come up:
- Data residency: where prompts and outputs are processed and stored (for example, "only in the EU").
- Retention: how long the provider keeps your prompts and outputs, and whether they are used for training. Many providers offer reduced or zero retention on paid or enterprise tiers.
- Processor agreements: contracts such as a data processing agreement (DPA) under GDPR, or a business associate agreement for US health data.
- Sub-processors: who else touches the data. A gateway that sends overflow to a second provider adds that provider to the list.
Brightlane sells to EU companies, and ticket T-1023 even asks to move a workspace to the EU region. If Brightlane promises EU processing, its fallback chain may only contain EU-processing endpoints. That is a design constraint on Part C, not an afterthought.
| Situation | Use this | Why |
|---|---|---|
| Customer contracts promise EU processing | A cloud catalog with an EU geography, or EU self-hosting | Region pinning is contractual; a US fallback would breach it |
| Tickets contain personal data you must not retain | A provider tier with zero or short retention, plus redaction before logging (Part D) | Retention is set by the provider; your logs are set by you |
| Regulated data that cannot leave your network | Self-hosting | No provider tier removes the transfer itself |
| No special constraints | Any hosted API | Pick on quality, latency, and cost |
Vendor lock-in and a provider abstraction
Lock-in is the cost of switching providers. Some of it is unavoidable (prompts tuned for one model degrade on another, as Module 4 showed), but most of it is self-inflicted: provider SDK calls scattered through the codebase, provider-specific parameters, and no way to route traffic somewhere else in an emergency.
The fix is a gateway: one internal function that every feature calls, which decides which provider serves the request. Because supportdesk.llm.chat already hides three providers behind one signature, the gateway only has to hold a list of callables with that signature. That also makes it testable: a fake provider with the same signature can fail on command.
Here is the whole gateway. It is the longest file in the module because every later part reuses it.
examples/m13_gateway.py
"""A small LLM gateway: one front door in front of several chat providers.
Every provider is a callable with the same signature as supportdesk.llm.chat,
so a real provider (llm.chat with provider="groq") and a fake one used in
tests are interchangeable. The gateway adds what production needs and a single
call does not have: fallback chains, retries with exponential backoff and
jitter that honor Retry-After, per-user quotas, a kill switch, an exact
response cache, and prefix-cache statistics.
Time is injectable (clock and sleep), so the demo and tests run instantly and
print the same numbers every time.
"""
from __future__ import annotations
import hashlib
import json
import random
import time
from collections import OrderedDict, defaultdict
from dataclasses import dataclass, field
from typing import Any, Callable
from supportdesk.llm import ChatResult
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages
ChatFn = Callable[..., ChatResult]
# Errors ------------------------------------------------------------------------
class ProviderError(Exception):
"""A failed provider call, already classified as retryable or not."""
retryable = False
def __init__(self, message: str, status: int | None = None, retry_after: float | None = None) -> None:
super().__init__(message)
self.status = status
self.retry_after = retry_after
class RateLimited(ProviderError):
retryable = True
class ServerError(ProviderError):
retryable = True
class BadRequest(ProviderError):
retryable = False
class QuotaExceeded(Exception):
pass
class KillSwitchOn(Exception):
pass
class AllProvidersFailed(Exception):
pass
def classify(exc: Exception) -> ProviderError:
"""Map an exception from the openai SDK (or a fake) onto the gateway's error types."""
if isinstance(exc, ProviderError):
return exc
import openai # imported here so fakes never need the SDK
if isinstance(exc, openai.RateLimitError):
header = exc.response.headers.get("retry-after") if exc.response is not None else None
try:
retry_after = float(header) if header is not None else None
except ValueError:
retry_after = None # an HTTP date instead of seconds; fall back to our own backoff
return RateLimited(str(exc), 429, retry_after)
if isinstance(exc, (openai.APITimeoutError, openai.APIConnectionError)):
return ServerError(str(exc), None)
if isinstance(exc, openai.APIStatusError):
if exc.status_code >= 500:
return ServerError(str(exc), exc.status_code)
return BadRequest(str(exc), exc.status_code)
raise exc
# Configuration -------------------------------------------------------------------
@dataclass(frozen=True)
class Route:
"""One step of a fallback chain: a provider name, the callable, and the model it serves."""
provider: str
call: ChatFn
model: str
@dataclass(frozen=True)
class RetryPolicy:
max_attempts: int = 3 # tries per provider before falling back
base_s: float = 0.5 # first backoff ceiling
cap_s: float = 8.0 # backoff never exceeds this
max_retry_after_s: float = 10.0 # a longer Retry-After means "go to the next provider"
deadline_s: float = 30.0 # total time budget for one gateway call
def backoff(self, attempt: int, rng: random.Random) -> float:
"""Full jitter: a random wait between 0 and min(cap, base * 2**attempt)."""
return rng.uniform(0, min(self.cap_s, self.base_s * 2 ** attempt))
@dataclass
class Quota:
usd_per_day: float = 0.50
requests_per_minute: int = 10
class QuotaBook:
"""Per-user spend per day and requests per minute, checked before and charged after a call."""
def __init__(self, default: Quota, clock: Callable[[], float]) -> None:
self.default = default
self.overrides: dict[str, Quota] = {}
self.clock = clock
self.spent: dict[tuple[str, int], float] = defaultdict(float)
self.recent: dict[str, list[float]] = defaultdict(list)
def _day(self) -> int:
return int(self.clock() // 86_400)
def check(self, user: str) -> None:
quota = self.overrides.get(user, self.default)
now = self.clock()
self.recent[user] = [t for t in self.recent[user] if now - t < 60]
if len(self.recent[user]) >= quota.requests_per_minute:
raise QuotaExceeded(f"{user}: {quota.requests_per_minute} requests per minute reached")
if self.spent[(user, self._day())] >= quota.usd_per_day:
raise QuotaExceeded(f"{user}: daily budget {quota.usd_per_day:.2f} USD reached")
self.recent[user].append(now)
def charge(self, user: str, usd: float) -> None:
self.spent[(user, self._day())] += usd
class ExactCache:
"""Response cache keyed on the full request. Least recently used entries are evicted first."""
def __init__(self, max_entries: int = 1000, ttl_s: float = 3600, clock: Callable[[], float] = time.monotonic) -> None:
self.max_entries, self.ttl_s, self.clock = max_entries, ttl_s, clock
self.store: OrderedDict[str, tuple[float, ChatResult]] = OrderedDict()
self.hits = self.misses = 0
@staticmethod
def key(route: str, messages: list[dict[str, Any]], kwargs: dict[str, Any]) -> str:
payload = json.dumps({"route": route, "messages": messages, "kwargs": kwargs}, sort_keys=True, default=str)
return hashlib.sha256(payload.encode()).hexdigest()
def get(self, key: str) -> ChatResult | None:
item = self.store.get(key)
if item is None or self.clock() - item[0] > self.ttl_s:
self.store.pop(key, None)
self.misses += 1
return None
self.store.move_to_end(key)
self.hits += 1
return item[1]
def put(self, key: str, result: ChatResult) -> None:
self.store[key] = (self.clock(), result)
self.store.move_to_end(key)
while len(self.store) > self.max_entries:
self.store.popitem(last=False)
class PrefixStats:
"""Tracks whether a request's stable prefix (every message but the last) was seen recently.
Providers with prompt caching bill a recently seen prefix at the cached rate.
The gateway cannot see the provider's cache, but it can measure how often
your traffic *could* hit it, which is what prompt layout decides.
"""
def __init__(self, ttl_s: float = 300, clock: Callable[[], float] = time.monotonic) -> None:
self.ttl_s, self.clock = ttl_s, clock
self.seen: dict[str, float] = {}
self.requests = self.hits = self.prefix_tokens_hit = self.input_tokens = 0
def observe(self, messages: list[dict[str, Any]]) -> bool:
prefix = messages[:-1]
key = hashlib.sha256(json.dumps(prefix, sort_keys=True).encode()).hexdigest()
now = self.clock()
hit = bool(prefix) and key in self.seen and now - self.seen[key] <= self.ttl_s
self.seen[key] = now
self.requests += 1
self.input_tokens += count_messages(messages)
if hit:
self.hits += 1
self.prefix_tokens_hit += count_messages(prefix) - 3 # minus the reply priming tokens
return hit
@dataclass
class GatewayResult:
result: ChatResult
route: str
provider: str
cache: str # "exact" or "miss"
prefix_hit: bool
cost_usd: float
attempts: list[dict[str, Any]] = field(default_factory=list)
# The gateway ------------------------------------------------------------------------
class Gateway:
def __init__(self, routes: dict[str, list[Route]], *, policy: RetryPolicy | None = None,
quota: Quota | None = None, clock: Callable[[], float] = time.monotonic,
sleep: Callable[[float], None] = time.sleep, seed: int = 0,
on_event: Callable[[dict[str, Any]], None] | None = None) -> None:
self.routes = routes
self.policy = policy or RetryPolicy()
self.clock, self.sleep = clock, sleep
self.rng = random.Random(seed)
self.quotas = QuotaBook(quota or Quota(), clock)
self.cache = ExactCache(clock=clock)
self.prefix = PrefixStats(clock=clock)
self.killed: dict[str, str] = {} # route name or "*" -> reason
self.disabled_providers: set[str] = set()
self.on_event = on_event or (lambda event: None)
# Operator controls
def kill(self, route: str = "*", reason: str = "") -> None:
self.killed[route] = reason or "no reason given"
def revive(self, route: str = "*") -> None:
self.killed.pop(route, None)
def disable_provider(self, provider: str) -> None:
self.disabled_providers.add(provider)
def enable_provider(self, provider: str) -> None:
self.disabled_providers.discard(provider)
# The one method callers use
def chat(self, messages: list[dict[str, Any]], *, route: str, user: str, **kwargs: Any) -> GatewayResult:
for name in ("*", route):
if name in self.killed:
self.on_event({"type": "killed", "route": route, "reason": self.killed[name]})
raise KillSwitchOn(f"route {route!r} is switched off: {self.killed[name]}")
self.quotas.check(user)
prefix_hit = self.prefix.observe(messages)
cacheable = kwargs.get("temperature", 0.0) in (0, 0.0) and not kwargs.get("tools")
key = ExactCache.key(route, messages, kwargs) if cacheable else ""
if cacheable:
cached = self.cache.get(key)
if cached is not None:
self.on_event({"type": "cache_hit", "route": route})
return GatewayResult(cached, route, cached.provider, "exact", prefix_hit, 0.0)
started = self.clock()
attempts: list[dict[str, Any]] = []
for step in self.routes[route]:
if step.provider in self.disabled_providers:
attempts.append({"provider": step.provider, "outcome": "disabled"})
continue
for attempt in range(self.policy.max_attempts):
if self.clock() - started > self.policy.deadline_s:
raise AllProvidersFailed(f"deadline {self.policy.deadline_s}s exceeded; attempts={attempts}")
try:
result = step.call(messages, model=step.model, **kwargs)
except Exception as exc: # noqa: BLE001 (classify re-raises what it does not know)
error = classify(exc)
record = {"provider": step.provider, "attempt": attempt + 1,
"outcome": type(error).__name__, "status": error.status}
attempts.append(record)
self.on_event({"type": "error", "route": route, **record})
if not error.retryable:
raise
if error.retry_after is not None and error.retry_after > self.policy.max_retry_after_s:
record["action"] = f"retry-after {error.retry_after:g}s too long, fall back"
break
if attempt + 1 == self.policy.max_attempts:
record["action"] = "attempts used up, fall back"
break
delay = self.policy.backoff(attempt, self.rng)
if error.retry_after is not None:
delay = max(delay, error.retry_after)
if self.clock() - started + delay > self.policy.deadline_s:
record["action"] = "wait would break deadline, fall back"
break
record["action"] = f"sleep {delay:.2f}s"
self.sleep(delay)
continue
usd = cost_usd(result.usage, step.model) if step.model in PRICES else 0.0
self.quotas.charge(user, usd)
attempts.append({"provider": step.provider, "attempt": attempt + 1, "outcome": "ok"})
self.on_event({"type": "ok", "route": route, "provider": step.provider,
"latency_ms": result.latency_ms, "usd": usd})
if cacheable:
self.cache.put(key, result)
return GatewayResult(result, route, step.provider, "miss", prefix_hit, usd, attempts)
raise AllProvidersFailed(f"every provider failed for route {route!r}: {attempts}")
# Fakes for demos and tests ----------------------------------------------------------
class FakeClock:
"""Simulated time: sleep() advances the clock instantly and remembers each wait."""
def __init__(self, start: float = 1_000_000.0) -> None:
self.now = start
self.sleeps: list[float] = []
def __call__(self) -> float:
return self.now
def sleep(self, seconds: float) -> None:
self.sleeps.append(round(seconds, 3))
self.now += seconds
class FakeProvider:
"""A provider that fails on a script, then answers through ScriptedLLM (not a model).
script entries: "ok", "429" (no Retry-After), "429:2.5" (Retry-After 2.5 s),
"500", "timeout", "400". After the script runs out, outcome_fn decides (or "ok").
latency_fn, if given, draws each call's latency in ms (for realistic logs).
"""
def __init__(self, name: str, script: list[str] | None = None, latency_ms: float = 400.0,
clock: FakeClock | None = None, latency_fn: Callable[[], float] | None = None,
outcome_fn: Callable[[], str] | None = None) -> None:
self.name = name
self.script = list(script or [])
self.latency_ms = latency_ms
self.latency_fn = latency_fn
self.outcome_fn = outcome_fn
self.clock = clock
self.calls = 0
self.llm = ScriptedLLM(responder=self._respond, model=f"{name}-stand-in")
def _respond(self, messages: list[dict[str, Any]], kwargs: dict[str, Any]) -> str:
subject = messages[-1]["content"].splitlines()[0][:60]
return f"[{self.name}] draft for: {subject}"
def __call__(self, messages: list[dict[str, Any]], **kwargs: Any) -> ChatResult:
self.calls += 1
latency = self.latency_fn() if self.latency_fn else self.latency_ms
if self.clock is not None:
self.clock.now += latency / 1000
outcome = self.script.pop(0) if self.script else (self.outcome_fn() if self.outcome_fn else "ok")
if outcome.startswith("429"):
retry_after = float(outcome.split(":")[1]) if ":" in outcome else None
raise RateLimited(f"{self.name}: 429 Too Many Requests", 429, retry_after)
if outcome == "500":
raise ServerError(f"{self.name}: 500 Internal Server Error", 500)
if outcome == "timeout":
raise ServerError(f"{self.name}: request timed out", None)
if outcome == "400":
raise BadRequest(f"{self.name}: 400 unsupported parameter", 400)
result = self.llm(messages, **kwargs)
result.model = kwargs.get("model", result.model)
result.provider = self.name
result.latency_ms = round(latency, 1)
return result
def ticket_messages(subject: str, body: str) -> list[dict[str, str]]:
"""Stable system prompt first, the volatile ticket last (cache-aware layout from Module 7)."""
system = ("You are Brightlane's support assistant. Draft a short, polite reply for a human agent "
"to review. Cite help-center articles by id. Never promise refunds; agents decide those.")
return [{"role": "system", "content": system},
{"role": "user", "content": f"Subject: {subject}\n\n{body}"}]
def demo() -> None:
clock = FakeClock()
groq = FakeProvider("groq", ["429:1.5", "500", "ok"], latency_ms=350, clock=clock)
gemini = FakeProvider("gemini", [], latency_ms=900, clock=clock)
local = FakeProvider("ollama", [], latency_ms=2500, clock=clock)
routes = {"draft": [Route("groq", groq, "openai/gpt-oss-120b"),
Route("gemini", gemini, "gemini-3.5-flash"),
Route("ollama", local, "qwen3:8b")]}
gw = Gateway(routes, quota=Quota(usd_per_day=0.002, requests_per_minute=5), clock=clock, sleep=clock.sleep)
msgs = ticket_messages("Charged twice this month", "My card was charged 288 USD twice. Please refund.")
print("1) groq fails with 429 (Retry-After 1.5 s), then 500, then succeeds")
r = gw.chat(msgs, route="draft", user="u-17")
for a in r.attempts:
print(" ", a)
print(f" served by {r.provider}: {r.result.text!r} cost={r.cost_usd:.6f} USD")
print(f" simulated sleeps={clock.sleeps} simulated elapsed={clock.now - 1_000_000:.2f} s")
print("2) the identical request again (temperature 0)")
r = gw.chat(msgs, route="draft", user="u-17")
print(f" cache={r.cache} provider={r.provider} cost={r.cost_usd} groq calls so far={groq.calls}")
print("3) groq rate-limited for a long time (Retry-After 30 s): fall back at once")
groq.script = ["429:30"]
other = ticket_messages("Export to CSV", "How do I export my tasks to CSV?")
r = gw.chat(other, route="draft", user="u-17")
for a in r.attempts:
print(" ", a)
print(f" served by {r.provider}: {r.result.text!r}")
print("4) one user sends 7 requests in the same minute (limit 5 per minute)")
for i in range(7):
try:
gw.chat(ticket_messages(f"Question {i}", "Where is the billing page?"), route="draft", user="u-99")
print(f" request {i}: ok")
except QuotaExceeded as exc:
print(f" request {i}: refused ({exc})")
print("5) incident: the operator flips the kill switch on the draft route")
gw.kill("draft", "drafts quoting wrong refund window, INC-311")
try:
gw.chat(msgs, route="draft", user="u-17")
except KillSwitchOn as exc:
print(f" refused: {exc}; the ticket goes to the human queue")
gw.revive("draft")
print(f"exact cache: hits={gw.cache.hits} misses={gw.cache.misses}; "
f"prefix seen before on {gw.prefix.hits} of {gw.prefix.requests} requests")
if __name__ == "__main__":
demo()
Code explained
- In simple words: a receptionist for model calls: it checks whether the line is open, whether this caller has budget left, whether we already know the answer, and then tries providers in order until one answers.
- What happens:
ProviderError,RateLimited,ServerError,BadRequest: the three kinds of failure that need different handling. Rate limits and server errors are retryable; a bad request (your bug, such as an unsupported parameter) is not, and retrying or falling back would only hide it.classify(exc): turns exceptions from theopenaiSDK (whichllm.chatuses for all three providers) into those types. For a 429 it reads theretry-afterheader in seconds; if the header holds an HTTP date instead, it ignores it and uses its own backoff.Route: one step in a fallback chain: provider name, callable, and model id. The model id is passed to the callable and used for pricing.RetryPolicy.backoff: exponential backoff with full jitter. The wait ceiling doubles per attempt (0.5 s, 1 s, 2 s, capped at 8 s) and the actual wait is a random value below that ceiling, so a thousand clients that failed together do not retry together.max_retry_after_ssays when a provider's requested wait is too long to be worth it, anddeadline_sbounds the whole call.QuotaandQuotaBook: per-user requests per minute (a sliding 60-second window) and dollars per day, checked before the call and charged after it with the real cost frompricing.cost_usd.ExactCache: a least-recently-used map from a SHA-256 of the full request (route, messages, parameters) to the result, with a time-to-live. Only deterministic requests (temperature 0, no tools) are cached, because caching a sampled answer freezes one random draw.PrefixStats: measures how often a request's stable prefix (everything but the last message) was seen within the last five minutes. That is the traffic a provider's prompt cache could bill at the cached rate (Module 2). The gateway cannot see the provider's cache, but it can see whether your prompt layout gives the cache a chance.Gateway.chat: kill switch, quota, prefix stats, exact cache, then the chain. For each provider it tries up tomax_attemptstimes; on a retryable error it sleeps for the larger of the jittered backoff andRetry-After, unless that wait is too long or would break the deadline, in which case it falls back at once. Every attempt is recorded, andon_eventlets a tracer or metrics system listen.kill,revive,disable_provider: operator controls. A kill switch stops a route (or everything, with"*"); disabling a provider drains it from every chain without a deploy.FakeClock,FakeProvider: simulated time and a provider that fails on a script ("429:1.5"means a 429 withRetry-After: 1.5) and otherwise answers throughScriptedLLM.latency_fnandoutcome_fndraw latencies and failures at random for realistic logs in Part D. Simulated time makes a 30-second outage run in microseconds and print the same numbers every time.ticket_messages: the stable system prompt first, the volatile ticket last, following Module 7's cache-aware layout.
- Comes out: running
python examples/m13_gateway.pywalks through five scenarios. The replies are ScriptedLLM text, not model output; the attempts, waits, and refusals are the real behavior of the gateway.
To use real providers, pass llm.chat with the provider fixed. This script builds the chain from whichever keys you have set:
"""Wire the gateway to real providers through supportdesk.llm.chat.
Only providers whose API key is set join the chain; Ollama joins when
OLLAMA_BASE_URL is set or USE_OLLAMA=1. Run with no keys and it tells you so.
"""
from __future__ import annotations
import functools
import os
import time
from m13_gateway import AllProvidersFailed, Gateway, Route, ticket_messages
from supportdesk import llm
def build_routes() -> dict[str, list[Route]]:
chain = []
for provider in ("groq", "gemini"):
if os.environ.get(llm.PROVIDERS[provider]["key_env"]):
chain.append(Route(provider, functools.partial(llm.chat, provider=provider),
llm.PROVIDERS[provider]["default_model"]))
if os.environ.get("OLLAMA_BASE_URL") or os.environ.get("USE_OLLAMA") == "1":
chain.append(Route("ollama", functools.partial(llm.chat, provider="ollama"), "qwen3:8b"))
return {"draft": chain}
def main() -> None:
routes = build_routes()
print("fallback chain:", [r.provider for r in routes["draft"]] or "empty (set GROQ_API_KEY, GEMINI_API_KEY or USE_OLLAMA=1)")
if not routes["draft"]:
return
gateway = Gateway(routes)
started = time.perf_counter()
try:
r = gateway.chat(ticket_messages("Charged twice this month", "My card was charged 288 USD twice."),
route="draft", user="maya", max_tokens=300)
print(f"served by {r.provider} ({r.result.model}) in {r.result.latency_ms:.0f} ms, cost {r.cost_usd:.6f} USD")
print(r.result.text)
except AllProvidersFailed as exc:
print(f"all providers failed after {time.perf_counter() - started:.1f} s")
print(str(exc)[:230] + " ...")
if __name__ == "__main__":
main()
Code explained
- In simple words: the same gateway, with real providers plugged in where the fakes were.
- What happens:
functools.partial(llm.chat, provider="groq")produces a callable with exactly the signature the gateway expects; the gateway passesmodel=and your parameters through. A provider joins the chain only if its key is set (or, for Ollama, if you say it is running), so a missing key never turns into a confusing error in the middle of a fallback. - Comes out: first with no keys at all, then with
USE_OLLAMA=1while no Ollama server is running (a real failure, captured here):textfallback chain: empty (set GROQ_API_KEY, GEMINI_API_KEY or USE_OLLAMA=1)textfallback chain: ['ollama'] all providers failed after 5.8 s every provider failed for route 'draft': [{'provider': 'ollama', 'attempt': 1, 'outcome': 'ServerError', 'status': None, 'action': 'sleep 0.42s'}, {'provider': 'ollama', 'attempt': 2, 'outcome': 'ServerError', 'status': None, 'act ...The gateway slept only about 1.2 s in total, yet the call took 5.8 s. The rest is retries you did not see:
make_clientinllm.pycreates the OpenAI client withmax_retries=2, so each of the gateway's three attempts was really three connection attempts with the SDK's own backoff. Retries multiply across layers: 3 gateway attempts times 3 SDK attempts is 9 calls to a dead server. In production, retry in exactly one layer. If you own the client, set the SDK'smax_retries=0and let the gateway decide; if you cannot, lower the gateway'smax_attempts. With a key set, the same script prints the provider, model, latency, cost, and the draft (your output will differ).