CourseLarge Language Models · Module 13: Deployment, Operations, and Economics · part 66 of 80
Part 66 · Module 13: Deployment, Operations, and Economics

Part A: Access model

21 min read·22 Sept 2026

Module 13: Deployment, Operations, and Economics

By the end of this module, you'll have:

  • A small gateway in front of every model call, with fallback chains, retries that honor Retry-After, per-user quotas, a kill switch, and an exact-match cache, tested against fake providers that fail on cue.
  • A sizing sheet for self-hosting: weights and KV-cache memory computed exactly for TinyLM and from the published configs of Llama 3.1 8B, Qwen3-8B, and gpt-oss-120b, plus measured quantization and batching tradeoffs on TinyLM.
  • A total-cost-of-ownership calculator that compares a rented GPU with API prices for Brightlane's real ticket sizes and tells you the break-even volume.
  • Three caching layers and a cost-aware router, each with measured hit rates, false hits, accuracy, and cost.
  • Safe logs, JSONL traces, a latency and spend dashboard, and drift checks computed from those traces, plus an incident runbook.
  • A change-management kit: a deprecation checker, a paired regression test for model upgrades, and feature flags with percentage rollout and automatic rollback.

Prerequisites: Modules 1 to 3 (the llm.chat helper, tokens and pricing, prefill and decode), Module 7 (cache-aware prompt layout), Module 8 (tracing agent steps), Module 10 (eval sets and noise), and Module 11 (PII handling and denial-of-wallet). Working Python and a terminal.

Where we are: Module 12 made the assistant read screenshots and documents. Everything so far has been about getting good answers. This module is about keeping them coming when a provider rate-limits you at 9 a.m., when finance asks what the assistant costs, and when a model you depend on is switched off.

How this module is organized

PartWhat it covers
SetupWorking copy, packages, and how the examples run
Part A: Access modelHosted APIs vs self-hosting, cloud catalogs, privacy and residency, lock-in, and the gateway
Part B: Self-hostingMemory arithmetic, serving engines, quantization, batching and utilization, total cost vs API pricing
Part C: Operating in productionLatency budgets and streaming, rate limits and backoff, fallback chains, caching, cost-aware routing, quotas and kill switches
Part D: ObservabilitySafe logging, tracing, dashboards, drift detection, incident response
Part E: Change managementDeprecations, regression-testing an upgrade, feature flags, rollback plans
Module LabThree simulated days of operating the assistant, end to end

Setup

The examples run in order in one working copy of the course repository. Each is a complete script in examples/, and later scripts import earlier ones from the same folder (m13_gateway.py is used everywhere). No API key is needed. Where a real model would be called, a fake provider built on ScriptedLLM stands in: it replays text and fails on a script you control. That is exactly what you want for testing retries, fallbacks, and quotas, and it says nothing about model quality.

bash
cp -r supportdesk ~/work/m13 && cd ~/work/m13
python3.11 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
pip install scikit-learn==1.9.1 scipy==1.17.1 matplotlib==3.11.2
export PYTHONPATH=.
python -m pytest -q tests/test_m13_gateway.py tests/test_m13_ops.py

Code explained

  • In simple words: make a private copy of the project, install the pinned packages this module adds, and run its tests.
  • What happens: the cp gives you a sandbox so nothing you try here touches the canonical code. The three extra packages are pinned to the versions used to produce every output below: scikit-learn for TF-IDF and logistic regression, SciPy for statistical tests, and Matplotlib for one chart. PYTHONPATH=. lets scripts import supportdesk.*; scripts in examples/ also find each other because Python puts a script's own folder on the import path. The tests exercise the gateway and the operations helpers.
  • Comes out: all tests pass. Timings differ per machine.

A note on the numbers in this module. Everything labeled "measured" was produced by the code shown, on a 2-core CPU that was shared with other jobs while this module was written. Timings moved by 20 to 50 percent between runs, so read their shape, not their digits. Token counts, memory arithmetic, cost math, accuracy on the ticket set, and the statistics are deterministic and will match your run exactly.

Part A: Access model

Hosted APIs vs open-weight self-hosting

There are two ways to get tokens out of a model. With a hosted API you send a request to a provider (Groq, Google, OpenAI, Anthropic, and others) and pay per token. With self-hosting you download open weights (the trained parameters, published under a license) and run them on hardware you rent or own, paying for the hardware whether or not anyone is using it.

The course's helper already speaks to both: groq and gemini are hosted APIs, and ollama is a self-hosted server on your own machine. The interesting question is not which is "better" but which fits a given workload.

SituationUse thisWhy
Traffic is small or spiky, and you need the strongest modelsHosted APIYou pay only for tokens used; the best proprietary models are only available this way
Steady, high volume on a task an open model handles wellSelf-hosting (or a dedicated endpoint)A busy GPU can beat per-token prices; Part B computes where
Data may not leave your network, or a regulator requires a specific locationSelf-hosting, or a cloud catalog in your regionYou control where prompts go and what is retained
You need to fine-tune weights and serve many adaptersSelf-hosting an open modelYou can only modify weights you have (Module 9)
Prototype, or a team with no one on call for GPUsHosted APIOperations effort is part of the cost; Part B prices it

Cloud-provider model catalogs

Between "call a startup's API" and "run your own GPUs" sits a third option: the model catalog of the cloud you already use. You get many models (proprietary and open) behind one bill, one identity system, and your cloud's existing compliance agreements. The names change often, so here they are as of September 2026:

  • Amazon Bedrock (AWS). Offers models from many vendors. Its cross-Region inference routes requests to other AWS Regions for capacity; geographic inference profiles keep that routing inside one geography, such as the US or the EU, which matters for residency (AWS docs: cross-Region inference).
  • Gemini Enterprise Agent Platform (Google Cloud), the product formerly called Vertex AI. Google's product page now reads "Gemini Enterprise Agent Platform (formerly Vertex AI)"; its Model Garden lists Google, third-party, and open models (Google Cloud). Many tutorials and SDKs still say Vertex AI.
  • Microsoft Foundry (Azure), renamed from Azure AI Foundry in late 2025 (Microsoft Foundry product page). It hosts OpenAI models plus a catalog of others.

Check the current name and the regional availability of each model before you design around it. A model being "in the catalog" does not mean it is available in your Region.

Privacy, residency, and compliance as deciding factors

For many teams these decide the access model before price or quality is discussed. Four terms come up:

  • Data residency: where prompts and outputs are processed and stored (for example, "only in the EU").
  • Retention: how long the provider keeps your prompts and outputs, and whether they are used for training. Many providers offer reduced or zero retention on paid or enterprise tiers.
  • Processor agreements: contracts such as a data processing agreement (DPA) under GDPR, or a business associate agreement for US health data.
  • Sub-processors: who else touches the data. A gateway that sends overflow to a second provider adds that provider to the list.

Brightlane sells to EU companies, and ticket T-1023 even asks to move a workspace to the EU region. If Brightlane promises EU processing, its fallback chain may only contain EU-processing endpoints. That is a design constraint on Part C, not an afterthought.

SituationUse thisWhy
Customer contracts promise EU processingA cloud catalog with an EU geography, or EU self-hostingRegion pinning is contractual; a US fallback would breach it
Tickets contain personal data you must not retainA provider tier with zero or short retention, plus redaction before logging (Part D)Retention is set by the provider; your logs are set by you
Regulated data that cannot leave your networkSelf-hostingNo provider tier removes the transfer itself
No special constraintsAny hosted APIPick on quality, latency, and cost

Vendor lock-in and a provider abstraction

Lock-in is the cost of switching providers. Some of it is unavoidable (prompts tuned for one model degrade on another, as Module 4 showed), but most of it is self-inflicted: provider SDK calls scattered through the codebase, provider-specific parameters, and no way to route traffic somewhere else in an emergency.

The fix is a gateway: one internal function that every feature calls, which decides which provider serves the request. Because supportdesk.llm.chat already hides three providers behind one signature, the gateway only has to hold a list of callables with that signature. That also makes it testable: a fake provider with the same signature can fail on command.

.

Here is the whole gateway. It is the longest file in the module because every later part reuses it.

examples/m13_gateway.py

python
"""A small LLM gateway: one front door in front of several chat providers.

Every provider is a callable with the same signature as supportdesk.llm.chat,
so a real provider (llm.chat with provider="groq") and a fake one used in
tests are interchangeable. The gateway adds what production needs and a single
call does not have: fallback chains, retries with exponential backoff and
jitter that honor Retry-After, per-user quotas, a kill switch, an exact
response cache, and prefix-cache statistics.

Time is injectable (clock and sleep), so the demo and tests run instantly and
print the same numbers every time.
"""
from __future__ import annotations

import hashlib
import json
import random
import time
from collections import OrderedDict, defaultdict
from dataclasses import dataclass, field
from typing import Any, Callable

from supportdesk.llm import ChatResult
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages

ChatFn = Callable[..., ChatResult]


# Errors ------------------------------------------------------------------------

class ProviderError(Exception):
    """A failed provider call, already classified as retryable or not."""

    retryable = False

    def __init__(self, message: str, status: int | None = None, retry_after: float | None = None) -> None:
        super().__init__(message)
        self.status = status
        self.retry_after = retry_after


class RateLimited(ProviderError):
    retryable = True


class ServerError(ProviderError):
    retryable = True


class BadRequest(ProviderError):
    retryable = False


class QuotaExceeded(Exception):
    pass


class KillSwitchOn(Exception):
    pass


class AllProvidersFailed(Exception):
    pass


def classify(exc: Exception) -> ProviderError:
    """Map an exception from the openai SDK (or a fake) onto the gateway's error types."""
    if isinstance(exc, ProviderError):
        return exc
    import openai  # imported here so fakes never need the SDK

    if isinstance(exc, openai.RateLimitError):
        header = exc.response.headers.get("retry-after") if exc.response is not None else None
        try:
            retry_after = float(header) if header is not None else None
        except ValueError:
            retry_after = None  # an HTTP date instead of seconds; fall back to our own backoff
        return RateLimited(str(exc), 429, retry_after)
    if isinstance(exc, (openai.APITimeoutError, openai.APIConnectionError)):
        return ServerError(str(exc), None)
    if isinstance(exc, openai.APIStatusError):
        if exc.status_code >= 500:
            return ServerError(str(exc), exc.status_code)
        return BadRequest(str(exc), exc.status_code)
    raise exc


# Configuration -------------------------------------------------------------------

@dataclass(frozen=True)
class Route:
    """One step of a fallback chain: a provider name, the callable, and the model it serves."""

    provider: str
    call: ChatFn
    model: str


@dataclass(frozen=True)
class RetryPolicy:
    max_attempts: int = 3          # tries per provider before falling back
    base_s: float = 0.5            # first backoff ceiling
    cap_s: float = 8.0             # backoff never exceeds this
    max_retry_after_s: float = 10.0  # a longer Retry-After means "go to the next provider"
    deadline_s: float = 30.0       # total time budget for one gateway call

    def backoff(self, attempt: int, rng: random.Random) -> float:
        """Full jitter: a random wait between 0 and min(cap, base * 2**attempt)."""
        return rng.uniform(0, min(self.cap_s, self.base_s * 2 ** attempt))


@dataclass
class Quota:
    usd_per_day: float = 0.50
    requests_per_minute: int = 10


class QuotaBook:
    """Per-user spend per day and requests per minute, checked before and charged after a call."""

    def __init__(self, default: Quota, clock: Callable[[], float]) -> None:
        self.default = default
        self.overrides: dict[str, Quota] = {}
        self.clock = clock
        self.spent: dict[tuple[str, int], float] = defaultdict(float)
        self.recent: dict[str, list[float]] = defaultdict(list)

    def _day(self) -> int:
        return int(self.clock() // 86_400)

    def check(self, user: str) -> None:
        quota = self.overrides.get(user, self.default)
        now = self.clock()
        self.recent[user] = [t for t in self.recent[user] if now - t < 60]
        if len(self.recent[user]) >= quota.requests_per_minute:
            raise QuotaExceeded(f"{user}: {quota.requests_per_minute} requests per minute reached")
        if self.spent[(user, self._day())] >= quota.usd_per_day:
            raise QuotaExceeded(f"{user}: daily budget {quota.usd_per_day:.2f} USD reached")
        self.recent[user].append(now)

    def charge(self, user: str, usd: float) -> None:
        self.spent[(user, self._day())] += usd


class ExactCache:
    """Response cache keyed on the full request. Least recently used entries are evicted first."""

    def __init__(self, max_entries: int = 1000, ttl_s: float = 3600, clock: Callable[[], float] = time.monotonic) -> None:
        self.max_entries, self.ttl_s, self.clock = max_entries, ttl_s, clock
        self.store: OrderedDict[str, tuple[float, ChatResult]] = OrderedDict()
        self.hits = self.misses = 0

    @staticmethod
    def key(route: str, messages: list[dict[str, Any]], kwargs: dict[str, Any]) -> str:
        payload = json.dumps({"route": route, "messages": messages, "kwargs": kwargs}, sort_keys=True, default=str)
        return hashlib.sha256(payload.encode()).hexdigest()

    def get(self, key: str) -> ChatResult | None:
        item = self.store.get(key)
        if item is None or self.clock() - item[0] > self.ttl_s:
            self.store.pop(key, None)
            self.misses += 1
            return None
        self.store.move_to_end(key)
        self.hits += 1
        return item[1]

    def put(self, key: str, result: ChatResult) -> None:
        self.store[key] = (self.clock(), result)
        self.store.move_to_end(key)
        while len(self.store) > self.max_entries:
            self.store.popitem(last=False)


class PrefixStats:
    """Tracks whether a request's stable prefix (every message but the last) was seen recently.

    Providers with prompt caching bill a recently seen prefix at the cached rate.
    The gateway cannot see the provider's cache, but it can measure how often
    your traffic *could* hit it, which is what prompt layout decides.
    """

    def __init__(self, ttl_s: float = 300, clock: Callable[[], float] = time.monotonic) -> None:
        self.ttl_s, self.clock = ttl_s, clock
        self.seen: dict[str, float] = {}
        self.requests = self.hits = self.prefix_tokens_hit = self.input_tokens = 0

    def observe(self, messages: list[dict[str, Any]]) -> bool:
        prefix = messages[:-1]
        key = hashlib.sha256(json.dumps(prefix, sort_keys=True).encode()).hexdigest()
        now = self.clock()
        hit = bool(prefix) and key in self.seen and now - self.seen[key] <= self.ttl_s
        self.seen[key] = now
        self.requests += 1
        self.input_tokens += count_messages(messages)
        if hit:
            self.hits += 1
            self.prefix_tokens_hit += count_messages(prefix) - 3  # minus the reply priming tokens
        return hit


@dataclass
class GatewayResult:
    result: ChatResult
    route: str
    provider: str
    cache: str                     # "exact" or "miss"
    prefix_hit: bool
    cost_usd: float
    attempts: list[dict[str, Any]] = field(default_factory=list)


# The gateway ------------------------------------------------------------------------

class Gateway:
    def __init__(self, routes: dict[str, list[Route]], *, policy: RetryPolicy | None = None,
                 quota: Quota | None = None, clock: Callable[[], float] = time.monotonic,
                 sleep: Callable[[float], None] = time.sleep, seed: int = 0,
                 on_event: Callable[[dict[str, Any]], None] | None = None) -> None:
        self.routes = routes
        self.policy = policy or RetryPolicy()
        self.clock, self.sleep = clock, sleep
        self.rng = random.Random(seed)
        self.quotas = QuotaBook(quota or Quota(), clock)
        self.cache = ExactCache(clock=clock)
        self.prefix = PrefixStats(clock=clock)
        self.killed: dict[str, str] = {}          # route name or "*" -> reason
        self.disabled_providers: set[str] = set()
        self.on_event = on_event or (lambda event: None)

    # Operator controls
    def kill(self, route: str = "*", reason: str = "") -> None:
        self.killed[route] = reason or "no reason given"

    def revive(self, route: str = "*") -> None:
        self.killed.pop(route, None)

    def disable_provider(self, provider: str) -> None:
        self.disabled_providers.add(provider)

    def enable_provider(self, provider: str) -> None:
        self.disabled_providers.discard(provider)

    # The one method callers use
    def chat(self, messages: list[dict[str, Any]], *, route: str, user: str, **kwargs: Any) -> GatewayResult:
        for name in ("*", route):
            if name in self.killed:
                self.on_event({"type": "killed", "route": route, "reason": self.killed[name]})
                raise KillSwitchOn(f"route {route!r} is switched off: {self.killed[name]}")
        self.quotas.check(user)
        prefix_hit = self.prefix.observe(messages)

        cacheable = kwargs.get("temperature", 0.0) in (0, 0.0) and not kwargs.get("tools")
        key = ExactCache.key(route, messages, kwargs) if cacheable else ""
        if cacheable:
            cached = self.cache.get(key)
            if cached is not None:
                self.on_event({"type": "cache_hit", "route": route})
                return GatewayResult(cached, route, cached.provider, "exact", prefix_hit, 0.0)

        started = self.clock()
        attempts: list[dict[str, Any]] = []
        for step in self.routes[route]:
            if step.provider in self.disabled_providers:
                attempts.append({"provider": step.provider, "outcome": "disabled"})
                continue
            for attempt in range(self.policy.max_attempts):
                if self.clock() - started > self.policy.deadline_s:
                    raise AllProvidersFailed(f"deadline {self.policy.deadline_s}s exceeded; attempts={attempts}")
                try:
                    result = step.call(messages, model=step.model, **kwargs)
                except Exception as exc:  # noqa: BLE001  (classify re-raises what it does not know)
                    error = classify(exc)
                    record = {"provider": step.provider, "attempt": attempt + 1,
                              "outcome": type(error).__name__, "status": error.status}
                    attempts.append(record)
                    self.on_event({"type": "error", "route": route, **record})
                    if not error.retryable:
                        raise
                    if error.retry_after is not None and error.retry_after > self.policy.max_retry_after_s:
                        record["action"] = f"retry-after {error.retry_after:g}s too long, fall back"
                        break
                    if attempt + 1 == self.policy.max_attempts:
                        record["action"] = "attempts used up, fall back"
                        break
                    delay = self.policy.backoff(attempt, self.rng)
                    if error.retry_after is not None:
                        delay = max(delay, error.retry_after)
                    if self.clock() - started + delay > self.policy.deadline_s:
                        record["action"] = "wait would break deadline, fall back"
                        break
                    record["action"] = f"sleep {delay:.2f}s"
                    self.sleep(delay)
                    continue
                usd = cost_usd(result.usage, step.model) if step.model in PRICES else 0.0
                self.quotas.charge(user, usd)
                attempts.append({"provider": step.provider, "attempt": attempt + 1, "outcome": "ok"})
                self.on_event({"type": "ok", "route": route, "provider": step.provider,
                               "latency_ms": result.latency_ms, "usd": usd})
                if cacheable:
                    self.cache.put(key, result)
                return GatewayResult(result, route, step.provider, "miss", prefix_hit, usd, attempts)
        raise AllProvidersFailed(f"every provider failed for route {route!r}: {attempts}")


# Fakes for demos and tests ----------------------------------------------------------

class FakeClock:
    """Simulated time: sleep() advances the clock instantly and remembers each wait."""

    def __init__(self, start: float = 1_000_000.0) -> None:
        self.now = start
        self.sleeps: list[float] = []

    def __call__(self) -> float:
        return self.now

    def sleep(self, seconds: float) -> None:
        self.sleeps.append(round(seconds, 3))
        self.now += seconds


class FakeProvider:
    """A provider that fails on a script, then answers through ScriptedLLM (not a model).

    script entries: "ok", "429" (no Retry-After), "429:2.5" (Retry-After 2.5 s),
    "500", "timeout", "400". After the script runs out, outcome_fn decides (or "ok").
    latency_fn, if given, draws each call's latency in ms (for realistic logs).
    """

    def __init__(self, name: str, script: list[str] | None = None, latency_ms: float = 400.0,
                 clock: FakeClock | None = None, latency_fn: Callable[[], float] | None = None,
                 outcome_fn: Callable[[], str] | None = None) -> None:
        self.name = name
        self.script = list(script or [])
        self.latency_ms = latency_ms
        self.latency_fn = latency_fn
        self.outcome_fn = outcome_fn
        self.clock = clock
        self.calls = 0
        self.llm = ScriptedLLM(responder=self._respond, model=f"{name}-stand-in")

    def _respond(self, messages: list[dict[str, Any]], kwargs: dict[str, Any]) -> str:
        subject = messages[-1]["content"].splitlines()[0][:60]
        return f"[{self.name}] draft for: {subject}"

    def __call__(self, messages: list[dict[str, Any]], **kwargs: Any) -> ChatResult:
        self.calls += 1
        latency = self.latency_fn() if self.latency_fn else self.latency_ms
        if self.clock is not None:
            self.clock.now += latency / 1000
        outcome = self.script.pop(0) if self.script else (self.outcome_fn() if self.outcome_fn else "ok")
        if outcome.startswith("429"):
            retry_after = float(outcome.split(":")[1]) if ":" in outcome else None
            raise RateLimited(f"{self.name}: 429 Too Many Requests", 429, retry_after)
        if outcome == "500":
            raise ServerError(f"{self.name}: 500 Internal Server Error", 500)
        if outcome == "timeout":
            raise ServerError(f"{self.name}: request timed out", None)
        if outcome == "400":
            raise BadRequest(f"{self.name}: 400 unsupported parameter", 400)
        result = self.llm(messages, **kwargs)
        result.model = kwargs.get("model", result.model)
        result.provider = self.name
        result.latency_ms = round(latency, 1)
        return result


def ticket_messages(subject: str, body: str) -> list[dict[str, str]]:
    """Stable system prompt first, the volatile ticket last (cache-aware layout from Module 7)."""
    system = ("You are Brightlane's support assistant. Draft a short, polite reply for a human agent "
              "to review. Cite help-center articles by id. Never promise refunds; agents decide those.")
    return [{"role": "system", "content": system},
            {"role": "user", "content": f"Subject: {subject}\n\n{body}"}]


def demo() -> None:
    clock = FakeClock()
    groq = FakeProvider("groq", ["429:1.5", "500", "ok"], latency_ms=350, clock=clock)
    gemini = FakeProvider("gemini", [], latency_ms=900, clock=clock)
    local = FakeProvider("ollama", [], latency_ms=2500, clock=clock)
    routes = {"draft": [Route("groq", groq, "openai/gpt-oss-120b"),
                        Route("gemini", gemini, "gemini-3.5-flash"),
                        Route("ollama", local, "qwen3:8b")]}
    gw = Gateway(routes, quota=Quota(usd_per_day=0.002, requests_per_minute=5), clock=clock, sleep=clock.sleep)
    msgs = ticket_messages("Charged twice this month", "My card was charged 288 USD twice. Please refund.")

    print("1) groq fails with 429 (Retry-After 1.5 s), then 500, then succeeds")
    r = gw.chat(msgs, route="draft", user="u-17")
    for a in r.attempts:
        print("  ", a)
    print(f"   served by {r.provider}: {r.result.text!r}  cost={r.cost_usd:.6f} USD")
    print(f"   simulated sleeps={clock.sleeps}  simulated elapsed={clock.now - 1_000_000:.2f} s")

    print("2) the identical request again (temperature 0)")
    r = gw.chat(msgs, route="draft", user="u-17")
    print(f"   cache={r.cache} provider={r.provider} cost={r.cost_usd}  groq calls so far={groq.calls}")

    print("3) groq rate-limited for a long time (Retry-After 30 s): fall back at once")
    groq.script = ["429:30"]
    other = ticket_messages("Export to CSV", "How do I export my tasks to CSV?")
    r = gw.chat(other, route="draft", user="u-17")
    for a in r.attempts:
        print("  ", a)
    print(f"   served by {r.provider}: {r.result.text!r}")

    print("4) one user sends 7 requests in the same minute (limit 5 per minute)")
    for i in range(7):
        try:
            gw.chat(ticket_messages(f"Question {i}", "Where is the billing page?"), route="draft", user="u-99")
            print(f"   request {i}: ok")
        except QuotaExceeded as exc:
            print(f"   request {i}: refused ({exc})")

    print("5) incident: the operator flips the kill switch on the draft route")
    gw.kill("draft", "drafts quoting wrong refund window, INC-311")
    try:
        gw.chat(msgs, route="draft", user="u-17")
    except KillSwitchOn as exc:
        print(f"   refused: {exc}; the ticket goes to the human queue")
    gw.revive("draft")

    print(f"exact cache: hits={gw.cache.hits} misses={gw.cache.misses}; "
          f"prefix seen before on {gw.prefix.hits} of {gw.prefix.requests} requests")


if __name__ == "__main__":
    demo()

Code explained

  • In simple words: a receptionist for model calls: it checks whether the line is open, whether this caller has budget left, whether we already know the answer, and then tries providers in order until one answers.
  • What happens:
    • ProviderError, RateLimited, ServerError, BadRequest: the three kinds of failure that need different handling. Rate limits and server errors are retryable; a bad request (your bug, such as an unsupported parameter) is not, and retrying or falling back would only hide it.
    • classify(exc): turns exceptions from the openai SDK (which llm.chat uses for all three providers) into those types. For a 429 it reads the retry-after header in seconds; if the header holds an HTTP date instead, it ignores it and uses its own backoff.
    • Route: one step in a fallback chain: provider name, callable, and model id. The model id is passed to the callable and used for pricing.
    • RetryPolicy.backoff: exponential backoff with full jitter. The wait ceiling doubles per attempt (0.5 s, 1 s, 2 s, capped at 8 s) and the actual wait is a random value below that ceiling, so a thousand clients that failed together do not retry together. max_retry_after_s says when a provider's requested wait is too long to be worth it, and deadline_s bounds the whole call.
    • Quota and QuotaBook: per-user requests per minute (a sliding 60-second window) and dollars per day, checked before the call and charged after it with the real cost from pricing.cost_usd.
    • ExactCache: a least-recently-used map from a SHA-256 of the full request (route, messages, parameters) to the result, with a time-to-live. Only deterministic requests (temperature 0, no tools) are cached, because caching a sampled answer freezes one random draw.
    • PrefixStats: measures how often a request's stable prefix (everything but the last message) was seen within the last five minutes. That is the traffic a provider's prompt cache could bill at the cached rate (Module 2). The gateway cannot see the provider's cache, but it can see whether your prompt layout gives the cache a chance.
    • Gateway.chat: kill switch, quota, prefix stats, exact cache, then the chain. For each provider it tries up to max_attempts times; on a retryable error it sleeps for the larger of the jittered backoff and Retry-After, unless that wait is too long or would break the deadline, in which case it falls back at once. Every attempt is recorded, and on_event lets a tracer or metrics system listen.
    • kill, revive, disable_provider: operator controls. A kill switch stops a route (or everything, with "*"); disabling a provider drains it from every chain without a deploy.
    • FakeClock, FakeProvider: simulated time and a provider that fails on a script ("429:1.5" means a 429 with Retry-After: 1.5) and otherwise answers through ScriptedLLM. latency_fn and outcome_fn draw latencies and failures at random for realistic logs in Part D. Simulated time makes a 30-second outage run in microseconds and print the same numbers every time.
    • ticket_messages: the stable system prompt first, the volatile ticket last, following Module 7's cache-aware layout.
  • Comes out: running python examples/m13_gateway.py walks through five scenarios. The replies are ScriptedLLM text, not model output; the attempts, waits, and refusals are the real behavior of the gateway.

To use real providers, pass llm.chat with the provider fixed. This script builds the chain from whichever keys you have set:

python
"""Wire the gateway to real providers through supportdesk.llm.chat.

Only providers whose API key is set join the chain; Ollama joins when
OLLAMA_BASE_URL is set or USE_OLLAMA=1. Run with no keys and it tells you so.
"""
from __future__ import annotations

import functools
import os
import time

from m13_gateway import AllProvidersFailed, Gateway, Route, ticket_messages
from supportdesk import llm


def build_routes() -> dict[str, list[Route]]:
    chain = []
    for provider in ("groq", "gemini"):
        if os.environ.get(llm.PROVIDERS[provider]["key_env"]):
            chain.append(Route(provider, functools.partial(llm.chat, provider=provider),
                               llm.PROVIDERS[provider]["default_model"]))
    if os.environ.get("OLLAMA_BASE_URL") or os.environ.get("USE_OLLAMA") == "1":
        chain.append(Route("ollama", functools.partial(llm.chat, provider="ollama"), "qwen3:8b"))
    return {"draft": chain}


def main() -> None:
    routes = build_routes()
    print("fallback chain:", [r.provider for r in routes["draft"]] or "empty (set GROQ_API_KEY, GEMINI_API_KEY or USE_OLLAMA=1)")
    if not routes["draft"]:
        return
    gateway = Gateway(routes)
    started = time.perf_counter()
    try:
        r = gateway.chat(ticket_messages("Charged twice this month", "My card was charged 288 USD twice."),
                         route="draft", user="maya", max_tokens=300)
        print(f"served by {r.provider} ({r.result.model}) in {r.result.latency_ms:.0f} ms, cost {r.cost_usd:.6f} USD")
        print(r.result.text)
    except AllProvidersFailed as exc:
        print(f"all providers failed after {time.perf_counter() - started:.1f} s")
        print(str(exc)[:230] + " ...")


if __name__ == "__main__":
    main()

Code explained

  • In simple words: the same gateway, with real providers plugged in where the fakes were.
  • What happens: functools.partial(llm.chat, provider="groq") produces a callable with exactly the signature the gateway expects; the gateway passes model= and your parameters through. A provider joins the chain only if its key is set (or, for Ollama, if you say it is running), so a missing key never turns into a confusing error in the middle of a fallback.
  • Comes out: first with no keys at all, then with USE_OLLAMA=1 while no Ollama server is running (a real failure, captured here):

    text
    fallback chain: empty (set GROQ_API_KEY, GEMINI_API_KEY or USE_OLLAMA=1)
    

    text
    fallback chain: ['ollama']
    all providers failed after 5.8 s
    every provider failed for route 'draft': [{'provider': 'ollama', 'attempt': 1, 'outcome': 'ServerError', 'status': None, 'action': 'sleep 0.42s'}, {'provider': 'ollama', 'attempt': 2, 'outcome': 'ServerError', 'status': None, 'act ...
    

    The gateway slept only about 1.2 s in total, yet the call took 5.8 s. The rest is retries you did not see: make_client in llm.py creates the OpenAI client with max_retries=2, so each of the gateway's three attempts was really three connection attempts with the SDK's own backoff. Retries multiply across layers: 3 gateway attempts times 3 SDK attempts is 9 calls to a dead server. In production, retry in exactly one layer. If you own the client, set the SDK's max_retries=0 and let the gateway decide; if you cannot, lower the gateway's max_attempts. With a key set, the same script prints the provider, model, latency, cost, and the draft (your output will differ).