CourseLarge Language Models · Module 13: Deployment, Operations, and Economics · part 68 of 80
Part 68 · Module 13: Deployment, Operations, and Economics

Part C: Operating in production

32 min read·22 Sept 2026

The gateway from Part A is the place where operations happen. This part tunes each of its behaviors with a measurement: how long users wait, how retries behave in a crowd, what each cache layer is worth and what it gets wrong, and when to pay for a stronger model.

Latency budgets and streaming

A latency budget splits the time a user will tolerate into allowances for each step, so that a slow step is a visible overrun rather than a vague "it feels slow". Two numbers matter for anything a person watches:

  • Time to first token (TTFT): how long until the first word appears. It is dominated by queueing, network, and prefill (the model reading the whole prompt in one pass).
  • Total time: TTFT plus one decode step per output token. Long replies are slow replies.

Streaming (sending tokens as they are generated, llm.stream_chat in the course helper) does not make the total shorter. It moves the moment the user sees progress from "total time" to "TTFT", which for a 300-token draft is the difference between staring at a spinner for several seconds and reading along after a fraction of one.

Here is a budget for the agent-facing draft. The targets are Brightlane's own design choices, not measurements; the observability section measures against them.

StepBudget (p95)Why this size
Gateway checks, quota, cache lookup10 msIn-process work; anything more is a bug
Triage call (short JSON output)1,500 msShort output, so mostly prefill and network
KB search50 msLocal BM25 over 12 articles
Draft call, time to first token (streamed)1,000 msThe agent sees text appear within a second
Draft call, total4,000 msAbout 300 tokens at hosted decode speeds
Retries and fallbackWhatever is left, capped by the gateway deadlineA retry that breaks the budget should fall back instead

TinyLM makes the mechanics visible. The script streams a reply, then separates the cost of prompt length from the cost of output length, and finally measures prefix KV reuse: computing the cache of a shared prompt prefix once and starting every request from it. That is what vLLM's automatic prefix caching and SGLang's RadixAttention do, and what hosted providers bill at their cached-input rate.

examples/m13_latency.py

python
"""Latency on a self-hosted model: time to first token, streaming, and prefix KV reuse (TinyLM).

1. Streaming: yield each token as soon as it exists and record when the first one arrived.
2. Where time goes: vary prompt length and output length separately.
3. Prefix caching: compute the KV cache of a shared prompt prefix once and reuse it,
   the idea behind vLLM's automatic prefix caching and SGLang's RadixAttention.
"""
from __future__ import annotations

import os
import time
from collections.abc import Iterator

import torch
import torch.nn.functional as F

from supportdesk.tinylm import load

torch.manual_seed(0)
torch.set_num_threads(1)          # one thread: steadier timings on a shared 2-core machine
MODEL, TOKENIZER = load()


@torch.no_grad()
def stream_tinylm(prompt_ids: list[int], max_new_tokens: int, stats: dict) -> Iterator[str]:
    """Greedy decoding that yields text pieces; fills stats with ttft_ms and total_ms."""
    started = time.perf_counter()
    caches = [dict() for _ in MODEL.blocks]
    logits = MODEL(torch.tensor([prompt_ids]), caches)[0, -1]
    for step in range(max_new_tokens):
        next_id = int(logits.argmax())
        if step == 0:
            stats["ttft_ms"] = (time.perf_counter() - started) * 1000
        yield TOKENIZER.decode([next_id])
        logits = MODEL(torch.tensor([[next_id]]), caches, start_pos=len(prompt_ids) + step)[0, -1]
    stats["total_ms"] = (time.perf_counter() - started) * 1000


def timed(prompt_ids: list[int], new_tokens: int, repeats: int = 15) -> tuple[float, float]:
    """Fastest of `repeats` runs (the timeit convention): other processes can only add time."""
    runs = []
    for _ in range(repeats):
        stats: dict = {}
        for _piece in stream_tinylm(prompt_ids, new_tokens, stats):
            pass
        runs.append((stats["ttft_ms"], stats["total_ms"]))
    return min(r[0] for r in runs), min(r[1] for r in runs)


@torch.no_grad()
def forward_after_prefix(idx: torch.Tensor, caches: list[dict], start_pos: int) -> torch.Tensor:
    """TinyGPT.forward for several new tokens on top of a cache, with a correct causal mask.

    TinyGPT's own forward only masks when the cache is empty, so feeding a chunk of
    several tokens after a cached prefix lets each token see the tokens after it.
    This version uses the same weights and builds the offset mask explicitly.
    """
    t = idx.shape[1]
    x = MODEL.tok_emb(idx) + MODEL.pos_emb(torch.arange(start_pos, start_pos + t))
    for block, cache in zip(MODEL.blocks, caches):
        b, _, d = x.shape
        q, k, v = block.qkv(block.ln1(x)).split(d, dim=2)
        q, k, v = (z.view(b, t, block.n_heads, d // block.n_heads).transpose(1, 2) for z in (q, k, v))
        k, v = torch.cat([cache["k"], k], dim=2), torch.cat([cache["v"], v], dim=2)
        cache["k"], cache["v"] = k, v
        mask = torch.ones(t, k.shape[2], dtype=torch.bool).tril(diagonal=k.shape[2] - t)
        y = F.scaled_dot_product_attention(q, k, v, attn_mask=mask)
        x = x + block.proj(y.transpose(1, 2).contiguous().view(b, t, d))
        x = x + block.mlp(block.ln2(x))
    return MODEL.head(MODEL.ln_f(x))


@torch.no_grad()
def prefill(ids: list[int], prefix_caches: list[dict] | None = None, prefix_len: int = 0):
    """Prefill ids; with prefix_caches, only the tokens after the cached prefix are computed."""
    if prefix_caches is None:
        caches = [dict() for _ in MODEL.blocks]
        return MODEL(torch.tensor([ids]), caches)[0, -1], caches
    caches = [{"k": c["k"].clone(), "v": c["v"].clone()} for c in prefix_caches]   # never mutate the shared copy
    return forward_after_prefix(torch.tensor([ids[prefix_len:]]), caches, prefix_len)[0, -1], caches


def main() -> None:
    print(f"machine load (1-minute average) {os.getloadavg()[0]:.2f} on {os.cpu_count()} cores; "
          f"torch threads {torch.get_num_threads()}")
    text = ("Customer (Ben): Good morning, how much does the Team plan cost?\n"
            "Agent (Dara): Team costs 12 USD per user per month, or 10 USD billed annually.\n"
            "Customer (Ben): Thank you. One more thing: I was charged twice for the Team plan this month "
            "and I would like the duplicate refunded please. The invoice number is on my statement.\nAgent (Dara):")
    ids = TOKENIZER.encode(text).ids
    short = TOKENIZER.encode("Customer (Chen): I was charged twice this month.\nAgent (Dara):").ids
    stats: dict = {}
    pieces = list(stream_tinylm(short, 24, stats))
    print("streamed:", repr("".join(pieces)))
    print(f"first token after {stats['ttft_ms']:.1f} ms, last after {stats['total_ms']:.1f} ms")

    print(f"\nfastest of 15 runs; long prompt has {len(ids)} tokens")
    print("prompt tokens  output tokens  TTFT ms  total ms")
    for p_len, n_out in ((8, 16), (32, 16), (64, 16), (88, 16), (32, 4), (32, 32), (32, 64)):
        ttft, total = timed(ids[-p_len:], n_out)
        print(f"{len(ids[-p_len:]):>13} {n_out:>14} {ttft:>8.1f} {total:>9.1f}")

    system = ("Agent notes: Team costs 12 USD per user per month or 10 USD annually. Business costs 24 USD. "
              "Reset links expire after 30 minutes. Duplicate charges are always refunded.\n")
    tickets = ["Customer (Mo): I forgot my password.\nAgent (Lena):",
               "Customer (Ana): How much is Business?\nAgent (Kofi):",
               "Customer (Chen): I was charged twice.\nAgent (Dara):"]
    prefix_ids = TOKENIZER.encode(system).ids
    _, prefix_caches = prefill(prefix_ids)
    print(f"\nshared prefix: {len(prefix_ids)} tokens, cached once")
    for ticket in tickets:
        full = TOKENIZER.encode(system + ticket).ids
        assert full[:len(prefix_ids)] == prefix_ids, "prefix must tokenize identically"
        cold, warm = [], []
        for _ in range(25):
            s = time.perf_counter()
            logits_cold, _ = prefill(full)
            cold.append((time.perf_counter() - s) * 1000)
            s = time.perf_counter()
            logits_warm, _ = prefill(full, prefix_caches, len(prefix_ids))
            warm.append((time.perf_counter() - s) * 1000)
        diff = (logits_cold - logits_warm).abs().max().item()
        print(f"  {len(full):>3} tokens: full prefill {min(cold):5.2f} ms, "
              f"reuse prefix {min(warm):5.2f} ms (fastest of 25), max logit difference {diff:.1e}")


if __name__ == "__main__":
    main()

Code explained

  • In simple words: a stopwatch on each phase of generation, plus a demonstration of skipping work the model has already done.
  • What happens:
    • stream_tinylm: greedy decoding as a generator. It records TTFT when the first token exists and total time at the end, the same two numbers llm.StreamStats records for hosted models.
    • timed: the fastest of 15 runs, the timeit convention: background processes can only add time, so the minimum is the cleanest estimate of the work itself.
    • forward_after_prefix: runs several new tokens on top of a cached prefix with a correct causal mask. It exists because of a real limitation in the canonical TinyGPT.forward: when a cache already holds tokens and several new tokens arrive at once, it applies no causal mask, so each new token could attend to the tokens after it. generate() never does that (it decodes one token at a time), so it is not a bug for the course's use, but prefix reuse needs it. This function uses the same weights and builds the offset mask explicitly.
    • prefill: full prefill, or prefill of only the new tokens on top of a copy of the cached prefix (copied so one request can never corrupt the shared prefix).
    • main: the three experiments. The last one checks that the logits with and without reuse agree, so the speedup is not bought with a wrong answer.
  • Comes out:

SituationUse thisWhy
A person watches the reply appear (agent draft, chat)Streaming, with TTFT as the budgeted metricPerceived speed is time to first token
The output is parsed by code (triage JSON, tool calls)No streaming; cap max_tokensYou need the whole object anyway; streaming adds parsing complexity
Long, stable instructions or context in every requestPut them first and keep them byte-identical (prefix caching)Cuts prefill time and billed input
Latency budget blown by long outputsShorter output format, lower max_tokens, or a faster modelDecode time is linear in output tokens

Rate limits, quotas, and backoff

Every hosted provider enforces rate limits: requests per minute, tokens per minute, and sometimes tokens per day, per API key or per organization. Exceed one and you get HTTP 429 Too Many Requests, often with a Retry-After header saying how many seconds to wait. The gateway's RetryPolicy already does the right things, which Part A's scenarios showed on a single request: honor Retry-After, back off exponentially with jitter, and fall back if the wait is too long.

Why jitter matters only shows up in a crowd. At 9 a.m. every support agent opens the queue, and 500 draft requests arrive in the same second against a provider limit of 100 per second. The simulation compares three retry strategies; the third uses the gateway's own RetryPolicy.backoff.

examples/m13_backoff.py

python
"""Why backoff needs jitter: 500 agents hit a rate-limited provider at the same moment.

The provider accepts at most 10 requests per 100 ms slot (100 per second) and
answers 429 to the rest. Each client retries on its own schedule. We count how
many requests the provider had to reject and when each client finally got in.
Simulated time, seeded: the numbers are exact for this setup.
"""
from __future__ import annotations

import heapq
import random
import statistics

from m13_gateway import RetryPolicy

CLIENTS, CAPACITY_PER_SLOT, SLOT_S = 500, 10, 0.1


def simulate(strategy: str, seed: int = 0) -> dict[str, float]:
    rng = random.Random(seed)
    policy = RetryPolicy(base_s=0.5, cap_s=8.0)
    queue = [(0.0, client, 0) for client in range(CLIENTS)]      # (send time, client, attempt)
    heapq.heapify(queue)
    used: dict[int, int] = {}
    done, rejected = [], 0
    while queue:
        t, client, attempt = heapq.heappop(queue)
        slot = int(t / SLOT_S + 1e-9)
        if used.get(slot, 0) < CAPACITY_PER_SLOT:
            used[slot] = used.get(slot, 0) + 1
            done.append(t)
            continue
        rejected += 1
        if strategy == "fixed 1 s":
            wait = 1.0
        elif strategy == "exponential, no jitter":
            wait = min(policy.cap_s, policy.base_s * 2 ** attempt)
        else:
            wait = policy.backoff(attempt, rng)                    # the gateway's full jitter
        heapq.heappush(queue, (t + wait, client, attempt + 1))
    return {"rejected": rejected, "p50": statistics.median(done), "p95": sorted(done)[int(0.95 * len(done))],
            "last": max(done)}


if __name__ == "__main__":
    print(f"{CLIENTS} clients at t=0, capacity {CAPACITY_PER_SLOT / SLOT_S:.0f} requests/s "
          f"(the burst needs at least {CLIENTS / CAPACITY_PER_SLOT * SLOT_S:.1f} s)")
    print(f"{'strategy':<24}{'429s sent':>10}{'p50 wait s':>11}{'p95 wait s':>11}{'last in s':>10}")
    for strategy in ("fixed 1 s", "exponential, no jitter", "exponential, full jitter"):
        r = simulate(strategy)
        print(f"{strategy:<24}{r['rejected']:>10}{r['p50']:>11.2f}{r['p95']:>11.2f}{r['last']:>10.2f}")

Code explained

  • In simple words: 500 people push through one door that lets 10 through every tenth of a second, and we count how many bounce off before everyone is in.
  • What happens: a priority queue holds each client's next send time. A request that lands in a 100 ms slot with capacity left is served; otherwise it gets a 429 and is rescheduled by its strategy: a fixed 1 second wait, exponential backoff without jitter (0.5, 1, 2, 4, 8, 8 seconds), or exponential backoff with full jitter (a random wait between zero and that ceiling). Time is simulated, so the run is instant and exact.
  • Comes out:

    text
    500 clients at t=0, capacity 100 requests/s (the burst needs at least 5.0 s)
    strategy                 429s sent p50 wait s p95 wait s last in s
    fixed 1 s                    12250      24.50      47.00     49.00
    exponential, no jitter       12250     171.50     351.50    367.50
    exponential, full jitter      1607       2.48       6.70      9.87
    

    Without jitter, clients that failed together retry together. The fixed and the exponential strategies both send exactly 12,250 rejected requests (490 + 480 + ... + 10: every round, the whole crowd hits one slot and only 10 get in), and exponential backoff without jitter makes it far worse in time, because the synchronized crowd now waits 8 seconds between its collisions: the last client gets in after 6 minutes. Full jitter spreads the retries across the window, cutting rejected requests by 87 percent and getting everyone in within 10 seconds of a burst that needs at least 5. This is the thundering herd problem, and jitter is the fix.

Quotas are the other side of rate limits: limits you impose on your own users, so one runaway script or one attacker (Module 11's denial of wallet) cannot spend the whole organization's budget or rate limit. The gateway's QuotaBook has two: requests per minute (protects capacity) and dollars per day (protects the bill). Scenario 4 in Part A showed the per-minute limit refusing the sixth request.

SituationUse thisWhy
Provider returns 429 with a short Retry-AfterWait at least that long, with jitter on topThe provider told you when capacity returns
429 with a long Retry-After, and another provider is allowedFall back immediatelyWaiting 30 s breaks every latency budget
Many clients share one limitFull jitter, one retry layer onlySynchronized retries multiply the load that caused the 429
One user or script could flood the systemPer-user requests-per-minute and dollars-per-day quotasContains both accidents and denial-of-wallet attacks

Fallback chains across providers and models

A fallback chain is an ordered list of routes: if the first provider fails in a retryable way, try the next. The gateway implements it; Part A's scenario 3 showed a long Retry-After on Groq sending the request to Gemini with no wait at all. Four things decide whether a chain actually helps:

  1. Independence. Two routes to the same underlying model in the same region fail together. A useful chain crosses providers or at least regions.
  2. Availability arithmetic. If each provider is independently unavailable 0.5 percent of the time, a two-provider chain is unavailable about 0.005 x 0.005 = 0.0025 percent of the time. Correlated failures (a shared cloud region, a shared upstream) break that multiplication, which is why the first point matters.
  3. Equivalence. The fallback model is a different model. Its output format, refusal behavior, and quality differ, so every model in a chain must pass the same eval suite (Module 10) with the prompt it will actually receive. Some teams keep a prompt variant per model.
  4. Constraints travel with the request. Part A's residency rule applies here: a request that must stay in the EU may only fall back to EU routes. Build chains per data class, not one global chain.

What never falls back: a BadRequest (HTTP 400). It means your request is wrong, and sending it elsewhere only hides the bug, as the gateway test test_bad_request_is_not_retried_or_fallen_back pins down.

Caching layers: exact, prefix, semantic

There are three caches, and they save different things:

LayerKeyWhat it savesRisk
ExactHash of the whole requestThe entire call (cost and latency)Serving a stale answer after a KB or prompt change
PrefixThe stable leading tokens of the promptPrefill work and billed input tokens (at the provider or your engine)Almost none; it only needs a stable layout
SemanticSimilarity of meaning to an earlier requestThe entire call, for similar but not identical requestsAnswering a different question with a confident wrong reply

Exact caching is in the gateway and only caches deterministic requests (temperature 0, no tools). Real tickets are almost never byte-identical, so its hit rate on raw tickets is near zero; it earns its keep on repeated internal calls such as re-triaging the same ticket after an edit, or evaluation runs. Put a version string for the prompt and the KB in the key, so a change invalidates old entries.

Prefix caching needs no code of yours beyond a disciplined layout, but the layout is easy to break. Module 7 taught cache-aware ordering (stable first, volatile last). This script measures what it is worth on the 72 real tickets, using the gateway's PrefixStats:

python
"""How prompt layout decides prefix-cache hits, measured on the 72 tickets with the gateway's PrefixStats.

Layout A keeps the long, stable instructions and help-center excerpt first and
the ticket last. Layout B puts the ticket inside the system message (a common
mistake), so no two requests share a prefix. Savings use gemini-3.5-flash
input prices from supportdesk.pricing (1.50 USD per 1M, cached 0.15 USD per 1M).
"""
from __future__ import annotations

from m13_gateway import FakeClock, PrefixStats
from supportdesk.data import get_article, load_tickets
from supportdesk.pricing import PRICES

INSTRUCTIONS = ("You are Brightlane's support assistant. Draft a short, polite reply for a human agent to review. "
                "Cite help-center articles by id. Never promise refunds; agents decide those.\n\n"
                + "\n\n".join(get_article(a).body for a in ("billing-plans", "billing-refunds", "account-login")))


def layout_a(ticket) -> list[dict]:
    return [{"role": "system", "content": INSTRUCTIONS}, {"role": "user", "content": ticket.text}]


def layout_b(ticket) -> list[dict]:
    return [{"role": "system", "content": f"Ticket {ticket.id}: {ticket.subject}\n\n{INSTRUCTIONS}"},
            {"role": "user", "content": ticket.body}]


if __name__ == "__main__":
    price = PRICES["gemini-3.5-flash"]
    for name, layout in (("A: stable prefix, ticket last", layout_a), ("B: ticket inside system", layout_b)):
        clock = FakeClock()
        stats = PrefixStats(ttl_s=300, clock=clock)
        for t in load_tickets():
            stats.observe(layout(t))
            clock.now += 2.0                      # one ticket every 2 seconds, inside the cache lifetime
        uncached = stats.input_tokens * price.input / 1e6
        cached = ((stats.input_tokens - stats.prefix_tokens_hit) * price.input + stats.prefix_tokens_hit * price.cached_input) / 1e6
        print(f"{name:<31} prefix hits {stats.hits:>2}/{stats.requests}, "
              f"{stats.prefix_tokens_hit / stats.input_tokens:5.1%} of input tokens cacheable, "
              f"input USD per 1k tickets {uncached * 1000 / stats.requests:.3f} without caching, "
              f"{cached * 1000 / stats.requests:.3f} with")

Code explained

  • In simple words: the same prompt contents in two orders, and a count of how often each order lets the provider reuse work.
  • What happens: layout A puts the instructions and three help-center articles in the system message and the ticket in the user message. Layout B puts the ticket id and subject at the top of the system message, a common mistake when someone "adds context" to the system prompt. Both run through PrefixStats with one ticket every 2 seconds (inside the 300-second window the class assumes; Groq's page says its cached data expires after 2 hours without use). Cost uses gemini-3.5-flash input prices from supportdesk.pricing: 1.50 USD per million, 0.15 USD cached.
  • Comes out:

    text
    A: stable prefix, ticket last   prefix hits 71/72, 91.0% of input tokens cacheable, input USD per 1k tickets 0.598 without caching, 0.109 with
    B: ticket inside system         prefix hits  0/72,  0.0% of input tokens cacheable, input USD per 1k tickets 0.604 without caching, 0.604 with
    

    The same tokens, reordered, cost 82 percent less on input under layout A, and zero percent less under layout B, because a prefix cache matches from the first token and one changed character at the start invalidates everything after it. Check your provider's minimum cacheable length (Groq's page gives 128 to 1,024 tokens depending on the model): a 30-token system prompt is never cached anywhere.

Semantic caching reuses the reply to an earlier, similar request. It is the one layer that can make the assistant wrong, so it needs the most careful measurement. The script uses TF-IDF vectors (each text becomes a vector of word and word-pair weights, with common words weighted down) and cosine similarity (how closely two vectors point the same way, from 0 to 1) as a cheap stand-in for an embedding model. It stores the 72 real tickets, then looks up two sets of queries written for this module: 20 paraphrases of specific earlier tickets (reusing that ticket's reply is correct) and 15 near misses that look similar but need a different answer ("How much is Team?" against a stored "How much is Business?").

examples/m13_semantic_cache.py

python
"""A semantic cache for ticket replies, and an honest measurement of its hit and false-hit rates.

The cache stores the reply drafted for every earlier ticket. A new ticket that
is similar enough (cosine similarity of TF-IDF vectors, a cheap stand-in for an
embedding model) reuses the stored reply. The query set below was written for
this module: paraphrases that should reuse a specific earlier reply, and near
misses that look alike but need a different answer.
"""
from __future__ import annotations

import re

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

from supportdesk.data import load_tickets

# (text, id of the earlier ticket whose reply is correct, or None when no cached reply is correct)
PARAPHRASES = [
    ("Double charge. I see two identical charges of 288 USD for our Team subscription this month, please refund one.", "T-1001"),
    ("Business plan price. What would the Business plan cost for a team of 14 on annual billing?", "T-1002"),
    ("How to cancel. We switched to a different tool. How can I cancel our Brightlane subscription?", "T-1003"),
    ("Password reset link expired. I clicked the reset password link this morning and it says it has expired.", "T-1009"),
    ("No access to 2FA. My phone was stolen so I can't get my two-factor code, and I never saved backup codes.", "T-1010"),
    ("SSO on Team plan? Is Google Workspace single sign-on included in the Team plan?", "T-1012"),
    ("Automation runs on Business. What is the automation run limit on the Business plan, and when does it reset?", "T-1015"),
    ("Slack posts stopped. Since Monday our board updates no longer show up in our Slack channel.", "T-1017"),
    ("CSV export missing comments. The CSV export of my board has no comments in it. How do I include them?", "T-1019"),
    ("Use app offline. Does the iPhone app work without an internet connection, for example on a flight?", "T-1022"),
    ("Delete my data. Under GDPR I request erasure of all my personal data.", "T-1024"),
    ("Is Brightlane down? Nobody in our office can load Brightlane, it just spins. Outage?", "T-1026"),
    ("MS Teams. Is there an integration with Microsoft Teams?", "T-1029"),
    ("Dark theme. Could you add a dark theme to the web app?", "T-1030"),
    ("Free for 3 people? We are three people. Can we use Brightlane for free?", "T-1042"),
    ("When does my refund arrive? My refund for the double charge was approved yesterday. How long until the money is back?", "T-1045"),
    ("Incident status. Where do I see whether you have an incident at the moment?", "T-1056"),
    ("Annual savings on Team. How much cheaper per user per month is Team if we pay yearly?", "T-1067"),
    ("iOS 16. Does the iPhone app support iOS 16?", "T-1053"),
    ("15 minute lockout. Why did I get locked out for 15 minutes? Can that be shorter?", "T-1048"),
]
NEAR_MISSES = [
    ("Wrong amount. My card was charged 288 USD this month for the Team plan but we only have 10 users, it should be 120.", None),
    ("How much is Team? We are 14 people. What would Team cost per month if we pay monthly?", None),
    ("Pause instead. Please don't cancel our plan, we just want to pause it for two months. How do I do it?", None),
    ("Invite link expired. The invitation link to join my team's workspace says expired. I clicked it the next morning.", None),
    ("SSO on Business? Does the Business plan include SSO with Okta?", None),
    ("Automation limit on Team. How many automation runs do we get on Team and when does it reset?", None),
    ("Too many Slack notifications. Board notifications post too often to #eng-updates in Slack. Can we reduce them?", None),
    ("Import from CSV. I want to import my tasks from a CSV file into a board. How do I get them in?", None),
    ("Copy of my data. Please send me a copy of all personal data you hold about me under GDPR article 15.", None),
    ("Mobile app spinner. Brightlane is not loading on my phone app only, just a spinner. The website works.", None),
    ("Refund denied. You denied my duplicate charge refund yesterday. Why?", None),
    ("Free for 30? Is it free for a team of 30?", None),
    ("Member resets owner 2FA. I'm a regular member. Can I reset the owner's 2FA device?", None),
    ("Annual discount on Business. How much do we save per user per month going annual on Business?", None),
    ("Android version. Will the app work on my Android 9 phone?", None),
]


SLOTS = {
    "plan": re.compile(r"\b(free|team|business|enterprise)\b"),
    "number": re.compile(r"\b\d+\b"),
    "platform": re.compile(r"\b(android|iphone|ios|web)\b"),
    "negation": re.compile(r"\b(don't|not|never|denied|no longer)\b"),
}


def slots(text: str) -> tuple[frozenset[str], ...]:
    """The details that change the right answer even when the wording is similar."""
    lower = text.lower()
    return tuple(frozenset(pattern.findall(lower)) for pattern in SLOTS.values())


class SemanticCache:
    def __init__(self, texts: list[str], ids: list[str], threshold: float, guard: bool = False) -> None:
        self.vectorizer = TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True)
        self.matrix = self.vectorizer.fit_transform(texts)
        self.ids, self.threshold, self.guard = ids, threshold, guard
        self.slots = [slots(t) for t in texts]

    def lookup(self, text: str) -> tuple[str | None, float]:
        """Return (id of the cached ticket to reuse, similarity), or (None, best similarity) on a miss."""
        scores = cosine_similarity(self.vectorizer.transform([text]), self.matrix)[0]
        best = int(scores.argmax())
        if self.guard and slots(text) != self.slots[best]:
            return None, float(scores[best])
        return (self.ids[best] if scores[best] >= self.threshold else None), float(scores[best])


def evaluate(threshold: float, guard: bool = False, verbose: bool = False) -> dict[str, int]:
    tickets = load_tickets()
    cache = SemanticCache([f"{t.subject}. {t.body}" for t in tickets], [t.id for t in tickets], threshold, guard)
    counts = {"correct_hit": 0, "wrong_hit": 0, "missed": 0, "near_miss_hit": 0}
    for text, expected in PARAPHRASES + NEAR_MISSES:
        hit, score = cache.lookup(text)
        if expected is not None:
            key = "missed" if hit is None else ("correct_hit" if hit == expected else "wrong_hit")
        else:
            key = "near_miss_hit" if hit is not None else "missed"
        counts[key] += 1
        if verbose and key in ("wrong_hit", "near_miss_hit"):
            print(f"  FALSE HIT {score:.2f}: {text[:62]!r} reused {hit}")
    return counts


def table(guard: bool) -> None:
    print(f"{'threshold':>9} {'correct hits':>13} {'wrong hits':>11} {'near-miss hits':>15} {'hit rate':>9} {'false-hit share':>16}")
    for threshold in (0.3, 0.4, 0.5, 0.6, 0.7, 0.8):
        c = evaluate(threshold, guard)
        hits = c["correct_hit"] + c["wrong_hit"] + c["near_miss_hit"]
        false = c["wrong_hit"] + c["near_miss_hit"]
        total = len(PARAPHRASES) + len(NEAR_MISSES)
        print(f"{threshold:>9.1f} {c['correct_hit']:>9}/{len(PARAPHRASES)} {c['wrong_hit']:>11} "
              f"{c['near_miss_hit']:>12}/{len(NEAR_MISSES)} {hits / total:>9.0%} {false / max(hits, 1):>16.0%}")


if __name__ == "__main__":
    print(f"{len(PARAPHRASES)} paraphrases (a reuse is correct), {len(NEAR_MISSES)} near misses (any reuse is wrong)")
    print("\nSimilarity only:")
    table(guard=False)
    print("\nFalse hits at threshold 0.5:")
    evaluate(0.5, verbose=True)
    print("\nSimilarity plus slot guard (plan, numbers, platform, negation must match):")
    table(guard=True)
    print("\nFalse hits with the guard at threshold 0.3:")
    evaluate(0.3, guard=True, verbose=True)

Code explained

  • In simple words: a lookup that says "we answered something like this before", and an audit of how often "like this" is actually the same question.
  • What happens:
    • PARAPHRASES and NEAR_MISSES: the labeled query set. Each paraphrase names the ticket whose reply is correct; near misses expect no reuse at all.
    • SLOTS and slots(text): a slot guard: small regexes that extract details which change the answer even when the wording is close (plan names, numbers, platforms, negations). A cached reply is reused only if these match exactly.
    • SemanticCache: fits a TF-IDF vectorizer on the stored texts (scikit-learn 1.9.1, word unigrams and bigrams) and returns the most similar stored ticket if it clears the threshold (and, with guard=True, if the slots match).
    • evaluate(threshold, guard): counts correct hits, wrong hits (a paraphrase matched the wrong ticket), near-miss hits (a false reuse), and misses. table sweeps the threshold.
  • Comes out:

SituationUse thisWhy
Identical internal requests (re-runs, evals, retries of the same call)Exact cache with versioned keysFree, and wrong only if you forget to version
Long stable instructions or contextPrefix caching through prompt layoutLarge savings, no correctness risk
High-volume, low-risk questions with few distinguishing details (store hours, status page link)Semantic cache with a high threshold and a slot guard, measured on near missesSaves whole calls where a wrong reuse is cheap
Answers that depend on plan, amount, account, or negation (most of Brightlane)No semantic cache for the final answer; cache retrieval results insteadA confident wrong answer costs more than the call it saved

Cost-aware routing: cheap model first, escalate on need

A router sends each request to the cheapest option that is likely to be good enough, and escalates to a stronger, more expensive model when the cheap one is unsure. It only works if the cheap tier's confidence is informative: when it says it is sure, it must usually be right.

The script makes this concrete for ticket categories. Tier 1 is a real classical classifier (TF-IDF plus logistic regression), scored with cross-validation (the tickets are split into 6 parts; each part is predicted by a model trained on the other 5, so every prediction is on a ticket the model never saw). Its accuracy and confidence are measured. Tier 2 is a ScriptedLLM stand-in that is right with a probability we set, 92 percent, and is billed at gemini-3.5-flash prices on real token counts. That 92 percent is an assumption, not a measurement: the math is the part to keep.

examples/m13_routing.py

python
"""Cost-aware routing: answer with a cheap model first, escalate to an expensive one when unsure.

Tier 1 is a REAL classical classifier (TF-IDF + logistic regression), scored
with cross-validation on all 72 tickets, so its accuracy and confidence are
measured. Tier 2 is a ScriptedLLM stand-in that is right with a probability
we SET (0.92), billed at gemini-3.5-flash prices on real token counts. Swap in
a real model and measure its accuracy before trusting any number here.
"""
from __future__ import annotations

import math
import random

import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold, cross_val_predict
from sklearn.pipeline import make_pipeline

from supportdesk.data import CATEGORIES, load_tickets
from supportdesk.pricing import cost_usd
from supportdesk.stand_in import ScriptedLLM

STRONG_ACCURACY = 0.92          # assumption for the stand-in; measure your real model
STRONG_MODEL = "gemini-3.5-flash"
SYSTEM = "Classify the Brightlane support ticket into one category: " + ", ".join(CATEGORIES) + ". Reply with the category only."


def wilson(k: int, n: int, z: float = 1.96) -> tuple[float, float]:
    """95 percent Wilson confidence interval for a proportion k/n."""
    if n == 0:
        return 0.0, 1.0
    p = k / n
    center = (p + z * z / (2 * n)) / (1 + z * z / n)
    half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
    return center - half, center + half


def cheap_tier(tickets) -> tuple[list[str], np.ndarray]:
    """Out-of-fold predictions and confidences: each ticket is scored by a model that never saw it."""
    texts = [f"{t.subject} {t.body}" for t in tickets]
    labels = [t.gold["category"] for t in tickets]
    model = make_pipeline(TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True), LogisticRegression(C=100, max_iter=3000))
    folds = StratifiedKFold(n_splits=6, shuffle=True, random_state=0)
    probs = cross_val_predict(model, texts, labels, cv=folds, method="predict_proba")
    classes = sorted(set(labels))
    return [classes[i] for i in probs.argmax(1)], probs.max(1)


def strong_stand_in(gold: dict[str, str], seed: int = 0) -> ScriptedLLM:
    """Right with probability STRONG_ACCURACY, otherwise a random wrong category. Not a model."""
    rng = random.Random(seed)

    def respond(messages, kwargs):
        truth = gold[messages[-1]["content"]]
        if rng.random() < STRONG_ACCURACY:
            return truth
        return rng.choice([c for c in CATEGORIES if c != truth])

    return ScriptedLLM(responder=respond, model=STRONG_MODEL)


def main() -> None:
    tickets = load_tickets()
    n = len(tickets)
    preds, conf = cheap_tier(tickets)
    correct = np.array([p == t.gold["category"] for p, t in zip(preds, tickets)])
    lo, hi = wilson(int(correct.sum()), n)
    print(f"cheap tier alone (cross-validated, n={n}): accuracy {correct.mean():.1%} (95% CI {lo:.0%} to {hi:.0%})")

    print("\nIs confidence informative? accuracy by confidence band:")
    for a, b in ((0.0, 0.4), (0.4, 0.5), (0.5, 0.6), (0.6, 1.01)):
        band = (conf >= a) & (conf < b)
        if band.any():
            print(f"  confidence {a:.1f} to {min(b, 1):.1f}: {int(correct[band].sum())}/{int(band.sum())} right")

    gold = {t.text: t.gold["category"] for t in tickets}
    strong_cost = sum(cost_usd(strong_stand_in(gold)([{"role": "system", "content": SYSTEM},
                                                      {"role": "user", "content": t.text}]).usage, STRONG_MODEL)
                      for t in tickets) / n
    print(f"\nstrong tier: assumed accuracy {STRONG_ACCURACY:.0%}, cost {strong_cost * 100_000:.2f} USD per 100k tickets "
          f"(real token counts, {STRONG_MODEL} prices)")

    print(f"\n{'threshold':>9} {'escalated':>10} {'expected acc':>13} {'simulated acc':>14} {'USD per 100k':>13} {'vs always':>10}")
    for threshold in (0.0, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 1.01):
        escalate = conf < threshold
        expected = (correct[~escalate].sum() + STRONG_ACCURACY * escalate.sum()) / n
        sims = []
        for seed in range(200):
            llm = strong_stand_in(gold, seed)
            right = correct[~escalate].sum()
            for t, esc in zip(tickets, escalate):
                if esc:
                    reply = llm([{"role": "system", "content": SYSTEM}, {"role": "user", "content": t.text}]).text
                    right += reply == t.gold["category"]
            sims.append(right / n)
        label = {0.0: "never", 1.01: "always"}.get(threshold, f"< {threshold:.2f}")
        print(f"{label:>9} {int(escalate.sum()):>6}/{n} {expected:>13.1%} {np.mean(sims):>9.1%} +-{np.std(sims):.1%}"
              f" {escalate.mean() * strong_cost * 100_000:>13.2f} {escalate.mean():>10.0%}")


if __name__ == "__main__":
    main()

Code explained

  • In simple words: let the cheap clerk answer when confident, pass the rest to the expensive expert, and compute what that mix costs and scores.
  • What happens:
    • wilson(k, n): a 95 percent confidence interval for a proportion, better behaved than the textbook formula at small n.
    • cheap_tier: out-of-fold class probabilities from cross_val_predict; the prediction is the most probable class and the confidence is its probability.
    • strong_stand_in: the tier-2 stand-in (not a model): right with probability STRONG_ACCURACY, otherwise a random wrong category.
    • main: accuracy of the cheap tier with its interval, accuracy by confidence band (is confidence informative?), then for each escalation threshold the expected accuracy (cheap right on kept tickets + 0.92 x escalated tickets, divided by n), a 200-seed simulation that checks the formula, and the cost.
  • Comes out:

    text
    cheap tier alone (cross-validated, n=72): accuracy 48.6% (95% CI 37% to 60%)
    
    Is confidence informative? accuracy by confidence band:
      confidence 0.0 to 0.4: 5/20 right
      confidence 0.4 to 0.5: 8/20 right
      confidence 0.5 to 0.6: 11/16 right
      confidence 0.6 to 1.0: 11/16 right
    
    strong tier: assumed accuracy 92%, cost 11.53 USD per 100k tickets (real token counts, gemini-3.5-flash prices)
    
    threshold  escalated  expected acc  simulated acc  USD per 100k  vs always
        never      0/72         48.6%     48.6% +-0.0%          0.00         0%
       < 0.35     13/72         62.4%     62.5% +-1.3%          2.08        18%
       < 0.40     20/72         67.2%     67.4% +-1.6%          3.20        28%
       < 0.45     31/72         75.7%     75.9% +-2.0%          4.96        43%
       < 0.50     40/72         81.7%     81.9% +-2.3%          6.41        56%
       < 0.55     51/72         83.2%     83.4% +-2.5%          8.17        71%
       < 0.60     56/72         86.8%     87.0% +-2.6%          8.97        78%
       always     72/72         92.0%     92.3% +-3.0%         11.53       100%
    

    The cheap tier alone gets 48.6 percent (35 of 72, interval 37 to 60 percent). It is weak because 72 tickets in five languages is very little training data, but its confidence is informative: below 0.5 it is right 13 of 40 times, above 0.5 it is right 22 of 32 times. That is what makes routing work. Escalating everything under 0.5 sends 56 percent of tickets to the strong tier, costs 56 percent of the always-escalate bill, and reaches an expected 81.7 percent. The simulation (81.9 percent, plus or minus 2.3) agrees with the formula, which is the check that the formula is right.

    Whether that trade is good depends on your prices and your error costs, not on this table. For Brightlane the strong tier costs about 11.5 USD per 100,000 tickets, so saving 44 percent of it saves about 5 USD per 100,000 tickets while losing about 10 accuracy points: not worth it for triage. Routing pays off when the strong tier is expensive (a large reasoning model, long outputs) and the cheap tier is accurate on its confident share. Measure both tiers on the same labeled set before choosing a threshold, and keep the confidence bands in your dashboard: if confidence stops being informative after a change, the router silently gets worse.

SituationUse thisWhy
Cheap tier accurate on its confident share, strong tier expensiveConfidence-threshold routingMost traffic stays cheap without losing accuracy
Cheap tier's confidence not informative (flat accuracy across bands)No routing; pick one modelA threshold would escalate at random
Errors are expensive and the strong tier is cheap (Brightlane triage)Send everything to the strong tierThe saving does not cover the accuracy lost
Requests have a verifiable output (JSON schema, tests)Escalate on verification failure, not on confidenceA failed check is a better signal than a probability

Per-user quotas and kill switches

Quotas limit one user; a kill switch limits everyone. It is an operator control that turns a feature off instantly, without a deploy, when the feature is causing harm: drafts quoting a wrong refund window, a prompt injection spreading through the KB, a cost spike. Part A's scenario 5 showed the gateway refusing the draft route with the incident id in the message, and the ticket going to the human queue.

Design points that matter in an incident:

  • Granularity: the gateway supports one route (kill("draft")), everything (kill("*")), and draining one provider from every chain (disable_provider("groq")). Add per-tenant switches if one customer's data is the problem.
  • A safe off state: killing drafts must mean "tickets go to humans", not "the support page errors". Decide the off behavior when you build the feature, and test it.
  • Speed: the switch must be read at request time from something you can change in seconds (a config service, a feature-flag system, a database row), not baked into a container image.
  • Audit: record who flipped it, when, and why. The reason string travels with every refusal so logs explain themselves.
  • Quotas with sensible defaults: Brightlane's lab uses 20 requests per minute and a daily dollar cap per agent. Set defaults from measured normal use (the dashboard gives you the p99 per user) and keep overrides for known heavy users.