Part E: Change management
The models under an LLM system change more often than any other dependency. Providers retire models on a published schedule, update models behind the same name, and change prices. Your own team changes prompts weekly. This part turns each kind of change into a routine: a check that warns you, a test that says whether the change is safe, a flag that ships it gradually, and a plan to undo it.
Model deprecations and forced migrations
A deprecation is a provider's announcement that a model will be switched off on a shutdown date; after that date, requests naming it fail. You do not choose the timing, so this is a forced migration. The notice periods are short. Groq's deprecations page (console.groq.com/docs/deprecations, read 21 September 2026) shows, for example:
llama-3.1-8b-instantandllama-3.3-70b-versatile: announced 17 June 2026, shut down 16 August 2026 (60 days). Replacements:openai/gpt-oss-20bfor the 8B model, andopenai/gpt-oss-120borqwen/qwen3.6-27bfor the 70B model.qwen/qwen3-32bandmeta-llama/llama-4-scout-17b-16e-instruct: announced 17 June, shut down 17 July 2026 (30 days).moonshotai/kimi-k2-instruct-0905: users emailed 23 March, shut down 15 April 2026 (23 days).qwen/qwen3.6-27b: shut down 14 September 2026, replaced byqwen/qwen3.8-27b. So one of the recommended replacements from June was itself retired three months later.
This affects the course's own code: supportdesk/pricing.py still lists llama-3.1-8b-instant and llama-3.3-70b-versatile, which Groq retired on 16 August. Nothing in the course calls them by default, which is why nothing broke, but a learner who set LLM_MODEL=llama-3.3-70b-versatile from an older tutorial would now get an error. A check in CI finds this before a customer does:
examples/m13_deprecations.py
"""Check every model id the project depends on against known shutdown dates.
The schedule below is copied from Groq's deprecations page
(console.groq.com/docs/deprecations, read 21 Sep 2026). Keep it in version
control, update it when a provider emails you, and run this in CI so a
shutdown date turns into a failing check weeks ahead instead of an outage.
It also checks the recommended replacements, because a replacement can be
retired too.
"""
from __future__ import annotations
from datetime import date
from supportdesk.llm import PROVIDERS
from supportdesk.pricing import PRICES
SHUTDOWNS = {
# model id: (announced or None if the page gives no date, shutdown, recommended replacements)
"llama-3.1-8b-instant": (date(2026, 6, 17), date(2026, 8, 16), ["openai/gpt-oss-20b"]),
"llama-3.3-70b-versatile": (date(2026, 6, 17), date(2026, 8, 16), ["openai/gpt-oss-120b", "qwen/qwen3.6-27b"]),
"qwen/qwen3-32b": (date(2026, 6, 17), date(2026, 7, 17), ["openai/gpt-oss-120b"]),
"meta-llama/llama-4-scout-17b-16e-instruct": (date(2026, 6, 17), date(2026, 7, 17), ["openai/gpt-oss-120b"]),
"moonshotai/kimi-k2-instruct-0905": (date(2026, 3, 23), date(2026, 4, 15), ["openai/gpt-oss-120b"]),
"qwen/qwen3.6-27b": (None, date(2026, 9, 14), ["qwen/qwen3.8-27b"]),
"groq/compound": (date(2026, 8, 24), date(2026, 9, 21), []),
"groq/compound-mini": (date(2026, 8, 24), date(2026, 9, 21), []),
}
WARN_DAYS = 45
def live_replacements(model: str, today: date) -> list[str]:
"""Follow the replacement chain until a model that is not retired by `today`."""
found: list[str] = []
for candidate in SHUTDOWNS[model][2]:
if candidate in SHUTDOWNS and SHUTDOWNS[candidate][1] <= today:
found += [f"{r} (via {candidate}, itself retired)" for r in live_replacements(candidate, today)]
else:
found.append(candidate)
return found
def check(models: set[str], today: date) -> int:
problems = 0
for model in sorted(models):
if model not in SHUTDOWNS:
print(f" ok {model}")
continue
announced, shutdown, _ = SHUTDOWNS[model]
days = (shutdown - today).days
if days <= 0:
status, problems = "RETIRED ", problems + 1
elif days <= WARN_DAYS:
status, problems = "MIGRATE ", problems + 1
else:
status = "planned "
notice = f"{(shutdown - announced).days} days notice" if announced else "notice date not published"
print(f" {status} {model}: shutdown {shutdown} ({days:+d} days, {notice})")
print(f" replace with: {', '.join(live_replacements(model, today)) or 'no replacement named'}")
return problems
if __name__ == "__main__":
in_use = {p["default_model"] for p in PROVIDERS.values()} | set(PRICES)
today = date(2026, 9, 21)
print(f"models referenced by supportdesk (llm.PROVIDERS defaults and pricing.PRICES), checked {today}:")
problems = check(in_use, today)
print(f"{problems} model(s) need action")
raise SystemExit(1 if problems else 0)
Code explained
- In simple words: a calendar of shutdown dates, checked against every model id the project mentions, that fails the build when one is close or past.
- What happens:
SHUTDOWNS: the schedule, copied from the provider's page with the date it was read.Nonemarks an announcement date the page does not give.live_replacements(model, today): follows the replacement chain. If a recommended replacement is itself retired, it follows that model's replacement and says so.check(models, today): prints each model's status (ok,planned,MIGRATEwithin 45 days,RETIREDon or after the shutdown date) and counts problems.- The main block collects every model id from
llm.PROVIDERSandpricing.PRICESand exits with status 1 if anything needs action, which is what makes a CI job fail.
- Comes out:text
models referenced by supportdesk (llm.PROVIDERS defaults and pricing.PRICES), checked 2026-09-21: ok gemini-3.5-flash ok gemini-3.5-flash-lite RETIRED llama-3.1-8b-instant: shutdown 2026-08-16 (-36 days, 60 days notice) replace with: openai/gpt-oss-20b RETIRED llama-3.3-70b-versatile: shutdown 2026-08-16 (-36 days, 60 days notice) replace with: openai/gpt-oss-120b, qwen/qwen3.8-27b (via qwen/qwen3.6-27b, itself retired) ok openai/gpt-oss-120b ok openai/gpt-oss-20b ok qwen3:8b 2 model(s) need action exit code 1Two retired models, found automatically, and a replacement chain that would have sent you to a model that no longer exists either. Keep the schedule in version control, subscribe to every provider's deprecation notices (email and changelog), and check other providers' pages the same way: Google, OpenAI, and the cloud catalogs all publish retirement schedules.
Regression testing a model upgrade
Whether the migration is forced or chosen, the question is the same: is the new model at least as good on our tasks? Module 10 built the eval suite; the part that matters here is the statistics, because two versions scored on the same tickets are paired. Each ticket is scored by both, so the fair comparison looks only at the tickets where they disagree.
The script stands in two real classical classifiers for an "old" and a "new" model (word n-grams versus character n-grams, both TF-IDF with logistic regression), scores both out-of-fold on all 72 tickets, and compares them properly. Swap in two real models' predictions and the statistics stay the same.
examples/m13_upgrade.py
"""Regression-test an upgrade: compare two "model versions" on the same eval tickets, paired.
The two versions are REAL classical classifiers standing in for an old and a
new model (word n-gram vs character n-gram TF-IDF with logistic regression).
Every ticket is scored out-of-fold, so each version is judged on tickets it did
not train on. The statistics are the part to keep when you swap in real models.
"""
from __future__ import annotations
import random
import numpy as np
from scipy.stats import binomtest
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold, cross_val_predict
from sklearn.pipeline import make_pipeline
from supportdesk.data import load_tickets
VERSIONS = {
"v1 (word n-grams)": TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True),
"v2 (char n-grams)": TfidfVectorizer(analyzer="char_wb", ngram_range=(2, 5), sublinear_tf=True),
}
def predictions(tickets) -> dict[str, list[str]]:
texts = [f"{t.subject} {t.body}" for t in tickets]
labels = [t.gold["category"] for t in tickets]
folds = StratifiedKFold(n_splits=6, shuffle=True, random_state=0)
return {name: list(cross_val_predict(make_pipeline(vec, LogisticRegression(C=10, max_iter=3000)), texts, labels, cv=folds))
for name, vec in VERSIONS.items()}
def paired_bootstrap(a: np.ndarray, b: np.ndarray, rounds: int = 10_000, seed: int = 0) -> tuple[float, float]:
"""95 percent interval for accuracy(b) - accuracy(a), resampling tickets (keeping pairs together)."""
rng = np.random.default_rng(seed)
idx = rng.integers(0, len(a), size=(rounds, len(a)))
diffs = b[idx].mean(1) - a[idx].mean(1)
return float(np.percentile(diffs, 2.5)), float(np.percentile(diffs, 97.5))
def main() -> None:
tickets = load_tickets()
preds = predictions(tickets)
(n1, p1), (n2, p2) = preds.items()
a = np.array([p == t.gold["category"] for p, t in zip(p1, tickets)])
b = np.array([p == t.gold["category"] for p, t in zip(p2, tickets)])
print(f"n={len(tickets)} tickets, out-of-fold predictions")
print(f" {n1}: {a.sum()}/{len(a)} = {a.mean():.1%}")
print(f" {n2}: {b.sum()}/{len(b)} = {b.mean():.1%}")
fixed = int((~a & b).sum())
broken = int((a & ~b).sum())
print(f"\npaired view: both right {int((a & b).sum())}, both wrong {int((~a & ~b).sum())}, "
f"fixed by v2 {fixed}, broken by v2 {broken}")
test = binomtest(fixed, fixed + broken, 0.5)
lo, hi = paired_bootstrap(a.astype(float), b.astype(float))
print(f"McNemar exact test on the {fixed + broken} discordant tickets: p = {test.pvalue:.2f}")
print(f"paired bootstrap 95% interval for the accuracy change: {lo:+.1%} to {hi:+.1%}")
rng = random.Random(0)
unpaired = []
for _ in range(10_000):
unpaired.append(np.mean([b[rng.randrange(len(b))] for _ in b]) - np.mean([a[rng.randrange(len(a))] for _ in a]))
print(f"unpaired bootstrap (ignores pairing, for contrast): {np.percentile(unpaired, 2.5):+.1%} "
f"to {np.percentile(unpaired, 97.5):+.1%}")
print("\nbroken by v2 (read these before shipping):")
for t, x, y, p_old, p_new in zip(tickets, a, b, p1, p2):
if x and not y:
print(f" {t.id} [{t.language}] gold={t.gold['category']:<15} v1={p_old:<15} v2={p_new:<15} {t.subject[:40]}")
print("\nby language (right / total):")
for lang in sorted({t.language for t in tickets}):
mask = np.array([t.language == lang for t in tickets])
print(f" {lang}: v1 {a[mask].sum()}/{mask.sum()} v2 {b[mask].sum()}/{mask.sum()}")
if __name__ == "__main__":
main()
Code explained
- In simple words: grade the old and the new version on the same exam, then count only the questions where their answers differ.
- What happens:
VERSIONS: two vectorizers; the character version reads sub-word pieces, which helps with typos and non-English words.predictions: out-of-fold predictions for each version with the same six folds, so both are judged on tickets they never trained on.paired_bootstrap: resamples tickets (keeping each ticket's two results together) 10,000 times and reports the middle 95 percent of the accuracy difference.main: accuracy of each version; the four paired cells (both right, both wrong, fixed by v2, broken by v2); the McNemar exact test, a binomial test on the discordant tickets asking whether "fixed" and "broken" could be a coin flip; the paired and, for contrast, an unpaired bootstrap; the tickets v2 broke; and a per-language breakdown.
- Comes out:text
n=72 tickets, out-of-fold predictions v1 (word n-grams): 35/72 = 48.6% v2 (char n-grams): 46/72 = 63.9% paired view: both right 32, both wrong 23, fixed by v2 14, broken by v2 3 McNemar exact test on the 17 discordant tickets: p = 0.01 paired bootstrap 95% interval for the accuracy change: +5.6% to +26.4% unpaired bootstrap (ignores pairing, for contrast): -1.4% to +31.9% broken by v2 (read these before shipping): T-1038 [en] gold=billing v1=billing v2=how_to Card declined T-1056 [en] gold=how_to v1=how_to v2=billing Status page T-1061 [de] gold=bug v1=bug v2=how_to Slack Benachrichtigungen by language (right / total): de: v1 3/3 v2 2/3 en: v1 31/64 v2 41/64 es: v1 1/2 v2 2/2 hi: v1 0/1 v2 1/1 ja: v1 0/2 v2 0/2v2 scores 63.9 percent against 48.6. The paired view shows why that is convincing: v2 fixed 14 tickets and broke 3; if the versions were equally good, a 14-to-3 split would be unlikely (McNemar p = 0.01). The paired bootstrap interval, +5.6 to +26.4 points, excludes zero. The unpaired bootstrap on the same data runs from -1.4 to +31.9 and includes zero, because it throws away the knowledge that both versions saw the same tickets. With small eval sets, pairing is often the difference between a clear answer and "we cannot tell".
Then read what it broke before you ship. T-1038 ("Card declined") moved from billing to how_to, T-1056 ("Status page") from how_to to billing, and T-1061, a German Slack notification bug, to how_to. A 15-point average gain does not tell you whether those three matter more to Brightlane than the 14 fixes; the list does. The per-language table is a warning about sample sizes: 1 to 3 tickets per non-English language cannot support any conclusion, so the upgrade gate needs more non-English cases before it can say anything about them (Module 10's advice on how many cases you need).
| Situation | Use this | Why |
|---|---|---|
| Two versions scored on the same cases (pass or fail) | McNemar exact test and a paired bootstrap | Uses the pairing; far more sensitive than comparing two accuracies |
| Scores are numeric (judge scores, edit distance) | Paired bootstrap or a paired t-test on per-case differences | Same idea for continuous values |
| Any upgrade that passes the statistics | Read every case the new version broke | An average cannot tell you which regressions are unacceptable |
| A slice (language, tier) has only a few cases | Report it as "not enough data", and add cases | Small slices hide real regressions |
Feature-flagging prompt and model changes
Passing the offline suite is necessary, not sufficient: real traffic contains things the suite does not. A feature flag lets you ship a change to a small share of users, compare, and widen or withdraw it by changing configuration instead of deploying code. For LLM systems the flag's value is typically a prompt version and a model id.
Two properties matter. The assignment must be sticky (the same agent always sees the same variant, so their experience is consistent and your comparison is clean), and ramping up must keep existing treatment users in treatment. Hashing flag name + user id into 10,000 buckets gives both.
examples/m13_flags.py
"""Feature flags for prompt and model changes: sticky percentage rollout, ramp, guardrail, rollback.
A flag maps each user to a variant by hashing (flag name, user id) into one of
10,000 buckets. The same user always lands in the same bucket, so ramping from
5 to 25 percent keeps the first 5 percent where they were. The flag config is
plain data, so a rollback is a config change, not a deploy.
"""
from __future__ import annotations
import hashlib
import json
import math
import random
from dataclasses import asdict, dataclass, field
@dataclass
class Flag:
name: str
control: dict # what everyone gets today
treatment: dict # the change under test
percent: float = 0.0 # share of users on treatment, 0 to 100
history: list[str] = field(default_factory=list)
def bucket(self, user: str) -> int:
digest = hashlib.sha256(f"{self.name}:{user}".encode()).hexdigest()
return int(digest[:8], 16) % 10_000
def variant(self, user: str) -> tuple[str, dict]:
if self.bucket(user) < self.percent * 100:
return "treatment", self.treatment
return "control", self.control
def set_percent(self, percent: float, reason: str) -> None:
self.history.append(f"{self.percent:g}% -> {percent:g}%: {reason}")
self.percent = percent
def two_proportion_z(bad_a: int, n_a: int, bad_b: int, n_b: int) -> float:
"""z statistic for 'treatment error rate is higher than control' (positive = treatment worse)."""
p = (bad_a + bad_b) / (n_a + n_b)
se = math.sqrt(p * (1 - p) * (1 / n_a + 1 / n_b)) or 1e-12
return (bad_b / n_b - bad_a / n_a) / se
def main() -> None:
flag = Flag("draft_prompt",
control={"prompt": "draft-v7", "model": "openai/gpt-oss-120b"},
treatment={"prompt": "draft-v8", "model": "openai/gpt-oss-120b"})
users = [f"agent-{i}" for i in range(5000)]
flag.set_percent(5, "start canary")
first = {u for u in users if flag.variant(u)[0] == "treatment"}
print(f"5% rollout: {len(first)} of {len(users)} users on treatment ({len(first) / len(users):.1%})")
flag.set_percent(25, "canary clean for 48 h")
second = {u for u in users if flag.variant(u)[0] == "treatment"}
print(f"25% rollout: {len(second)} users; all of the first {len(first)} still on treatment: {first <= second}")
print("\nguardrail: compare the rate of drafts agents rejected, control vs treatment")
rng = random.Random(3)
rates = {"control": 0.12, "treatment": 0.17} # simulated: v8 is worse
counts = {"control": [0, 0], "treatment": [0, 0]}
for day in range(1, 4):
for u in users:
name, _ = flag.variant(u)
if rng.random() < 0.2: # about 1,000 drafts a day
counts[name][1] += 1
counts[name][0] += rng.random() < rates[name]
(bc, nc), (bt, nt) = counts["control"], counts["treatment"]
z = two_proportion_z(bc, nc, bt, nt)
print(f" day {day}: control {bc}/{nc} = {bc / nc:.1%} rejected, treatment {bt}/{nt} = {bt / nt:.1%}, z = {z:+.2f}")
if z > 2.33: # one-sided, about 1 percent false-alarm rate
flag.set_percent(0, f"auto-rollback day {day}: rejection rate z={z:.2f}")
break
print(f"\nflag now: {flag.percent:g}% on treatment; history:")
for line in flag.history:
print(" ", line)
print("\nconfig as stored:", json.dumps({k: v for k, v in asdict(flag).items() if k != "history"}))
if __name__ == "__main__":
main()
Code explained
- In simple words: a dimmer switch for a prompt change, with a safety cut-out that turns it off by itself when agents start rejecting its drafts.
- What happens:
Flag: holds the control and treatment configurations and the treatment percentage.buckethashes the flag name and user id with SHA-256 into 0 to 9,999;variantcompares that bucket with the percentage. Including the flag name means different flags split users independently.set_percentrecords every change with a reason.two_proportion_z: the standard test of whether two rates differ, here the share of drafts agents rejected under control and treatment.main: ramps from 5 to 25 percent and checks stickiness, then simulates three days of about 1,000 drafts a day where the new prompt is secretly worse (17 percent rejected against 12). After each day it tests, and rolls back automatically if the treatment is worse with z above 2.33 (a one-sided test with about a 1 percent false-alarm rate).
- Comes out:text
5% rollout: 248 of 5000 users on treatment (5.0%) 25% rollout: 1201 users; all of the first 248 still on treatment: True guardrail: compare the rate of drafts agents rejected, control vs treatment day 1: control 82/727 = 11.3% rejected, treatment 45/240 = 18.8%, z = +2.97 flag now: 0% on treatment; history: 0% -> 5%: start canary 5% -> 25%: canary clean for 48 h 25% -> 0%: auto-rollback day 1: rejection rate z=2.97 config as stored: {"name": "draft_prompt", "control": {"prompt": "draft-v7", "model": "openai/gpt-oss-120b"}, "treatment": {"prompt": "draft-v8", "model": "openai/gpt-oss-120b"}, "percent": 0}Exactly the first 248 agents stayed in treatment when the flag went to 25 percent. After one day the treatment's rejection rate was 18.8 percent against 11.3 percent, z = 2.97, and the guardrail rolled back to 0 percent on its own, recording why. The simulated rates are ours; the mechanism is what you keep. In production, measure the guardrail metric from your traces (Module 10's online signals: accepted, edited, rejected drafts), and decide the threshold and the minimum sample before the experiment starts, not while watching it.
Rollback plans
A rollback returns the system to its last known good state. Plan it before the change, because during an incident nobody has time to work out what "the previous version" was. What needs to be true:
| What changed | How you roll back | What must exist beforehand |
|---|---|---|
| Prompt version | Set the flag to 0 percent, or point the route at the previous prompt id | Prompts stored with version ids (Module 4), never edited in place |
| Model (by choice) | Set the flag back, or reorder the fallback chain | The previous model still passes the suite and is still served |
| Model (forced by deprecation) | No rollback to the old model exists after the shutdown date | A second tested option in the chain before the date |
| Provider silently updated a model | Switch to a pinned, dated model id if the provider offers one; otherwise route to the fallback | Dated ids in config, drift alerts on outputs |
| KB or retrieval change | Revert the index to the previous snapshot | Versioned index snapshots |
| Code (gateway, parsers) | Normal deploy rollback | The usual release process |
Three habits make rollbacks boring. Change one thing at a time, so you know which to undo. Keep the previous configuration live and tested for a while after a change. And rehearse: flip the kill switch and a flag rollback in staging once a quarter, and time how long it takes.
Module Lab
The lab runs the whole operations stack for three simulated days of Brightlane support and reads everything back from the traces alone, the way an on-call engineer would. Every piece is imported from this module's examples: the flag assigns each agent a prompt version, the gateway handles fallback, retries, quotas, and the kill switch, the tracer writes redacted spans, and the report computes the dashboard, the error breakdown, category drift, and the per-variant comparison.
Three things go wrong on purpose, one per day:
- Day 1: a misconfigured script (
agent-bot) sends 30 tickets in 45 seconds. - Day 2: the primary provider rate-limits half its calls with
Retry-After: 25. - Day 3: a bad product release shifts the ticket mix toward bugs and login problems, and at midday an operator flips the kill switch on drafts (incident INC-312).
examples/m13_lab.py
"""Module 13 lab: operate the Brightlane assistant for three simulated days.
Combines the module's pieces: a feature flag picks the prompt version per
agent, the gateway handles fallback, retries, quotas, caching and the kill
switch, the tracer writes redacted spans, and the dashboard, drift check and
per-variant comparison are computed from those spans alone. Providers are
fakes built on ScriptedLLM, so replies are not model output; the operations
logic is real.
"""
from __future__ import annotations
import json
import random
import statistics
from collections import Counter, defaultdict
from pathlib import Path
from m13_drift import psi
from m13_flags import Flag
from m13_gateway import FakeClock, FakeProvider, Gateway, KillSwitchOn, Quota, Route, ticket_messages
from m13_observability import START, Tracer, hash_user, percentile, redact
from supportdesk.data import CATEGORIES, load_tickets
OUT = Path(__file__).resolve().parent / "out"
def run(path: Path, per_day: int = 200, seed: int = 11) -> None:
rng = random.Random(seed)
clock = FakeClock(START)
state = {"day": 1}
def groq_outcome() -> str:
return "429:25" if state["day"] == 2 and rng.random() < 0.5 else ("500" if rng.random() < 0.01 else "ok")
groq = FakeProvider("groq", clock=clock, latency_fn=lambda: rng.lognormvariate(6.0, 0.35), outcome_fn=groq_outcome)
gemini = FakeProvider("gemini", clock=clock, latency_fn=lambda: rng.lognormvariate(6.8, 0.30))
chain = [Route("groq", groq, "openai/gpt-oss-120b"), Route("gemini", gemini, "gemini-3.5-flash")]
gw = Gateway({"triage": chain, "draft": chain}, quota=Quota(usd_per_day=0.01, requests_per_minute=20),
clock=clock, sleep=clock.sleep, seed=seed)
flag = Flag("draft_prompt", {"prompt": "draft-v7"}, {"prompt": "draft-v8"}, percent=20)
tracer = Tracer(path, clock, seed)
tickets = load_tickets()
bug_heavy = [t for t in tickets if t.gold["category"] in ("bug", "account_access")]
for day in (1, 2, 3):
state["day"] = day
for i in range(per_day):
clock.now = START + (day - 1) * 86_400 + 8 * 3600 + i * 120
if day == 3 and i == per_day // 2:
gw.kill("draft", "INC-312: drafts cite a retired refund policy")
# Day 3 follows a bad release: more bug and login tickets arrive.
t = rng.choice(bug_heavy if day == 3 and rng.random() < 0.4 else tickets)
agent = f"agent-{rng.randint(1, 60)}"
if day == 1 and 50 <= i < 80:
agent = "agent-bot" # a misconfigured script: 30 requests in 45 seconds
clock.now = START + 8 * 3600 + 50 * 120 + (i - 50) * 1.5
variant, cfg = flag.variant(agent)
body = f"{t.body} (ref {day}-{i})"
try:
with tracer.span("handle_ticket", day=day, agent=hash_user(agent), variant=variant,
category=t.gold["category"]) as root:
with tracer.span("triage") as s:
r = gw.chat(ticket_messages(t.subject, body), route="triage", user=agent)
s.attrs.update(provider=r.provider, usd=r.cost_usd)
with tracer.span("draft", prompt=cfg["prompt"]) as s:
try:
r = gw.chat(ticket_messages(t.subject, body), route="draft", user=agent)
s.attrs.update(provider=r.provider, usd=r.cost_usd, excerpt=redact(body)[:60])
except KillSwitchOn:
s.attrs["outcome"] = "human_queue"
root.attrs["human_queue"] = True
except Exception: # noqa: BLE001 (recorded on the span)
pass
tracer.close()
def report(path: Path) -> None:
spans = [json.loads(line) for line in path.open(encoding="utf-8")]
traces = defaultdict(list)
for s in spans:
traces[s["trace_id"]].append(s)
days = defaultdict(lambda: defaultdict(list))
variants = defaultdict(list)
for trace in traces.values():
root = next(s for s in trace if s["parent_id"] is None)
d = days[root["day"]]
usd = sum(s.get("usd", 0.0) for s in trace)
d["latency"].append(root["duration_ms"])
d["error"].append(root["status"] != "ok")
d["usd"].append(usd)
d["human"].append(bool(root.get("human_queue")))
d["fallback"].append(any(s.get("provider") == "gemini" for s in trace))
d["category"].append(root["category"])
if root["status"] == "ok" and not root.get("human_queue"):
variants[root["variant"]].append(usd)
print(f"{'day':>3} {'tasks':>6} {'p50 ms':>7} {'p95 ms':>7} {'errors':>7} {'fallback':>9} {'to humans':>10} {'USD/1k':>8}")
for day, d in sorted(days.items()):
print(f"{day:>3} {len(d['latency']):>6} {statistics.median(d['latency']):>7.0f} {percentile(d['latency'], 0.95):>7.0f} "
f"{sum(d['error']) / len(d['error']):>7.1%} {sum(d['fallback']) / len(d['fallback']):>9.0%} "
f"{sum(d['human']):>10} {1000 * statistics.mean(d['usd']):>8.3f}")
errors = Counter(s.get("error") for s in spans if s["status"] == "error" and s["parent_id"] is None)
print("root-span errors by type:", dict(errors) or "none")
base, today = Counter(days[1]["category"]), Counter(days[3]["category"])
print(f"category drift day 3 vs day 1: PSI {psi(base, today, CATEGORIES):.3f}; "
f"bug {base['bug']} -> {today['bug']}, account_access {base['account_access']} -> {today['account_access']}")
for name, costs in sorted(variants.items()):
print(f"variant {name:<9}: {len(costs):>3} completed tasks, {1000 * statistics.mean(costs):.3f} USD per 1k")
if __name__ == "__main__":
log = OUT / "m13_lab_traces.jsonl"
run(log)
report(log)
Code explained
- In simple words: a three-day fire drill for the assistant, with the incident report written from the flight recorder.
- What happens:
run: builds the fake providers (Groq's outcome depends on the day), one fallback chain for both routes, a gateway with 20 requests per minute and 0.01 USD per day per agent, a flag with 20 percent ondraft-v8, and a tracer. Each of 200 tickets a day becomes a root span withtriageanddraftchildren, tagged with the hashed agent id, the variant, and the category. On day 3 the ticket pool leans toward bug and account-access tickets 40 percent of the time, and halfway through the day the draft route is killed; a killed draft is recorded ashuman_queue, not as an error.report: reads only the JSONL file. It prints the per-day dashboard (p50, p95, errors, fallback share, tickets sent to humans, cost per 1,000 tasks), counts root-span errors by type, computes PSI on the category mix of day 3 against day 1 withpsifrom the drift example, and compares cost for the two flag variants.
- Comes out:text
day tasks p50 ms p95 ms errors fallback to humans USD/1k 1 200 758 1272 10.0% 0% 0 0.035 2 200 1731 3152 0.0% 72% 0 0.243 3 200 643 1273 0.0% 0% 100 0.029 root-span errors by type: {'QuotaExceeded': 20} category drift day 3 vs day 1: PSI 0.245; bug 30 -> 31, account_access 46 -> 89 variant control : 387 completed tasks, 0.125 USD per 1k variant treatment: 93 completed tasks, 0.119 USD per 1kEach incident shows up where it should, and each looks different:
- Day 1: a 10.0 percent error rate, all 20 of type
QuotaExceeded. The bot's first 10 tickets used its 20 calls for the minute (every ticket makes two gateway calls, triage and draft), and the remaining 20 tickets were refused. Latency and cost are normal. Diagnosis from the traces: one hashed agent id owns every error. Action: fix the script, and keep the quota. - Day 2: no errors at all, but fallback share 72 percent, p50 latency more than doubled (758 to 1,731 ms), and cost per 1,000 tasks up about seven-fold. This is Part D's lesson again: the fallback chain kept every ticket flowing, and only the fallback, latency, and cost panels reveal the incident.
- Day 3: 100 tickets went to the human queue after the kill switch, and category PSI reached 0.245 (just under the 0.25 "act" line), driven by account-access tickets almost doubling (46 to 89). The bug count barely moved (30 to 31): in this run the release's extra tickets happened to land mostly on login problems, a reminder that a planted shift is realized through random draws.
- Variants: 387 completed tasks on control, 93 on treatment (20 percent), with similar cost per task. Cost is only the first check; the quality comparison needs the draft rejection rates from Part E, which the fake providers cannot produce. With real models, add them to the root span and let the flag guardrail decide.
- To extend the lab, add a fourth day where the fallback model is killed too (
disable_provider("gemini")) and confirm the runbook's containment step still leaves every ticket with a human; then add a drift alert that combines chi-square and PSI as Part D recommends.
Project Milestone
The Brightlane assistant is now operable. In your working copy you should have:
examples/m13_gateway.py: the single front door for every model call, with fallback chains, retries with full jitter that honorRetry-After, a deadline, per-user request and dollar quotas, a kill switch per route or globally, provider draining, an exact cache for deterministic calls, and prefix-cache statistics.examples/m13_real_gateway.pywires it to real providers throughllm.chat.examples/m13_memory.py,m13_quantize.py,m13_batching.py,m13_tco.py: a sizing and cost sheet that answers "should we self-host?" with Brightlane's measured token counts (answer today: no, by a factor of several hundred).examples/m13_latency.py,m13_backoff.py,m13_prefix_layout.py,m13_semantic_cache.py,m13_routing.py: measured decisions on streaming, retries, cache layout (82 percent input saving from layout alone), semantic caching (not for policy answers without a slot guard), and routing (not worth it for triage at current prices).examples/m13_observability.pyandm13_drift.py: redacted, hashed JSONL traces; a dashboard with p50 and p95 latency, fallback share, and cost per task; and drift checks combining chi-square and PSI.examples/out/m13_dashboard.pngis the chart.examples/m13_deprecations.py,m13_upgrade.py,m13_flags.py: a CI check for shutdown dates, a paired upgrade test, and sticky percentage flags with an automatic rollback guardrail.examples/m13_lab.py: the three-day drill.tests/test_m13_gateway.pyandtests/test_m13_ops.py: 21 tests, all passing, covering retries, fallback, deadlines, quotas, caching, the kill switch, error classification from realopenaiexceptions, memory math, redaction, tracing, PSI, flags, jitter, deprecation chains, and Groq's cache pricing.- An incident runbook (Part D) and a rollback table (Part E) in your team's docs.
Before moving on, run the gateway with a real key (python examples/m13_real_gateway.py), decide which single layer owns retries in your deployment (the SDK's max_retries or the gateway, Part A), and replace the simulated traffic in m13_observability.py with a day of your own traces.
Interview Questions
1. When does self-hosting an open model become cheaper than a hosted API? When the GPUs stay busy. A rented GPU costs the same per hour at 0 percent and 100 percent utilization, while an API charges per token. Compute the fixed monthly cost (GPU hours times replicas, plus engineering time) and divide by the API cost per request to get the break-even volume, then check that your serving engine can sustain the throughput that volume needs at your latency target. For Brightlane (30,000 tickets a month), the API costs about 8 USD a month for gpt-oss-120b while two rented L4s cost about 2,700 USD, and break-even is around 20 million tickets a month. Privacy or residency can justify self-hosting at any volume, but then it should be budgeted as a compliance cost, not sold as a saving.
2. How do you estimate the GPU memory a model needs? Weights plus KV cache plus overhead. Weights are parameters times bytes per parameter (2 for bf16, about 0.5 for 4-bit). KV cache per token is 2 times layers times KV heads times head size times bytes, multiplied by the tokens of every concurrent sequence. For Llama 3.1 8B that is 128 KiB per token, so a single 128k-token conversation needs 16 GiB of cache, more than the 15 GiB of bf16 weights. Grouped-query attention, sliding-window layers, and cache quantization reduce it. What remains after weights and overhead decides how many users fit at once.
3. What does a serving engine like vLLM do that a simple generate loop does not? Continuous batching (new requests join the running batch between decode steps), a paged KV cache (memory handed out in blocks as needed rather than reserved for the longest possible reply), prefix caching (shared prompt prefixes computed once), plus quantized kernels, speculative decoding, and an OpenAI-compatible HTTP API. Together they raise throughput many times over a naive loop at a controlled latency cost. On TinyLM, batching 32 requests raised throughput about 7 times while each request took nearly 6 times longer: batch size is the dial between cost and latency.
4. How would you decide whether 4-bit quantization is acceptable? Measure it on the task. Compare an aggregate metric (perplexity or the eval suite score) and a set of specific outputs that matter, because quantization error is small on average and concentrated in a few places. On TinyLM, int8 halved the size, changed perplexity by 0.02 percent, and changed the top prediction at 0.8 percent of positions while three key replies stayed identical. For a real 4-bit model, run the full regression suite with paired statistics against the unquantized model, read the cases it broke, and check slices such as non-English tickets.
5. Why add jitter to exponential backoff? Because clients that failed together retry together. Without jitter, a burst of 500 requests against a 100-per-second limit produced 12,250 rejected requests in the simulation, and exponential backoff without jitter took six minutes to clear because the crowd stayed synchronized. Full jitter (a random wait between zero and the exponential ceiling) cut rejections by 87 percent and cleared the burst in about 10 seconds. Also honor Retry-After, cap the total wait with a deadline, and retry in exactly one layer: SDK retries multiplied by gateway retries multiply the load.
6. What are the risks of a semantic cache, and how would you evaluate one? It can answer a different question with a confident, wrong, cached reply. Similar wording does not mean the same question: "How much is Team?" and "How much is Business?" are very similar texts. Evaluate with a labeled set of paraphrases (reuse is correct) and near misses (any reuse is wrong), sweep the threshold, and report hit rate and false-hit share. In this module, TF-IDF similarity alone gave 12 false hits out of 15 near misses at the threshold where it caught 11 of 20 paraphrases. A slot guard that requires plan names, numbers, platforms, and negations to match cut false hits to 1 while keeping 8 correct hits. For policy answers, cache retrieval results rather than final replies.
7. Your error rate is flat but users say the assistant got slow and the bill went up. What happened? Probably a provider incident hidden by the fallback chain. Retries and fallbacks keep requests succeeding, so the error rate stays at zero, while latency rises (waits and slower fallback models) and cost rises (the fallback is often more expensive, and retries bill input tokens again). The simulated incident in this module showed exactly that: 0 percent errors, 41 percent fallback, p95 latency doubled, cost per task five-fold. Check fallback share and retries per task on the dashboard, then open the slowest trace. Alert on those metrics, not only on errors.
8. How do you detect that a provider silently changed the model behind the same name? Monitor output distributions against a known-good baseline: reply length, refusal rate, predicted categories, tool-call patterns, and online quality signals, and run a fixed daily sample through the eval suite. Compare with a significance test (chi-square) and an effect size (PSI) together, because at low volume PSI is dominated by noise and at high volume chi-square flags tiny, harmless shifts. Prevent it where you can by pinning dated model ids, and log the exact model id returned in each response so a change is visible in traces.
9. A new model scores 64 percent against the old one's 49 percent on 72 tickets. Is it better? Look at the pairs. Both were scored on the same tickets, so count the tickets where they disagree: here the new one fixed 14 and broke 3. McNemar's exact test gives p = 0.01 and a paired bootstrap interval of +5.6 to +26.4 points, so the improvement is real. An unpaired comparison of the same data would include zero and be inconclusive. Then read the 3 broken tickets, and note that slices with 1 to 3 cases (each non-English language here) cannot support any conclusion.
10. How do you roll out a prompt change safely? Version the prompt, pass the offline regression suite, then ship behind a feature flag with sticky percentage assignment (hash of flag name and user id), starting at a few percent. Decide in advance the guardrail metric (for example the draft rejection rate), the test, and the minimum sample. Ramp up while the guardrail holds, and roll back automatically when it fails; the rollback is a configuration change with a recorded reason. In the simulation, a worse prompt was rolled back after one day at z = 2.97.
11. What is your plan when a provider announces a model shutdown in 30 days? The deprecation check should already have flagged it. Pick candidates (the provider's recommended replacement and at least one alternative, checking that the replacement is not itself scheduled for retirement), run the regression suite with paired statistics, read the broken cases, adjust prompts per model if needed, and ship behind a flag well before the date. Keep a second tested model in the fallback chain, because after the shutdown date there is no rollback to the old model.
12. What should you log for an LLM request, and what should you never log? Always log metadata: timestamp, route, provider, exact model id, token counts, cost, latency, attempts, cache status, prompt version, and flag variant, with user ids as a keyed hash (HMAC) so they can be joined but not reversed. Log content as a redacted excerpt by default, with full text only in a restricted store with short retention. Never log secrets. Redact before writing, at the gateway, because logs and traces are copied to more systems than anyone expects.
Other Tools and Providers
| What this module used | Alternatives | When to consider them |
|---|---|---|
Our m13_gateway.py | LiteLLM (proxy and SDK), Portkey, OpenRouter, Cloudflare AI Gateway, cloud catalogs (Amazon Bedrock, Google's Gemini Enterprise Agent Platform, Microsoft Foundry) | Many providers, many teams, or a need for managed keys, budgets, and logs without writing your own |
| TinyLM batching and prefix reuse | vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama | Any real self-hosted serving |
torch.ao.quantization.quantize_dynamic | torchao, AWQ, GPTQ, bitsandbytes, llama.cpp GGUF quantizations, FP8 in vLLM or TensorRT-LLM | Quantizing real models for serving |
| RunPod and Lambda on-demand GPUs | AWS, Google Cloud, Azure, CoreWeave, reserved or spot capacity | Reserved capacity lowers hourly prices if utilization is high; spot suits batch work |
Our JSONL Tracer | OpenTelemetry (with its generative-AI semantic conventions), Langfuse, Arize Phoenix, LangSmith, Datadog LLM Observability | Traces across services, retention, search, and team dashboards |
| TF-IDF semantic cache | Embedding models with a vector store (Redis, pgvector), GPTCache | Production semantic caching, after the near-miss evaluation |
| PSI and chi-square with SciPy | Evidently, NannyML, whylogs, Kolmogorov-Smirnov tests for numeric features | Scheduled drift reports over many features |
Flag class | LaunchDarkly, Unleash, GrowthBook, Flagsmith, OpenFeature | Flags shared across services with audit trails and a UI |
| Deprecation dict in code | Provider changelogs and deprecation pages, email notices | Always subscribe; the dict is how the notice reaches CI |
Coming Up in Module 14
The assistant is now built, evaluated, secured, adapted, and operated: every call goes through one gateway, costs and latencies are measured rather than guessed, and every change ships behind a test and a flag. Module 14, Applications, Patterns, and Frontiers, steps back from Brightlane's plumbing to ask which LLM applications work at all: the proven patterns (classification at scale, grounded question answering, drafting with review), where humans belong in the loop, when a regex or a classical model is the right answer (you have already seen a logistic regression reach 64 percent on triage), how to judge whether a use case clears the reliability bar, and how to read new model releases critically, so that the system you operate today stays adaptable as the models under it keep changing.