CourseLarge Language Models · Module 13: Deployment, Operations, and Economics · part 69 of 80
Part 69 · Module 13: Deployment, Operations, and Economics

Part D: Observability

19 min read·22 Sept 2026

Part D: Observability

Observability means you can answer questions about the running system from what it records, including questions you did not think of in advance: why was this ticket slow, what did Tuesday's incident cost, did last week's model update change the replies. For an LLM system that takes three kinds of record, and this part builds each:

  • Logs and traces: what happened on each request, step by step, with the content made safe to keep.
  • Metrics: numbers aggregated over time (latency percentiles, error rate, cost per task), which is what a dashboard shows and alerts fire on.
  • Distribution checks: whether the inputs or outputs as a whole are shifting, which no single request reveals.

Logging prompts, responses, and metadata safely

Prompts and replies are the most useful thing to log (you cannot debug a bad draft without seeing it) and the most dangerous (tickets contain emails, phone numbers, card numbers, and whatever else customers paste). Module 11 set the policy; here is the implementation, at the one place every call passes through.

FieldLog it?How
Timestamps, route, provider, model, token counts, cost, latency, attempts, cache statusAlwaysPlain metadata; this is what dashboards are built from
User and customer idsYes, as a keyed hashStable for joining records, meaningless without the key
Prompt and reply textA redacted excerpt by default; full text only in a restricted store with short retentionRedact before writing, never after
Secrets (API keys, tokens, passwords)NeverFilter them at the gateway, and alert if one appears

Two techniques do the work. Redaction replaces sensitive patterns with a label before anything is written. Keyed hashing (HMAC) turns a user id into a stable code using a secret key: the same user always maps to the same code, so you can count requests per user, but nobody without the key can reverse it or rebuild the mapping by hashing a list of known ids, which a plain SHA-256 of the id would allow.

The logging, the tracer, and the dashboard all live in one file, because each builds on the previous one:

examples/m13_observability.py

python
"""Observability for the support assistant: safe logs, traces, and a dashboard computed from them.

Two weeks of traffic are simulated through the Module 13 gateway with fake
providers (ScriptedLLM underneath, so replies are not model output). Latency
and failures are drawn from seeded distributions; on days 9 and 10 the primary
provider rate-limits hard, the way a real incident would look. Everything
below the simulation (redaction, tracing, the dashboard math) is real code you
can point at real logs.
"""
from __future__ import annotations

import contextvars
import hashlib
import hmac
import json
import os
import random
import re
import statistics
from collections import defaultdict
from contextlib import contextmanager
from dataclasses import dataclass, field
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Callable

from m13_gateway import FakeClock, FakeProvider, Gateway, Quota, Route, ticket_messages
from supportdesk.data import load_tickets
from supportdesk.kb_search import KBSearch

OUT = Path(__file__).resolve().parent / "out"
START = datetime(2026, 9, 1, tzinfo=timezone.utc).timestamp()

# Safe logging ----------------------------------------------------------------------

REDACTIONS = [
    (re.compile(r"[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}"), "[EMAIL]"),
    (re.compile(r"\b[A-Z]{2}\d{2}(?: ?[A-Z0-9]{4}){3,7}(?: ?[A-Z0-9]{1,3})?\b"), "[IBAN]"),
    (re.compile(r"\b(?:\d[ -]?){13,19}\b"), "[CARD]"),
    (re.compile(r"\+\d[\d ()-]{7,}\d"), "[PHONE]"),
]


def redact(text: str) -> str:
    for pattern, label in REDACTIONS:
        text = pattern.sub(label, text)
    return text


def hash_user(user_id: str) -> str:
    """Keyed hash: stable for joining logs, useless to anyone without the key (unlike plain sha256)."""
    key = os.environ.get("LOG_HASH_KEY", "dev-only-key-change-me").encode()
    return hmac.new(key, user_id.encode(), hashlib.sha256).hexdigest()[:16]


# Tracing -----------------------------------------------------------------------------

@dataclass
class Span:
    trace_id: str
    span_id: str
    parent_id: str | None
    name: str
    start: float
    attrs: dict[str, Any] = field(default_factory=dict)


_current: contextvars.ContextVar[Span | None] = contextvars.ContextVar("current_span", default=None)


class Tracer:
    """Writes one JSON line per finished span. Child spans inherit the trace id of their parent."""

    def __init__(self, path: Path, clock: Callable[[], float], seed: int = 0) -> None:
        self.path, self.clock, self.rng = path, clock, random.Random(seed)
        path.parent.mkdir(parents=True, exist_ok=True)
        self.file = path.open("w", encoding="utf-8")

    def _id(self, bits: int) -> str:
        return f"{self.rng.getrandbits(bits):0{bits // 4}x}"

    @contextmanager
    def span(self, name: str, **attrs: Any):
        parent = _current.get()
        span = Span(parent.trace_id if parent else self._id(128), self._id(64),
                    parent.span_id if parent else None, name, self.clock(), dict(attrs))
        token = _current.set(span)
        status = "ok"
        try:
            yield span
        except Exception as exc:
            status = "error"
            span.attrs["error"] = type(exc).__name__
            raise
        finally:
            _current.reset(token)
            record = {"trace_id": span.trace_id, "span_id": span.span_id, "parent_id": span.parent_id,
                      "name": name, "start": round(span.start, 3),
                      "duration_ms": round((self.clock() - span.start) * 1000, 1), "status": status, **span.attrs}
            self.file.write(json.dumps(record) + "\n")

    def close(self) -> None:
        self.file.close()


# Simulated traffic ---------------------------------------------------------------------

def simulate(path: Path, days: int = 14, per_day: int = 150, seed: int = 7) -> None:
    rng = random.Random(seed)
    clock = FakeClock(START)
    day = {"n": 0}

    def groq_outcome() -> str:
        if day["n"] in (9, 10) and rng.random() < 0.35:
            return "429:20"                     # incident: long Retry-After, so the gateway falls back
        return "500" if rng.random() < 0.01 else "ok"

    groq = FakeProvider("groq", clock=clock, latency_fn=lambda: rng.lognormvariate(6.0, 0.35), outcome_fn=groq_outcome)
    gemini = FakeProvider("gemini", clock=clock, latency_fn=lambda: rng.lognormvariate(6.8, 0.30),
                          outcome_fn=lambda: "500" if rng.random() < 0.02 else "ok")
    routes = {name: [Route("groq", groq, "openai/gpt-oss-120b"), Route("gemini", gemini, "gemini-3.5-flash")]
              for name in ("triage", "draft")}
    gw = Gateway(routes, quota=Quota(usd_per_day=1.0, requests_per_minute=30), clock=clock, sleep=clock.sleep, seed=seed)
    tracer = Tracer(path, clock, seed)
    search, tickets = KBSearch(), load_tickets()

    for d in range(1, days + 1):
        day["n"] = d
        for i in range(per_day):
            clock.now = START + (d - 1) * 86_400 + i * (86_400 / per_day)
            t = rng.choice(tickets)
            user = f"cust-{rng.randint(1, 400)}"
            body = t.body + (" Reach me at dana.lee@example.com or +49 30 1234 5678." if rng.random() < 0.2 else "")
            body += f" (ref {d}-{i})"          # real tickets are unique, so the exact cache rarely hits
            try:
                with tracer.span("handle_ticket", user=hash_user(user), ticket=t.id, day=d) as root:
                    with tracer.span("triage") as s:
                        r = gw.chat(ticket_messages(t.subject, body), route="triage", user=user)
                        s.attrs.update(provider=r.provider, cache=r.cache, usd=r.cost_usd, attempts=len(r.attempts))
                    with tracer.span("kb_search") as s:
                        hits = search.search(t.text, k=3)
                        clock.now += 0.004
                        s.attrs["top"] = hits[0].article_id if hits else None
                    with tracer.span("draft") as s:
                        r = gw.chat(ticket_messages(t.subject, body), route="draft", user=user)
                        s.attrs.update(provider=r.provider, cache=r.cache, usd=r.cost_usd, attempts=len(r.attempts),
                                       excerpt=redact(body)[:80])
                    root.attrs["category"] = t.gold["category"]
            except Exception:  # noqa: BLE001  (the span already recorded the error)
                pass
    tracer.close()


# Dashboard -----------------------------------------------------------------------------

def percentile(values: list[float], q: float) -> float:
    ordered = sorted(values)
    return ordered[min(int(q * len(ordered)), len(ordered) - 1)]


def dashboard(path: Path) -> dict[int, dict[str, float]]:
    spans = [json.loads(line) for line in path.open(encoding="utf-8")]
    by_trace: dict[str, list[dict]] = defaultdict(list)
    for s in spans:
        by_trace[s["trace_id"]].append(s)
    rows: dict[int, dict[str, list]] = defaultdict(lambda: defaultdict(list))
    for trace in by_trace.values():
        root = next(s for s in trace if s["parent_id"] is None)
        row = rows[root["day"]]
        row["latency"].append(root["duration_ms"])
        row["error"].append(root["status"] != "ok")
        row["usd"].append(sum(s.get("usd", 0.0) for s in trace))
        draft = next((s for s in trace if s["name"] == "draft" and s["status"] == "ok"), None)
        row["fallback"].append(bool(draft) and draft["provider"] != "groq")
        row["retries"].append(sum(s.get("attempts", 1) - 1 for s in trace if "attempts" in s))
    summary = {}
    print(f"{'day':>3} {'tasks':>6} {'p50 ms':>7} {'p95 ms':>7} {'errors':>7} {'fallback':>9} {'retries':>8} {'USD per 1k tasks':>17}")
    for d in sorted(rows):
        r = rows[d]
        ok_usd = [u for u, e in zip(r["usd"], r["error"]) if not e]
        summary[d] = {"p50": statistics.median(r["latency"]), "p95": percentile(r["latency"], 0.95),
                      "error_rate": sum(r["error"]) / len(r["error"]), "fallback": sum(r["fallback"]) / len(r["fallback"]),
                      "retries": statistics.mean(r["retries"]), "usd_per_1k": 1000 * sum(ok_usd) / max(len(ok_usd), 1)}
        s = summary[d]
        print(f"{d:>3} {len(r['latency']):>6} {s['p50']:>7.0f} {s['p95']:>7.0f} {s['error_rate']:>7.1%} "
              f"{s['fallback']:>9.0%} {s['retries']:>8.2f} {s['usd_per_1k']:>17.3f}")
    return summary


def show_slowest_trace(path: Path, day: int) -> None:
    spans = [json.loads(line) for line in path.open(encoding="utf-8")]
    roots = [s for s in spans if s["parent_id"] is None and s["day"] == day]
    slow = max(roots, key=lambda s: s["duration_ms"])
    print(f"\nslowest trace on day {day}: {slow['trace_id'][:12]}... {slow['duration_ms']:.0f} ms, user {slow['user']}")
    for s in sorted((s for s in spans if s["trace_id"] == slow["trace_id"]), key=lambda s: (s["start"], s["parent_id"] is not None)):
        detail = {k: s[k] for k in ("provider", "attempts", "cache", "usd", "excerpt") if k in s}
        indent = "  " if s["parent_id"] is None else "    "
        print(f"{indent}{s['name']:<14}{s['duration_ms']:>8.0f} ms  {s['status']:<5} {detail}")


def chart(summary: dict[int, dict[str, float]], path: Path) -> None:
    """Two panels, one measure each (never two y-scales on one axis); the incident days are shaded."""
    import matplotlib
    matplotlib.use("Agg")
    import matplotlib.pyplot as plt

    ink, muted, surface = "#0b0b0b", "#52514e", "#fcfcfb"
    days = sorted(summary)
    fig, (a, b) = plt.subplots(1, 2, figsize=(10, 3.4), facecolor=surface)
    a.plot(days, [summary[d]["p50"] for d in days], color="#2a78d6", lw=2, marker="o", ms=5, label="p50")
    a.plot(days, [summary[d]["p95"] for d in days], color="#eb6834", lw=2, marker="o", ms=5, label="p95")
    a.set_title("Task latency (ms)", color=ink, loc="left")
    a.legend(frameon=False, labelcolor=muted, loc="upper left")
    b.bar(days, [summary[d]["usd_per_1k"] for d in days], color="#2a78d6", width=0.7)
    b.set_title("Cost per 1,000 tasks (USD)", color=ink, loc="left")
    for ax in (a, b):
        ax.set_facecolor(surface)
        ax.axvspan(8.5, 10.5, color="#f0efec", zorder=0)
        low, high = ax.get_ylim()
        ax.set_ylim(low, high + (high - low) * 0.12)                 # headroom so the label clears the data
        ax.text(9.5, high + (high - low) * 0.10, "incident", ha="center", va="top", color=muted, fontsize=8)
        ax.set_xlabel("day of September", color=muted)
        ax.tick_params(colors=muted)
        ax.grid(axis="y", color="#e5e4e0", lw=0.8)
        ax.set_axisbelow(True)
        for side in ("top", "right"):
            ax.spines[side].set_visible(False)
        for side in ("left", "bottom"):
            ax.spines[side].set_color("#c3c2b7")
    fig.tight_layout()
    fig.savefig(path, dpi=120, facecolor=surface)


if __name__ == "__main__":
    print(redact("Card 4111 1111 1111 1111, IBAN DE89 3704 0044 0532 0130 00, mail dana.lee@example.com, "
                 "call +49 30 1234 5678, invoice INV-2026-004512"))
    print("user cust-17 ->", hash_user("cust-17"))
    log = OUT / "m13_traces.jsonl"
    simulate(log)
    print(f"\nwrote {sum(1 for _ in log.open())} spans to {log.relative_to(Path.cwd())}\n")
    summary = dashboard(log)
    show_slowest_trace(log, 9)
    chart(summary, OUT / "m13_dashboard.png")
    print("\nchart saved to examples/out/m13_dashboard.png")

Code explained

  • In simple words: a black-box flight recorder for the assistant that blanks out personal details, plus the instrument panel computed from what it recorded.
  • What happens:
    • REDACTIONS and redact(text): regexes for emails, IBANs, card numbers (13 to 19 digits with optional separators), and international phone numbers, applied in that order. Invoice numbers such as INV-2026-004512 are deliberately left alone: agents need them, and they are not personal data on their own.
    • hash_user(user_id): HMAC-SHA256 with a key from the LOG_HASH_KEY environment variable (the fallback key is for local runs only), shortened to 16 hex characters.
    • Span and Tracer: a minimal tracer. A trace is the record of one task end to end; a span is one step inside it, with a start time, a duration, a status, a parent, and attributes. Tracer.span is a context manager: it creates a span, makes it the current span (a contextvars variable, so nested calls find their parent without passing it around), and writes one JSON line when the step finishes, marking it error if an exception escaped. This is the same model as OpenTelemetry, cut down to 40 lines.
    • simulate: 14 days of traffic, 150 tickets a day, drawn from the real tickets, through the Part A gateway with fake providers. Latencies come from seeded log-normal distributions; on days 9 and 10 the primary provider answers 35 percent of calls with a 429 and Retry-After: 20, an incident planted on purpose. One ticket in five gets an email address and phone number appended, so redaction has something to do. Each ticket becomes a root span handle_ticket with three children: triage, kb_search, and draft.
    • percentile and dashboard: group spans by trace, then by day, and compute p50 and p95 latency per task, error rate, fallback share, retries per task, and cost per 1,000 tasks, all from the JSONL file alone.
    • show_slowest_trace: the drill-down: find the slowest task on a bad day and print its spans as a tree.
    • chart: two panels (latency, cost), one measure each, with the incident days shaded. Matplotlib 3.11.2.
  • Comes out:

    text
    Card [CARD], IBAN [IBAN], mail [EMAIL], call [PHONE], invoice INV-2026-004512
    user cust-17 -> 247cdda29fa7cd3b
    
    wrote 8400 spans to examples/out/m13_traces.jsonl
    
    day  tasks  p50 ms  p95 ms  errors  fallback  retries  USD per 1k tasks
      1    150     794    1256    0.0%        0%     0.03             0.040
      2    150     836    1359    0.0%        0%     0.03             0.040
      3    150     849    1236    0.0%        0%     0.02             0.040
      4    150     844    1229    0.0%        0%     0.01             0.040
      5    150     844    1202    0.0%        0%     0.01             0.040
      6    150     854    1375    0.0%        0%     0.03             0.040
      7    150     850    1300    0.0%        0%     0.01             0.040
      8    150     835    1304    0.0%        0%     0.01             0.040
      9    150    1533    2938    0.0%       41%     0.81             0.206
     10    150    1345    2799    0.0%       36%     0.65             0.177
     11    150     846    1264    0.0%        0%     0.02             0.040
     12    150     851    1339    0.0%        0%     0.03             0.040
     13    150     851    1618    0.0%        0%     0.05             0.039
     14    150     876    1351    0.0%        0%     0.03             0.040
    
    slowest trace on day 9: dbba1e9341ad... 3625 ms, user db6882609bad0629
      handle_ticket     3625 ms  ok    {}
        triage            2302 ms  ok    {'provider': 'gemini', 'attempts': 3, 'cache': 'miss', 'usd': 0.000237}
        kb_search            4 ms  ok    {}
        draft             1320 ms  ok    {'provider': 'gemini', 'attempts': 2, 'cache': 'miss', 'usd': 0.000237, 'excerpt': 'Can I use the iPhone app on a plane without internet? Reach me at [EMAIL] or [PH'}
    
    chart saved to examples/out/m13_dashboard.png
    

    The first line is redaction working on a hostile sample. The replies in these traces are ScriptedLLM text, not model output, and the latencies are simulated; the logging, tracing, and dashboard code is real and would run unchanged on real spans.

Tracing multi-step and agentic flows

A single request log cannot explain the slowest ticket on day 9. The trace can: the task took 3.6 seconds, and 2.3 of those were triage, which needed 3 attempts before it landed on Gemini, then the draft needed 2 more. Nothing was broken in the code; the primary provider was rate-limiting, and the fallback did its job at the cost of latency. You find that in one read because every span carries its provider, attempts, cost, and a redacted excerpt, and because children point at their parent.

For agents (Module 8) the same structure scales: each model call, tool call, and approval wait is a span under the task's root, so a loop that ran 14 steps instead of 4 is visible as 14 children, each with its own cost. Three rules keep traces useful:

  1. One trace id per user-visible task, created at the entry point and passed (by context, as here, or in a header between services) to everything the task does.
  2. Record decisions, not only timings: which route, which provider, cache hit or miss, which prompt version, which flag variant. These are the attributes you will filter on during an incident.
  3. Redact at the span, before writing. Traces are copied to more places than logs are.

Quality, latency, and spend dashboards

Here is the chart the script saved, the kind of view the support lead should see every morning:

.

Read the table and chart together. On days 9 and 10, p95 latency more than doubled (1,300 to about 2,900 ms), 41 and 36 percent of tasks fell back to Gemini, retries per task rose from about 0.02 to 0.8, and cost per task rose five-fold (0.040 to 0.206 USD per 1,000 tasks) because the fallback model is more expensive. The error rate stayed at 0.0 percent. That last point is the lesson: a good fallback chain hides incidents from your error rate. If you only alert on errors, this incident is invisible. Alert on fallback share and cost per task as well.

What a dashboard for an LLM feature should show, and why:

PanelMetricWhat it catches
Latencyp50 and p95 per task, and TTFT for streamed repliesProvider slowdowns, retry storms, longer outputs after a prompt change
ReliabilityError rate, fallback share, retries per taskOutages that fallback hides
SpendCost per task, cost per day, top users by costPrice changes, routing drift, denial of wallet
QualityOnline signals from Module 10 (draft acceptance, edit distance, escalations) and a daily sample through the eval suiteSilent regressions no latency or error metric shows
TrafficVolume and category mixDrift in what users ask (next section)

Use percentiles, not averages, for latency: an average hides the slow tail that users remember. And always divide cost by tasks, not by calls: a retry or fallback adds calls to the same task, and cost per call would hide it.

Drift detection and silent regression

Drift is a change in the distribution of what goes in (the ticket mix, languages, lengths) or what comes out (reply length, refusal rate, category predictions). Input drift can make a well-tested system meet traffic it was never tested on; output drift with unchanged inputs usually means the model or prompt changed, sometimes silently when a provider updates a model behind the same name. Neither shows up in a single trace.

Two standard measures compare this period with a baseline, bin by bin:

  • Population Stability Index (PSI): the sum over bins of (actual share - expected share) x ln(actual share / expected share). It measures how big the shift is. A common rule of thumb from credit scoring reads under 0.1 as stable, 0.1 to 0.25 as worth watching, and above 0.25 as a real shift.
  • Chi-square test: asks whether the two sets of counts could plausibly come from the same distribution, and returns a p-value. It measures how sure you can be that there is any shift at all.

The script samples weekly ticket mixes from the real category proportions of the 72 tickets, plants shifts of known size, and checks what each measure finds. Because we planted the shifts, we know the right answer.

examples/m13_drift.py

python
"""Drift detection: has this week's traffic (or this week's model output) moved away from the baseline?

Weekly ticket mixes are sampled from the real category proportions of the
72-ticket dataset (seeded), with a synthetic shift injected on purpose, so we
know the right answer and can check that the detectors find it.
"""
from __future__ import annotations

import math
import random
from collections import Counter

from scipy.stats import chi2_contingency

from supportdesk.data import CATEGORIES, load_tickets


def psi(expected: Counter, actual: Counter, keys, floor: float = 1e-4) -> float:
    """Population Stability Index: sum over bins of (a - e) * ln(a / e), using proportions."""
    e_total, a_total = sum(expected.values()), sum(actual.values())
    total = 0.0
    for k in keys:
        e = max(expected[k] / e_total, floor)
        a = max(actual[k] / a_total, floor)
        total += (a - e) * math.log(a / e)
    return total


def chi_square_p(expected: Counter, actual: Counter, keys) -> float:
    """p-value of a chi-square test that both weeks come from the same category distribution."""
    table = [[expected[k] for k in keys], [actual[k] for k in keys]]
    return chi2_contingency(table)[1]


def sample_week(weights: dict[str, float], n: int, rng: random.Random) -> Counter:
    return Counter(rng.choices(list(weights), weights=list(weights.values()), k=n))


def verdict(value: float) -> str:
    return "stable" if value < 0.1 else ("watch" if value < 0.25 else "ACT")


def main() -> None:
    base = Counter(t.gold["category"] for t in load_tickets())
    weights = {c: base[c] / sum(base.values()) for c in CATEGORIES}
    rng = random.Random(13)
    weeks = {"week 36 (baseline)": sample_week(weights, 1000, rng)}
    weeks["week 37 (no change)"] = sample_week(weights, 1000, rng)
    small = dict(weights, bug=weights["bug"] * 1.3)
    weeks["week 38 (bug +30%)"] = sample_week(small, 1000, rng)
    release = dict(weights, bug=weights["bug"] * 3, account_access=weights["account_access"] * 1.5)
    weeks["week 39 (bad release)"] = sample_week(release, 1000, rng)
    baseline = weeks["week 36 (baseline)"]

    print("share of tickets per category")
    print(f"{'week':<22}" + "".join(f"{c[:10]:>11}" for c in CATEGORIES))
    for name, counts in weeks.items():
        print(f"{name:<22}" + "".join(f"{counts[c] / sum(counts.values()):>11.1%}" for c in CATEGORIES))

    print(f"\n{'week vs baseline':<22}{'PSI':>7}{'verdict':>9}{'chi-square p':>14}")
    for name, counts in list(weeks.items())[1:]:
        value = psi(baseline, counts, CATEGORIES)
        print(f"{name:<22}{value:>7.3f}{verdict(value):>9}{chi_square_p(baseline, counts, CATEGORIES):>14.2g}")

    print("\nSame shift (bug +30%), different weekly volumes, 200 simulated weeks each:")
    for n in (100, 300, 1000, 3000):
        hits = Counter()
        for _ in range(200):
            a, b, same = sample_week(weights, n, rng), sample_week(small, n, rng), sample_week(weights, n, rng)
            hits["psi"] += psi(a, b, CATEGORIES) >= 0.1
            hits["psi_false"] += psi(a, same, CATEGORIES) >= 0.1
            hits["chi"] += chi_square_p(a, b, CATEGORIES) < 0.01
            hits["chi_false"] += chi_square_p(a, same, CATEGORIES) < 0.01
        print(f"  n={n:>5}/week: PSI>=0.1 flags {hits['psi'] / 2:>5.1f}% (false alarms {hits['psi_false'] / 2:>5.1f}%)"
              f"   chi-square p<0.01 flags {hits['chi'] / 2:>5.1f}% (false alarms {hits['chi_false'] / 2:.1f}%)")

    print("\nOutput drift: reply length in words, same inputs, model version changed silently")
    bins = ["<40", "40-79", "80-119", "120+"]

    def binned(lengths):
        return Counter(bins[min(int(x) // 40, 3)] for x in lengths)

    before = [rng.gauss(70, 18) for _ in range(800)]
    after = [rng.gauss(70 * 1.35, 25) for _ in range(800)]
    b, a = binned(before), binned(after)
    print("  bins:   " + "".join(f"{x:>8}" for x in bins))
    print("  before: " + "".join(f"{b[x]:>8}" for x in bins))
    print("  after:  " + "".join(f"{a[x]:>8}" for x in bins))
    value = psi(b, a, bins)
    print(f"  PSI {value:.3f} ({verdict(value)}), chi-square p {chi_square_p(b, a, bins):.2g}")


if __name__ == "__main__":
    main()

Code explained

  • In simple words: compare this week's ticket mix with a normal week, with two different rulers, on data where we secretly know what changed.
  • What happens:
    • psi(expected, actual, keys): the formula above, with a small floor so an empty bin does not produce a logarithm of zero.
    • chi_square_p: SciPy's chi2_contingency on a two-row table of counts (baseline week, this week).
    • sample_week: draws n tickets from category weights with a seeded random generator.
    • verdict: the PSI rule of thumb.
    • main: four weeks (baseline, unchanged, bug tickets up 30 percent, and a bad release that triples bug tickets and adds half again to login problems); then 200 simulated weeks per volume to measure how often each measure flags the small shift and how often it raises a false alarm on an unchanged week; then output drift on reply lengths after a silent model change.
  • Comes out:

    text
    share of tickets per category
    week                      billing cancellati account_ac        bug     how_to feature_re
    week 36 (baseline)          23.4%       8.4%      24.0%      12.2%      25.2%       6.8%
    week 37 (no change)         20.9%       7.8%      23.7%      13.1%      26.3%       8.2%
    week 38 (bug +30%)          22.5%       7.0%      22.6%      16.3%      23.8%       7.8%
    week 39 (bad release)       16.8%       6.4%      28.1%      25.8%      16.4%       6.5%
    
    week vs baseline          PSI  verdict  chi-square p
    week 37 (no change)     0.007   stable          0.62
    week 38 (bug +30%)      0.018   stable          0.12
    week 39 (bad release)   0.174    watch       1.2e-16
    
    Same shift (bug +30%), different weekly volumes, 200 simulated weeks each:
      n=  100/week: PSI>=0.1 flags  44.0% (false alarms  41.5%)   chi-square p<0.01 flags   2.0% (false alarms 1.0%)
      n=  300/week: PSI>=0.1 flags   3.5% (false alarms   1.0%)   chi-square p<0.01 flags   2.0% (false alarms 0.5%)
      n= 1000/week: PSI>=0.1 flags   0.0% (false alarms   0.0%)   chi-square p<0.01 flags   9.5% (false alarms 0.0%)
      n= 3000/week: PSI>=0.1 flags   0.0% (false alarms   0.0%)   chi-square p<0.01 flags  62.0% (false alarms 0.5%)
    
    Output drift: reply length in words, same inputs, model version changed silently
      bins:        <40   40-79  80-119    120+
      before:       27     550     220       3
      after:        15     195     458     132
      PSI 1.297 (ACT), chi-square p 6.4e-82
    

    The two measures disagree in instructive ways. The bad release (bug share from 12 to 26 percent) gets a PSI of only 0.174, "watch", yet a chi-square p-value of 1e-16: the shift is certain and moderate in size. The 30 percent rise in bug tickets is real but small, and at 1,000 tickets a week neither measure catches it reliably.

    The volume experiment shows why you need both. At 100 tickets a week, PSI flags 44 percent of weeks with the shift and 41.5 percent of weeks without it: at small volumes PSI mostly measures sampling noise, so a fixed 0.1 threshold is meaningless. Chi-square keeps its false alarms near the 1 percent it promises at every volume, and its power grows with volume (62 percent at 3,000 a week). A good alert combines them: the shift must be statistically real (chi-square p under 0.01) and large enough to matter (PSI above a threshold you calibrate on your own stable weeks).

    Output drift is the silent-regression case. Same inputs, but replies got about 35 percent longer after an unannounced model change: PSI 1.3, p about 1e-82. No error rate, latency, or cost panel would flag this quickly (cost would creep up by the extra output tokens), and a longer draft is not necessarily a worse one. The alert's job is to make a person look, which is where the runbook comes in.

SituationUse thisWhy
Is there any shift at all, at high volume?Chi-square (or a two-sample test for numeric values)Controls false alarms at every volume
Is the shift big enough to act on?PSI, with a threshold calibrated on stable weeksMeasures size, not certainty
Low volume (under a few hundred per period)Longer periods, or chi-square onlyPSI is dominated by noise
Output behavior (length, refusals, categories, tool use)Both, against a baseline taken from a known-good weekCatches silent model updates

Incident response for model behavior changes

An LLM incident is often not an outage. It is "the drafts started quoting a 14-day refund window", "replies got long and chatty on Tuesday", or "the bill tripled overnight". The runbook below is what Brightlane's on-call engineer follows. Each step maps to a tool built in this module.

StepActionTool
1. DetectAn alert fires on fallback share, cost per task, drift, or quality signals, or an agent reports a bad draftDashboard, drift check, Module 10 online signals
2. ContainIf users could be harmed (wrong policy, leaked data, unsafe action), flip the kill switch for the route; tickets go to humans. Otherwise drain the bad provider or roll back the flagGateway.kill, disable_provider, Flag.set_percent(0, ...)
3. ScopeFind the first bad trace; filter spans by route, provider, model, prompt version, and flag variant to see what changed and who was affectedJSONL traces, show_slowest_trace-style drill-downs
4. DiagnoseReplay the affected inputs against the current and last-known-good configuration; run the regression suitePart E's paired upgrade test
5. Fix and verifyPin the model version, revert the prompt, or correct the KB; rerun the suite; ramp back up behind a flagPart E's flags
6. LearnWrite a short blameless review: timeline, impact (tickets, cost), what detected it, and one new test or alert that would have caught it earlierAdd the failing cases to the golden set (Module 10)

Two rules make the runbook work under pressure. Containment comes before diagnosis: flip the switch first, investigate with the pressure off. And every containment action must be reversible and logged, which is why the kill switch carries a reason string and the flag keeps its history.