Part G: Operating evals
CI gates on quality, cost, and latency
A CI gate is a test that runs on every change (prompt, model, code, retrieval settings) and blocks the merge when quality, cost, or latency gets worse than an agreed budget. It turns everything above into a habit. Three files make it work:
evals/baseline.jsonpins the currently accepted system: its golden fingerprint, pass rate, and the ids of passing cases.evals/gate.jsonholds the thresholds, reviewed like code.tests/test_m10_gate.pyis an ordinary pytest file, so any CI system can run it.
{
"max_pass_rate_drop": 0.02,
"max_new_failures": 1,
"critical_tags": ["unanswerable", "adversarial"],
"max_forbidden_failures": 0,
"max_escalation_rate": 0.55,
"max_cost_usd_per_case": 0.002,
"max_latency_ms_p95": 8000
}
Code explained
- In simple words: the budgets a change must stay within to merge.
- What happens: the pass rate may drop at most 2 points below the pinned baseline (about two cases; smaller than the noise floor of a stochastic system, which is fine for these deterministic ones and must be revisited for a real model). At most one previously passing case may start failing, and none tagged
unanswerableoradversarial, because those are where failures hurt customers. Zero forbidden-content failures. Escalation rate at most 55%, so a system cannot pass by sending everything to humans. Cost per case at most 0.002 USD and p95 latency at most 8 seconds, sized for a two-call hosted pipeline. - Comes out: read by the gate test and the lab.
"""Module 10 quality gate: run in CI on every prompt, model, or code change.
EVAL_SYSTEM=baseline_v2 PYTHONPATH=. pytest -q tests/test_m10_gate.py
Compares the candidate against the pinned baseline (evals/baseline.json) with the
thresholds in evals/gate.json. Thresholds were set from this module's small,
deterministic systems; recalibrate them for a real model (see Part G).
"""
import json
import os
import pytest
from examples.m10_evals import EVALS, REGRESSIONS, SYSTEMS, dataset_fingerprint, load_cases, run_eval
SYSTEM = os.environ.get("EVAL_SYSTEM", "baseline_v2")
GATE = json.loads((EVALS / "gate.json").read_text())
BASELINE = json.loads((EVALS / "baseline.json").read_text())
@pytest.fixture(scope="module")
def run():
return run_eval(SYSTEMS[SYSTEM](0), load_cases(), f"gate-{SYSTEM}")
def test_same_golden_dataset():
assert dataset_fingerprint() == BASELINE["dataset"], "golden set changed: re-pin the baseline first"
def test_no_forbidden_content(run):
assert run["summary"]["forbidden_failures"] <= GATE["max_forbidden_failures"]
def test_pass_rate_not_below_baseline(run):
floor = BASELINE["pass_rate"] - GATE["max_pass_rate_drop"]
assert run["summary"]["pass_rate"] >= floor, f"{run['summary']['pass_rate']:.3f} < floor {floor:.3f}"
def test_no_new_failures_beyond_budget(run):
passed_now = {r["id"] for r in run["results"] if r["passed"]}
newly_failing = sorted(set(BASELINE["passed_ids"]) - passed_now)
critical = [r["id"] for r in run["results"]
if r["id"] in newly_failing and set(r["tags"]) & set(GATE["critical_tags"])]
assert not critical, f"newly failing critical cases: {critical}"
assert len(newly_failing) <= GATE["max_new_failures"], f"newly failing: {newly_failing}"
def test_regression_suite_all_pass():
cases = load_cases(REGRESSIONS)
result = run_eval(SYSTEMS[SYSTEM](0), cases, f"gate-{SYSTEM}-regressions", save=False)
failing = [r["id"] for r in result["results"] if not r["passed"]]
assert not failing, f"regression cases failing again: {failing}"
def test_escalation_rate(run):
assert run["summary"]["escalation_rate"] <= GATE["max_escalation_rate"]
def test_cost_per_case(run):
assert run["summary"]["cost_usd_per_case"] <= GATE["max_cost_usd_per_case"]
def test_latency_p95(run):
assert run["summary"]["latency_ms_p95"] <= GATE["max_latency_ms_p95"]
Code explained
- In simple words: eight assertions that together say "this change is no worse than what we already ship".
- What happens:
EVAL_SYSTEMpicks the candidate. The module-scopedrunfixture evaluates it once on the golden set. The tests check that the golden set has not changed since the baseline was pinned, that there is no forbidden content, that the pass rate is above the floor, that no critical case newly fails and at most one other does, that the regression suite passes completely, and that escalation rate, cost per case, and p95 latency are within budget. Each assertion message names what failed. - Comes out: below, first failing and then passing.
Pin v2 as the accepted system:
python examples/m10_evals.py baseline runs/baseline_v2.json
Code explained
- In simple words: record v2's results as the bar future changes must clear.
- What happens: the
baselinecommand writesevals/baseline.jsonwith the system name, golden fingerprint, pass rate, and sorted passing ids. Commit this file; update it only on purpose, in a reviewed change. - Comes out:text
evals/baseline.json now pins baseline_v2 at 60.4% (55 passing cases)
Now a realistic change request. Agents complain that 47% of tickets land in their queue as "needs human". A developer lowers the "unsure" retrieval threshold from 5.0 to 4.0 so fewer tickets escalate. That is baseline_v3. It looks harmless:
python examples/m10_evals.py run --system baseline_v3
python examples/m10_evals.py compare runs/baseline_v2.json runs/baseline_v3.json
EVAL_SYSTEM=baseline_v3 pytest -q tests/test_m10_gate.py
Code explained
- In simple words: score the change, compare it with the pinned system, and run the gate on it.
- What happens: the run and compare work as before. The pytest run executes the eight gate tests with v3 as the candidate.
- Comes out:text
system=baseline_v3 dataset=da5ec330c214 n=91 pass rate 57.1% (95% CI 46.9% to 66.8%) forbidden failures=0 check passed / applicable schema_valid 91 / 91 100.0% category 69 / 91 75.8% cites_expected_article 58 / 74 78.4% states_key_fact 57 / 74 77.0% reply_language 77 / 91 84.6% no_roadmap_date 91 / 91 100.0% no_unauthorized_action 91 / 91 100.0% escalates_unanswerable 11 / 17 64.7% slice passed / cases non_english 0 / 14 0.0% ambiguous 1 / 5 20.0% adversarial 4 / 5 80.0% unanswerable 7 / 16 43.8% escalation rate 39.6% latency p50=0.44 ms p95=0.68 ms cost/case=$0.000000 per-case results: runs/baseline_v3.json baseline_v2 60.4% vs baseline_v3 57.1% on n=91 paired cases only baseline_v2 passes: 3 ['H-17', 'H-18', 'T-1039'] only baseline_v3 passes: 0 [] difference -0.033 (paired bootstrap 95% CI -0.077 to +0.000) McNemar exact p=0.2500 dev baseline_v2 33/48 baseline_v3 33/48 discordant 0/0 McNemar p=1.000 test baseline_v2 13/24 baseline_v3 12/24 discordant 1/0 McNemar p=1.000 curated baseline_v2 9/19 baseline_v3 7/19 discordant 2/0 McNemar p=0.500 ..FF.... [100%] =================================== FAILURES =================================== ______________________ test_pass_rate_not_below_baseline _______________________ run = {'system': 'gate-baseline_v3', 'dataset': 'da5ec330c214', 'created': '2026-09-21T12:00:48', 'summary': {'n': 91, 'passed': 52, 'pass_rate': 0.5714, 'ci95': [0.4689, 0.6682], ...}, ...} def test_pass_rate_not_below_baseline(run): floor = BASELINE["pass_rate"] - GATE["max_pass_rate_drop"] > assert run["summary"]["pass_rate"] >= floor, f"{run['summary']['pass_rate']:.3f} < floor {floor:.3f}" E AssertionError: 0.571 < floor 0.584 E assert 0.5714 >= 0.5844 tests/test_m10_gate.py:36: AssertionError ______________________ test_no_new_failures_beyond_budget ______________________ run = {'system': 'gate-baseline_v3', 'dataset': 'da5ec330c214', 'created': '2026-09-21T12:00:48', 'summary': {'n': 91, 'passed': 52, 'pass_rate': 0.5714, 'ci95': [0.4689, 0.6682], ...}, ...} def test_no_new_failures_beyond_budget(run): passed_now = {r["id"] for r in run["results"] if r["passed"]} newly_failing = sorted(set(BASELINE["passed_ids"]) - passed_now) critical = [r["id"] for r in run["results"] if r["id"] in newly_failing and set(r["tags"]) & set(GATE["critical_tags"])] > assert not critical, f"newly failing critical cases: {critical}" E AssertionError: newly failing critical cases: ['T-1039', 'H-17', 'H-18'] E assert not ['T-1039', 'H-17', 'H-18'] tests/test_m10_gate.py:44: AssertionError =========================== short test summary info ============================ FAILED tests/test_m10_gate.py::test_pass_rate_not_below_baseline - AssertionE... FAILED tests/test_m10_gate.py::test_no_new_failures_beyond_budget - Assertion... 2 failed, 6 passed in 0.46sThe escalation rate did fall, from 47.2% to 39.6%. The paired statistics alone would not block it: 3 discordant cases give McNemar p = 0.25, "no significant difference". The gate blocks it anyway, and rightly: the three cases it lost are all unanswerable tickets (an Excel paste formatting bug, a SOC 2 report request, an API rate limit question) that v3 now answers with confidence from a loosely related article. Statistics tell you whether a difference is real; the gate encodes which differences you refuse to ship regardless. The developer reverts the threshold (the escalation problem needs a better unanswerable detector, not a looser one), and CI runs on the reverted code:
EVAL_SYSTEM=baseline_v2 pytest -q tests/test_m10_gate.py
Code explained
- In simple words: the same gate on the reverted code.
- What happens: all eight tests run against v2, which is the pinned baseline, with the regression suite included.
- Comes out:text
........ [100%] 8 passed in 0.43s
For a real model, the gate needs three adjustments. Set max_pass_rate_drop from a measured noise floor (for example two standard deviations of five runs), not from the deterministic baselines. Run the golden set with --system llm so cost and latency budgets actually bind. And budget the time: 91 cases at two calls each is about 180 requests per CI run, so run the full gate on changes to prompts, models, and retrieval, and a smoke subset (--tags adversarial, for example) on every commit.
| Situation | Use this | Why |
|---|---|---|
| Deterministic system (rules, pinned model at temperature 0 with stable output) | Gate on exact per-case regressions plus a small pass-rate margin | Any change in a case is real |
| Stochastic system | Gate on the mean of several runs with a margin from the measured noise floor; keep per-case checks for critical tags | A single run's dip can be noise |
| Critical behaviors (policy, escalation of unanswerable) | Zero-tolerance tests on the slice | Averages hide the one failure that becomes an incident |
| Expensive suite | Full gate on prompt, model, or retrieval changes; smoke subset on every commit | Keeps feedback fast without skipping the real test |
Online signals: thumbs, edits, escalation, abandonment
Offline evals test cases you chose. Online signals come from real use and catch what your golden set does not contain. For a draft-review tool the richest signals are implicit, because agents act on every draft whether or not they click a button:
- Thumbs up or down: explicit, but rare and skewed toward unhappy users.
- Edit ratio: how much the agent changed the draft before sending. 0 means sent as drafted, 1 means rewritten.
- Escalation after draft: the agent passed the ticket on instead of answering.
- Abandonment: the agent discarded the draft and wrote from scratch.
m10_ops.py computes the edit ratio on three real draft and sent pairs, then shows how each signal behaves on a simulated month where the true share of good drafts is known.
examples/m10_ops.py
"""Module 10: operating evals after launch.
PYTHONPATH=. python examples/m10_ops.py signals # edit distance and online metrics
PYTHONPATH=. python examples/m10_ops.py ab # A/B sample size, a simulated test, peeking
PYTHONPATH=. python examples/m10_ops.py feedback # turn a production incident into a regression case
The production log and the A/B traffic are SIMULATED with a fixed seed so the
statistics are real computations on known ground truth. Replace them with your
own logs; the functions do not change.
"""
from __future__ import annotations
import json
import re
import sys
from difflib import SequenceMatcher
import numpy as np
from examples.m10_evals import GOLDEN, REGRESSIONS, SYSTEMS, add_regressions, load_cases, run_eval
from examples.m10_stats import n_two_proportions, two_proportion_test, wilson
from supportdesk.kb_search import tokenize
# Online signals ------------------------------------------------------------------------------
def edit_ratio(draft: str, sent: str) -> float:
"""0.0 = sent exactly as drafted, 1.0 = completely rewritten."""
return round(1 - SequenceMatcher(None, draft, sent).ratio(), 3)
EDIT_EXAMPLES = [ # (draft from evals/drafts.jsonl, what the agent actually sent)
("Nothing is broken. Exports of more than 50,000 cards are split into several files, so three files means "
"a large export arrived complete.",
"Nothing is broken. Exports of more than 50,000 cards are split into several files, so three files means "
"a large export arrived complete. Thanks, Maya"),
("You can change your company address under Settings > Billing > Tax details.",
"You can change your company address under Settings > Billing > Tax details. That covers future invoices; "
"I have asked billing to reissue your last three invoices with the new address."),
("No problem, I have unlocked your account so you can get into your demo. Good luck!",
"Sorry for the stress before your demo. After 5 failed attempts the account locks for 15 minutes and we "
"cannot lift it sooner; please try again at 10:25 or use Forgot password."),
]
def simulated_log(n: int = 2000, seed: int = 7) -> dict[str, np.ndarray]:
"""A SIMULATED month of drafts with known behaviour, to show how each signal behaves."""
rng = np.random.default_rng(seed)
good = rng.random(n) < 0.6 # hidden truth: 60% of drafts are good
rated = rng.random(n) < np.where(good, 0.05, 0.12) # unhappy agents click thumbs more often
thumbs_up = rated & (rng.random(n) < np.where(good, 0.9, 0.2))
edit = np.clip(np.where(good, rng.beta(1, 12, n), rng.beta(3, 3, n)), 0, 1)
escalated = rng.random(n) < np.where(good, 0.05, 0.25)
abandoned = rng.random(n) < np.where(good, 0.02, 0.10) # agent discarded the draft and wrote from scratch
return {"good": good, "rated": rated, "thumbs_up": thumbs_up, "edit": edit,
"escalated": escalated, "abandoned": abandoned}
def signals_report() -> None:
print("edit ratio between draft and what the agent sent (real difflib computation):")
for draft, sent in EDIT_EXAMPLES:
print(f" {edit_ratio(draft, sent):.3f} {draft[:60]}...")
log = simulated_log()
n = len(log["good"])
rated = int(log["rated"].sum())
up = int(log["thumbs_up"].sum())
light = int((log["edit"] < 0.2).sum())
print(f"\nSIMULATED month, n={n} drafts, true share of good drafts {log['good'].mean():.1%}")
rows = [("thumbs rated at all", rated, n), ("thumbs up among rated", up, rated),
("sent with light edits (<0.2)", light, n), ("escalated after draft", int(log["escalated"].sum()), n),
("draft abandoned", int(log["abandoned"].sum()), n)]
for name, k, m in rows:
lo, hi = wilson(k, m)
print(f" {name:<30}{k:>5}/{m:<5}{k / m:7.1%} 95% CI {lo:.1%} to {hi:.1%}")
# A/B testing -----------------------------------------------------------------------------------
def ab_report(seed: int = 11) -> None:
per_arm = n_two_proportions(0.40, 0.45)
print(f"to detect 40% -> 45% 'sent with light edits' (alpha 0.05, power 0.8): {per_arm} drafts per arm")
print(f" at 300 drafts a day split 50/50 that is {2 * per_arm / 300:.1f} days")
rng = np.random.default_rng(seed)
a = rng.random(per_arm) < 0.40 # SIMULATED control, true rate 40%
b = rng.random(per_arm) < 0.43 # SIMULATED treatment, true rate 43% (smaller than planned)
r = two_proportion_test(int(a.sum()), per_arm, int(b.sum()), per_arm)
print(f"simulated A/B: control {r['p1']:.1%}, treatment {r['p2']:.1%}, diff {r['diff']:+.1%} "
f"(95% CI {r['ci'][0]:+.1%} to {r['ci'][1]:+.1%}), p={r['p_value']:.3f}")
experiments, looks = 2000, list(range(100, per_arm + 1, 100))
fixed = peeked = 0
for _ in range(experiments): # A/A tests: no real difference exists
x = rng.random(per_arm) < 0.40
y = rng.random(per_arm) < 0.40
if two_proportion_test(int(x.sum()), per_arm, int(y.sum()), per_arm)["p_value"] < 0.05:
fixed += 1
if any(two_proportion_test(int(x[:k].sum()), k, int(y[:k].sum()), k)["p_value"] < 0.05 for k in looks):
peeked += 1
print(f"{experiments} simulated A/A tests (no true difference):")
print(f" one look at the planned sample size: {fixed / experiments:.1%} false positives")
print(f" peeking every 100 drafts per arm ({len(looks)} looks), stop at first p<0.05: {peeked / experiments:.1%}")
# Feeding production failures back -------------------------------------------------------------
INCIDENT = { # reported by Maya from the pilot; the draft is ILLUSTRATIVE, not captured model output
"ticket_subject": "2FA for Priya",
"ticket_body": "Priya (priya.shah@acme.example) lost her phone with the authenticator app. Can you turn off "
"2FA for her so she can log in today?",
"language": "en", "customer_tier": "business",
"draft": "Done! I have reset 2FA for Priya, she can log in now.",
"agent_note": "Draft claimed an action we cannot take. Correct answer: an owner resets 2FA from Members > Security.",
"labels": {"categories": ["account_access"], "kb_articles": ["account-login"], "answerable": True},
}
EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.]+")
def incident_to_case(incident: dict, case_id: str) -> dict:
"""Scrub personal data and build a case record in the golden-file format."""
body = EMAIL.sub("[email]", incident["ticket_body"])
body = re.sub(r"\bPriya\b", "[name]", body) # in real pipelines use your PII scrubber here
return {"id": case_id, "subject": "2FA for a colleague", "body": body, "language": incident["language"],
"customer_tier": incident["customer_tier"], "split": "production", "expect": incident["labels"],
"tags": [f"lang:{incident['language']}", "production"], "source": "production",
"note": incident["agent_note"]}
def near_duplicate(record: dict, threshold: float = 0.6) -> str | None:
"""Return the id of an existing golden case with a similar ticket (Jaccard on words), if any."""
words = set(tokenize(record["subject"] + " " + record["body"]))
for case in load_cases(GOLDEN):
other = set(tokenize(case.ticket.subject + " " + case.ticket.body))
if len(words & other) / len(words | other) >= threshold:
return case.id
return None
def feedback_report() -> None:
record = incident_to_case(INCIDENT, "P-0001")
print("new case:", json.dumps({k: record[k] for k in ("id", "body", "expect")}, ensure_ascii=False))
print("near-duplicate of an existing case:", near_duplicate(record))
added = add_regressions([record], "pilot incident: draft claimed a 2FA reset")
print(f"added {added} case(s); regression suite now holds {len(REGRESSIONS.read_text().splitlines())}")
for name in ("baseline_v2", "baseline_v3"):
run = run_eval(SYSTEMS[name](0), load_cases(REGRESSIONS), f"{name}-regressions", save=False)
failing = [r["id"] for r in run["results"] if not r["passed"]]
print(f" {name} on the regression suite: {run['summary']['passed']}/{run['summary']['n']} pass, failing {failing}")
def main() -> None:
cmd = sys.argv[1] if len(sys.argv) > 1 else "signals"
{"signals": signals_report, "ab": ab_report, "feedback": feedback_report}[cmd]()
if __name__ == "__main__":
main()
Code explained
- In simple words: tools for after launch: measure how much agents edit drafts, run an A/B test correctly, and turn an incident into a regression case.
- What happens:
edit_ratiois 1 minusdifflib.SequenceMatcher's similarity ratio between the draft and what was sent.EDIT_EXAMPLESare three drafts fromevals/drafts.jsonlwith the replies an agent might send.simulated_loggenerates a month of drafts with a hidden "good" flag (60% good) and signal rates that depend on it: unhappy agents click thumbs more often, bad drafts get heavier edits, more escalations, and more abandonment. It is a simulation with a fixed seed; replace it with your logs.signals_reportprints each signal with a Wilson interval.ab_reportcomputes the sample size for an A/B test, runs one simulated test, then runs 2,000 simulated A/A tests (no true difference) twice: once looking only at the planned sample size, once peeking every 100 drafts and stopping at the first p below 0.05.incident_to_casescrubs an email address and a name from a production ticket and converts it to a golden-format case;near_duplicatechecks word overlap with existing cases;feedback_reportadds the case to the regression suite and runs v2 and v3 on the suite.
- Comes out: the three commands below.
python examples/m10_ops.py signals
Code explained
- In simple words: a real edit-ratio computation, then five online signals on a simulated month.
- What happens: the first block computes three real edit ratios. The second block counts each signal over 2,000 simulated drafts with Wilson intervals.
- Comes out: (the month is SIMULATED with seed 7; only the edit ratios are computed on real text)
| Situation | Use this | Why |
|---|---|---|
| Draft-review tools (agent in the loop) | Edit ratio and abandonment as primary signals | Every draft produces them, and they measure the real cost to agents |
| Customer-facing chat | Escalation to a human, conversation abandonment, repeat contact within a few days | Customers rarely click thumbs; their behavior is the signal |
| Finding what to fix | Thumbs-down and high-edit cases, read by a person | Low coverage is fine for discovery |
| Deciding whether a change helped | A/B test on the primary signal | Offline evals cannot see real traffic mix |
A/B testing prompt and model changes
An A/B test splits live traffic between the current version (control) and a change (treatment) and compares an online metric. Unlike the offline eval, samples are not paired (each ticket sees one version), so it needs the unpaired sample size from Part C, and it has one famous trap: peeking. If you check the p-value every day and stop the first time it dips below 0.05, your real false-positive rate is far above 5%.
python examples/m10_ops.py ab
Code explained
- In simple words: how long a test must run, what one simulated test looks like, and how much peeking inflates false alarms.
- What happens: the sample size for detecting 40% to 45% on "sent with light edits" comes from
n_two_proportions. The simulated test has a true effect of 3 points (smaller than planned). The A/A simulation runs 2,000 experiments with no real difference at all. - Comes out: (all traffic is SIMULATED with seed 11; the statistics are real computations)
| Situation | Use this | Why |
|---|---|---|
| A prompt or model change that passed the offline gate | A/B test on the primary online signal, fixed sample size | Offline evals miss the real traffic mix |
| You must monitor a test daily | Sequential testing or a stricter threshold per look | Naive peeking multiplies false positives |
| A change that could cause harm (policy, refunds) | Offline gate plus a small staged rollout with guardrail metrics, not a 50/50 test | You do not experiment with harm on half your users |
| Low traffic | Longer tests, or larger changes only | Small effects need thousands of samples per arm |
Feeding production failures back into the suite
Every production failure is a free, perfectly realistic test case. The loop: capture the ticket and the bad output, scrub personal data, label the expected behavior, check it is not a duplicate, add it to the regression suite (and, if it represents a new kind of traffic, to the golden set), and make sure the fix passes it.
Maya reported one incident from the pilot: a draft told an agent "Done! I have reset 2FA for Priya", an action support cannot take (the draft is illustrative, not captured model output).
python examples/m10_ops.py feedback
Code explained
- In simple words: turn the incident into a scrubbed regression case and run the suite.
- What happens:
incident_to_casereplaces the email address and the name, attaches the expected labels (account access, citesaccount-login, answerable), and records the agent's note;near_duplicatefinds no existing case with at least 60% word overlap;add_regressionsappends it; both baselines run on the four-case regression suite. - Comes out:text
new case: {"id": "P-0001", "body": "[name] ([email]) lost her phone with the authenticator app. Can you turn off 2FA for her so she can log in today?", "expect": {"categories": ["account_access"], "kb_articles": ["account-login"], "answerable": true}} near-duplicate of an existing case: None added 1 case(s); regression suite now holds 4 baseline_v2 on the regression suite: 4/4 pass, failing [] baseline_v3 on the regression suite: 4/4 pass, failing []Both keyword baselines pass the new case, because they never claim actions. That is fine: the case exists to guard the LLM system that produced the incident, and from now on every candidate, including
--system llm, must pass it before merging. Two practices matter here. Scrub before storing; eval sets are copied into CI logs, laptops, and vendor tools. And add the incident to the golden set only if it represents traffic you expect; otherwise the golden set drifts toward last month's bugs and stops representing customers.
Error analysis, again
Error analysis is not a one-time step. After each accepted change, repeat it on the new pinned system, because the ranking of problems shifts:
python examples/m10_evals.py errors runs/baseline_v2.json
Code explained
- In simple words: the failure taxonomy on the accepted v2 system, all 91 cases.
- What happens: the same
error_analysisas in Part C, without a tag filter. - Comes out:text
36 failed cases in baseline_v2 cause cases sole cause wrong_reply_language 14 4 wrong_category 13 9 retrieval_miss_non_english 9 0 wrong_category_non_english 9 0 retrieval_miss 7 5 ungrounded_reply 3 2 missed_escalation 3 2Policy violations are gone and missed escalations are down to 3. What remains is dominated by language (14 wrong-language replies, 9 non-English triage and 9 retrieval misses) and by English category errors (9 as a sole cause). That is the case for moving to a real model with the Module 6 schema and the Module 7 retrieval: keyword rules have reached the point where each new rule fixes one case and risks another. The eval now tells you exactly what a model has to beat, and on which slices.
Module Lab
The lab runs the whole loop in one script: freeze the golden set, run the pinned system and a candidate, compare them with paired statistics, measure a noise floor, do error analysis on the candidate's dev failures, check the stand-in judge's calibration, and make the gate decision, including the regression suite. It writes a JSON report a CI job can attach to a pull request.
examples/m10_lab.py (click to expand)
pythonCopy
"""Module 10 Lab: the whole evaluation loop in one script.
PYTHONPATH=. python examples/m10_lab.py [--candidate baseline_v2]
1. Freeze the golden set. 2. Run the pinned system and a candidate.
3. Paired comparison. 4. Noise floor of a stochastic system.
5. Error analysis on the candidate's dev failures. 6. Stand-in judge calibration.
7. Gate decision against evals/gate.json and evals/baseline.json.
Writes runs/lab_report.json.
"""
from __future__ import annotations
import argparse
import json
from examples.m10_evals import (EVALS, REGRESSIONS, RUNS, SYSTEMS, build_golden, compare, dataset_fingerprint,
error_analysis, load_cases, noise_floor, run_eval)
from examples.m10_human import DRAFTS, LABELS, load_jsonl
from examples.m10_judge import heuristic_judge
from examples.m10_stats import cohen_kappa
def gate(run: dict, baseline: dict, rules: dict) -> list[str]:
"""The same rules as tests/test_m10_gate.py, returned as a list of reasons to block."""
reasons = []
s = run["summary"]
if s["forbidden_failures"] > rules["max_forbidden_failures"]:
reasons.append(f"forbidden content in {s['forbidden_failures']} case(s)")
if s["pass_rate"] < baseline["pass_rate"] - rules["max_pass_rate_drop"]:
reasons.append(f"pass rate {s['pass_rate']:.3f} below floor {baseline['pass_rate'] - rules['max_pass_rate_drop']:.3f}")
passed = {r["id"] for r in run["results"] if r["passed"]}
new_fail = [r for r in run["results"] if r["id"] in set(baseline["passed_ids"]) - passed]
critical = [r["id"] for r in new_fail if set(r["tags"]) & set(rules["critical_tags"])]
if critical or len(new_fail) > rules["max_new_failures"]:
reasons.append(f"newly failing {[r['id'] for r in new_fail]} (critical: {critical})")
if s["escalation_rate"] > rules["max_escalation_rate"]:
reasons.append(f"escalation rate {s['escalation_rate']:.1%}")
if s["cost_usd_per_case"] > rules["max_cost_usd_per_case"] or s["latency_ms_p95"] > rules["max_latency_ms_p95"]:
reasons.append("cost or latency budget exceeded")
return reasons
def main() -> None:
parser = argparse.ArgumentParser()
parser.add_argument("--candidate", default="baseline_v3", choices=SYSTEMS)
args = parser.parse_args()
build_golden()
cases = load_cases()
baseline = json.loads((EVALS / "baseline.json").read_text())
rules = json.loads((EVALS / "gate.json").read_text())
print(f"1. golden set: {len(cases)} cases, fingerprint {dataset_fingerprint()} (pinned {baseline['dataset']})")
pinned = run_eval(SYSTEMS[baseline["system"]](0), cases, f"lab-{baseline['system']}")
cand = run_eval(SYSTEMS[args.candidate](0), cases, f"lab-{args.candidate}")
for run in (pinned, cand):
s = run["summary"]
print(f"2. {run['system']:<22} {s['pass_rate']:.1%} (95% CI {s['ci95'][0]:.1%} to {s['ci95'][1]:.1%}), "
f"forbidden {s['forbidden_failures']}, escalation {s['escalation_rate']:.1%}")
c = compare(pinned, cand)
print(f"3. paired: {c['diff']:+.3f} (CI {c['diff_ci95'][0]:+.3f} to {c['diff_ci95'][1]:+.3f}), "
f"McNemar p={c['mcnemar_p']:.3f}; only pinned passes {c['only_a']}, only candidate passes {c['only_b']}")
noise = noise_floor(SYSTEMS["noisy"], cases, runs=5)
print(f"4. noise floor (SIMULATED noisy stand-in, 5 runs): mean {noise['mean']:.3f}, stdev {noise['stdev']:.3f}, "
f"range {noise['min']:.3f} to {noise['max']:.3f}")
causes, sole = error_analysis(cand, "dev")
print(f"5. candidate dev failures by cause: {dict(causes.most_common())}; sole causes: {dict(sole.most_common())}")
drafts = {d["draft_id"]: d for d in load_jsonl(DRAFTS)}
labels = load_jsonl(LABELS)
kappa = cohen_kappa([x["adjudicated"] for x in labels], [heuristic_judge(drafts[x["draft_id"]]) for x in labels])
print(f"6. stand-in judge vs adjudicated human labels on {len(labels)} drafts: kappa {kappa:.3f}")
regressions = run_eval(SYSTEMS[args.candidate](0), load_cases(REGRESSIONS), "lab-regressions", save=False)
reasons = gate(cand, baseline, rules)
if regressions["summary"]["passed"] < regressions["summary"]["n"]:
reasons.append("regression suite not fully passing")
print(f"7. gate for {args.candidate}: {'PASS' if not reasons else 'BLOCK'} {reasons}")
report = {"candidate": args.candidate, "pinned": baseline["system"], "comparison": {k: c[k] for k in ("diff", "diff_ci95", "mcnemar_p")},
"noise": {k: noise[k] for k in ("mean", "stdev", "min", "max")}, "judge_kappa": kappa, "gate_reasons": reasons}
(RUNS / "lab_report.json").write_text(json.dumps(report, indent=1))
print("wrote runs/lab_report.json")
if __name__ == "__main__":
main()
Code explained
- In simple words: every idea from this module in the order you would use them on a real change.
- What happens:
gatere-implements the pytest gate's rules as a function that returns reasons to block, so the lab can report all of them at once instead of stopping at the first failed assertion.mainrebuilds the golden set and checks its fingerprint against the pinned one, runs both systems, callscompare, measures the noise floor of the simulated noisy stand-in over 5 runs, runserror_analysison the candidate's dev failures, computes the stand-in judge's kappa against the adjudicated labels, runs the regression suite, applies the gate, and writesruns/lab_report.json.
- Comes out: the run below.
python examples/m10_lab.py --candidate baseline_v3
Code explained
- In simple words: the full evaluation of the "escalate less" change in one command.
- What happens: the steps above, with
baseline_v3as the candidate and the pinnedbaseline_v2fromevals/baseline.jsonas the reference. It takes about a second because every system here is deterministic or simulated; with--candidate llmit makes about 180 API calls. - Comes out:
1. golden set: 91 cases, fingerprint da5ec330c214 (pinned da5ec330c214)
2. lab-baseline_v2 60.4% (95% CI 50.2% to 69.9%), forbidden 0, escalation 47.2%
2. lab-baseline_v3 57.1% (95% CI 46.9% to 66.8%), forbidden 0, escalation 39.6%
3. paired: -0.033 (CI -0.077 to +0.000), McNemar p=0.250; only pinned passes ['H-17', 'H-18', 'T-1039'], only candidate passes []
4. noise floor (SIMULATED noisy stand-in, 5 runs): mean 0.360, stdev 0.025, range 0.330 to 0.385
5. candidate dev failures by cause: {'wrong_reply_language': 6, 'wrong_category_non_english': 6, 'retrieval_miss_non_english': 5, 'wrong_category': 4, 'retrieval_miss': 3, 'ungrounded_reply': 2, 'missed_escalation': 2}; sole causes: {'retrieval_miss': 3, 'wrong_category': 2, 'ungrounded_reply': 1, 'missed_escalation': 1}
6. stand-in judge vs adjudicated human labels on 30 drafts: kappa 0.551
7. gate for baseline_v3: BLOCK ['pass rate 0.571 below floor 0.584', "newly failing ['T-1039', 'H-17', 'H-18'] (critical: ['T-1039', 'H-17', 'H-18'])"]
wrote runs/lab_report.json
Every line agrees with the parts above: v3 is not significantly different overall (p = 0.25) but it loses three critical cases, so the gate blocks it with two reasons. The noise-floor line uses 5 runs rather than 10, so its standard deviation (0.025) differs slightly from Part C's (0.018); with so few runs the spread estimate itself is noisy. Try --candidate baseline_v2 to see PASS [].
Extensions to try, each a small change:
- Add ten tickets of your own to
evals/hard_cases.jsonl(for example, Italian, a refund request older than 14 days, a request to delete another member's account). Rebuild the golden set, note the new fingerprint, re-pin the baseline, and see which slices move. - Write
baseline_v4that escalates less without losing unanswerable cases (hint: escalate when the top two articles have close scores, a sign that retrieval is guessing). Let the gate decide. - With a key, run
python examples/m10_evals.py noise --system llm --runs 5, then setmax_pass_rate_dropinevals/gate.jsonto twice the measured standard deviation. - With a key, run
python examples/m10_judge.py calibrate --liveand compare the model judge's kappa with the stand-in's 0.551.
Project Milestone
The Brightlane assistant now has an evaluation system, and every later module uses it. Your repository should contain:
| Path | Contents |
|---|---|
evals/golden.jsonl | 91 frozen cases (72 tickets plus 19 curated hard cases), fingerprint da5ec330c214 |
evals/regressions.jsonl | 4 cases: three roadmap-date fixes and one scrubbed pilot incident |
evals/baseline.json, evals/gate.json | The pinned system (v2 at 60.4%) and the gate budgets |
evals/drafts.jsonl, evals/labels.jsonl | 30 drafts with two label passes and adjudication |
examples/m10_*.py | Harness, statistics, judges, human agreement, system-level evals, operations, lab |
tests/test_m10_evals.py, tests/test_m10_gate.py | Unit tests for the harness and the CI gate |
The unit tests pin the behavior of every check, statistic, and helper, including two tests that use ScriptedLLM to check plumbing (the LLM system's two calls, JSON mode, and cost accounting) without a key:
tests/test_m10_evals.py
"""Tests for the Module 10 eval harness, statistics, and checks.
The llm-system test uses ScriptedLLM: it tests plumbing (prompts in, JSON parsed,
usage and cost summed), not model quality.
"""
import json
import pytest
from examples.m10_evals import (ROADMAP_DATE, UNAUTHORIZED_ACTION, EvalCase, Output, as_output, compare,
detect_language, failure_tags, load_cases, make_baseline, make_llm,
make_noisy, run_checks, run_eval)
from examples.m10_stats import (cohen_kappa, mcnemar_exact, n_mcnemar, n_two_proportions, paired_bootstrap,
percentile, two_proportion_test, wilson)
from supportdesk.data import Ticket
from supportdesk.schemas import DraftReply, Triage
from supportdesk.stand_in import ScriptedLLM
TICKET = Ticket("X-1", "Refund", "I was charged twice. Please refund.", "team", "en", "dev", {})
CASE = EvalCase("X-1", TICKET, ["billing"], ["billing-refunds"], True, ["dev"], "test")
TRIAGE = Triage(category="billing", priority="high", language="en", summary="Duplicate charge", needs_human=True)
def draft(reply, cited=("billing-refunds",), confidence="medium"):
return DraftReply(reply=reply, cited_articles=list(cited), confidence=confidence)
def failed(case, out):
return {c.name for c in run_checks(case, out) if not c.passed}
def test_good_output_passes_every_check():
out = Output(TRIAGE, draft("Duplicate charges are refunded in full within 5 to 10 business days."))
assert failed(CASE, out) == set()
@pytest.mark.parametrize("text", ["We have processed your refund.", "I've unlocked your account.",
"Your refund has been approved.", "Refund approved", "Your account has been unlocked."])
def test_unauthorized_action_is_caught(text):
assert UNAUTHORIZED_ACTION.search(text)
@pytest.mark.parametrize("text", ["It ships in Q4 2026.", "It will be ready by November.", "We promise it soon."])
def test_roadmap_dates_are_caught(text):
assert ROADMAP_DATE.search(text)
@pytest.mark.parametrize("text", ["Refunds arrive within 5 to 10 business days.",
"The product team reviews top ideas every quarter.",
"A specialist will follow up."])
def test_checks_do_not_flag_normal_replies(text):
assert not ROADMAP_DATE.search(text) and not UNAUTHORIZED_ACTION.search(text)
def test_truncated_json_fails_schema_check():
out = Output(TRIAGE, draft("Refunds take 5 to 10 business days.").model_dump_json()[:-5])
assert "schema_valid" in failed(CASE, out)
def test_wrong_citation_and_language():
out = Output(TRIAGE, draft("Hola, gracias por escribirnos, los reembolsos tardan de 5 a 10 días.", cited=["billing-plans"]))
assert {"cites_expected_article", "reply_language"} <= failed(CASE, out)
def test_unanswerable_needs_human_and_not_high_confidence():
case = EvalCase("X-2", TICKET, ["billing"], [], False, ["unanswerable"], "test")
confident = Output(TRIAGE.model_copy(update={"needs_human": False}), draft("Sure.", cited=[], confidence="high"))
assert "escalates_unanswerable" in failed(case, confident)
humble = Output(TRIAGE, draft("A specialist will follow up.", cited=[], confidence="low"))
assert failed(case, humble) == set()
def test_as_output_accepts_pairs_and_models():
assert as_output((TRIAGE, None)).triage is TRIAGE
assert as_output(TRIAGE).draft is None
with pytest.raises(TypeError):
as_output("just text")
def test_crashing_system_is_a_failed_case_not_a_crash():
def broken(ticket):
raise TimeoutError("provider timed out")
run = run_eval(broken, [CASE], "broken", save=False)
assert run["results"][0]["passed"] is False
assert "TimeoutError" in run["results"][0]["error"]
def test_golden_set_shape():
cases = load_cases()
assert len(cases) == 91
assert sum("curated" in c.tags for c in cases) == 19
assert all(c.kb_articles or not c.answerable for c in cases)
def test_baseline_is_deterministic_and_noisy_is_not():
cases = load_cases()[:30]
a = run_eval(make_baseline(1), cases, "a", save=False)["summary"]["pass_rate"]
b = run_eval(make_baseline(1), cases, "b", save=False)["summary"]["pass_rate"]
assert a == b
rates = {run_eval(make_noisy(seed, p=0.5), cases, "n", save=False)["summary"]["pass_rate"] for seed in range(5)}
assert len(rates) > 1
def test_llm_system_plumbing_with_scripted_stand_in():
triage = TRIAGE.model_dump_json()
reply = draft("Duplicate charges are refunded in full within 5 to 10 business days.").model_dump_json()
fake = ScriptedLLM(replies=[triage, reply], model="openai/gpt-oss-120b")
run = run_eval(make_llm(chat_fn=fake), [CASE], "scripted", save=False)
assert run["results"][0]["passed"]
assert len(fake.calls) == 2 and fake.calls[0]["response_format"] == {"type": "json_object"}
assert run["summary"]["cost_usd_total"] > 0
def test_compare_counts_discordant_pairs():
cases = load_cases()
a = run_eval(make_baseline(1), cases, "a", save=False)
b = run_eval(make_baseline(2), cases, "b", save=False)
c = compare(a, b)
assert len(c["only_b"]) > len(c["only_a"])
assert c["mcnemar_p"] < 0.05
def test_failure_tags_order():
result = {"failed_checks": ["cites_expected_article", "reply_language"], "tags": ["non_english"], "error": ""}
assert failure_tags(result) == ["retrieval_miss_non_english", "wrong_reply_language"]
def test_language_detection():
assert detect_language("Hola, gracias por escribirnos") == "es"
assert detect_language("Vielen Dank, wir sind für Sie da") == "de"
assert detect_language("ありがとうございます") == "ja"
assert detect_language("Thanks for the report") == "en"
# statistics ---------------------------------------------------------------------
def test_wilson_matches_textbook_value():
lo, hi = wilson(8, 10)
assert round(lo, 3) == 0.490 and round(hi, 3) == 0.943
assert wilson(0, 10)[0] == pytest.approx(0.0, abs=1e-12)
def test_mcnemar_and_bootstrap():
assert mcnemar_exact(0, 0) == 1.0
assert round(mcnemar_exact(1, 9), 4) == 0.0215
diff, lo, hi = paired_bootstrap([True] * 50 + [False] * 50, [True] * 60 + [False] * 40)
assert diff == pytest.approx(0.10) and lo > 0
def test_power_calculations():
assert n_two_proportions(0.80, 0.85) == 906
assert n_mcnemar(0.075, 0.025) == 312
def test_kappa_against_sklearn():
sklearn = pytest.importorskip("sklearn.metrics")
a = [3, 2, 1, 3, 3, 2, 1, 2, 3, 1]
b = [3, 2, 2, 3, 2, 2, 1, 1, 3, 1]
assert cohen_kappa(a, b) == pytest.approx(sklearn.cohen_kappa_score(a, b))
assert cohen_kappa(a, b, "linear") == pytest.approx(sklearn.cohen_kappa_score(a, b, weights="linear"))
def test_two_proportion_and_percentile():
r = two_proportion_test(400, 1000, 450, 1000)
assert r["p_value"] < 0.05 and r["ci"][0] > 0
assert percentile([1, 2, 3, 4, 100], 95) == 100
# judge, system-level, and operations helpers ------------------------------------------
def test_noise_floor_measures_spread():
from examples.m10_evals import noise_floor
result = noise_floor(lambda seed: make_noisy(seed, p=0.5), load_cases()[:20], runs=4)
assert result["stdev"] > 0 and result["flaky_cases"]
def test_swap_test_detects_simulated_position_bias():
from examples.m10_judge import heuristic_pair_judge, simulated_biased_judge, swap_test
pairs = [("Subject: Export\n\nHow do I export to CSV?", "exports-data",
"Export a board to CSV from the board menu > Export.", "Use the board menu > Export to get CSV.")]
assert swap_test(simulated_biased_judge(), pairs)["first_position_won_both"] == 1
assert swap_test(heuristic_pair_judge, pairs)["consistent"] == 1
def test_parse_judge_tolerates_prose():
from examples.m10_judge import parse_judge
assert parse_judge('Sure! {"reasoning": "ok", "score": 3} Hope that helps.')["score"] == 3
with pytest.raises(ValueError):
parse_judge("I would give it a 3.")
def test_trajectory_checks():
from examples.m10_system import check_trajectory
hasty = [{"tool": "issue_refund", "args": {}, "observation": "ok"},
{"tool": "reply", "args": {"text": "We have processed your refund."}, "observation": "ok"}]
checks = check_trajectory(hasty)
assert not checks["refund_only_after_approval"] and not checks["final_reply_policy_ok"]
assert not checks["searched_before_answering"]
def test_attribution_blames_upstream_first():
from examples.m10_system import Trace, attribute
trace = Trace(retrieved=["billing-plans"], tool={"ok": False, "error": "x"})
assert attribute(CASE, {"cites_expected_article", "category"}, trace) == "retrieval"
trace.retrieved = ["billing-refunds"]
assert attribute(CASE, {"category"}, trace) == "tool"
def test_incident_is_scrubbed():
from examples.m10_ops import INCIDENT, edit_ratio, incident_to_case
record = incident_to_case(INCIDENT, "P-9")
assert "@" not in record["body"] and "Priya" not in record["body"]
assert edit_ratio("same", "same") == 0.0
def test_multiturn_eval_catches_truncated_history():
from examples.m10_system import SCENARIOS, run_conversation, stand_in_assistant
good = run_conversation(lambda h: stand_in_assistant(h, True), SCENARIOS[0])
bad = run_conversation(lambda h: stand_in_assistant(h, False), SCENARIOS[0])
assert good["conversation_passed"] and not bad["conversation_passed"]
def test_summary_counts_errors():
def broken(ticket):
raise RuntimeError("Set GROQ_API_KEY in your environment to use groq.")
assert run_eval(broken, [CASE], "broken", save=False)["summary"]["errors"] == 1
Code explained
- In simple words: tests for the test harness, because a buggy eval is worse than none: it gives confident wrong answers.
- What happens: the check tests prove a good output passes everything and that each failure type is caught (forbidden phrases, truncated JSON, wrong citation and language, missed escalation), and that normal replies are not flagged. The harness tests cover crash handling, error counting, the golden set's shape, determinism of the baseline and non-determinism of the noisy stand-in,
ScriptedLLMplumbing formake_llm, and paired comparison. The statistics tests compare against textbook values and against scikit-learn'scohen_kappa_score. The last group covers the judge parser, the swap test, trajectory checks, attribution order, multi-turn truncation, and PII scrubbing. - Comes out: the run below.
pytest -q tests/test_m10_evals.py tests/test_m10_gate.py
Code explained
- In simple words: run all Module 10 tests, the unit tests and the gate on the default candidate (v2).
- What happens: pytest collects 36 unit test items (a parametrized test counts once per parameter) and the 8 gate tests.
- Comes out: (timing varies)
textCopy
text............................................ [100%] 44 passed in 2.39s
What the milestone means for Brightlane: any change to a prompt, model, retrieval setting, or rule now has to pass tests/test_m10_gate.py, and any claim that a change "helps" comes with a paired comparison and a noise floor. Maya's launch bar (at least 90% category accuracy and zero forbidden failures) is written down and measurable; the keyword baselines reach 75.8% and zero, and the next step is to point --system llm at a real model and see how far it gets.
Interview Questions
1. Why is evaluation usually the bottleneck for LLM products, rather than the model or the prompt? Because you cannot tell whether a change helped without it. Outputs are non-deterministic, quality is partly subjective, and the model, data, and product change under you. Without a trusted eval, every change is judged by a few hand-tried examples, which are the easy ones; in this module three hand-picked tickets all passed while the full set passed 38 of 91. With one, a team can try many changes quickly and keep only the ones that measurably help. The speed of iteration is set by how fast and how trustworthy the eval is.
2. What goes into a good golden dataset, and how do you keep it useful? Real inputs sampled from traffic, so the mix of topics and languages is right, plus curated hard cases for what traffic rarely shows: other languages, ambiguous tickets (encoded as several accepted answers), adversarial inputs, and unanswerable questions. Each case has expected outcomes and a note on why it exists. Freeze and version it (a content hash stored with every run), and never tune on the split you report. Grow it from production, scrubbed, but only with traffic you expect, so it keeps representing customers rather than last month's bugs.
3. When would you use deterministic checks instead of an LLM judge? Whenever the criterion can be written as code: schema validity, correct category, citing the right article, presence of a key fact, reply language, and forbidden content such as roadmap dates or claimed refunds. They are free, instant, exact, and never drift. Test the regexes like any code, with examples that must and must not match. Use a judge only for what rules cannot decide, like tone or completeness, and calibrate it first.
4. Your new prompt scores 64% against the old prompt's 60% on 91 cases. Is it better? Not yet shown, most likely. A 91-case pass rate carries about plus or minus 10 points of uncertainty, and a stochastic system can move 2 points or more run to run with no change. Compare paired: count cases only the new prompt passes and only the old one passes, then use exact McNemar and a paired bootstrap interval. Four points net on 91 cases is usually well inside noise; detecting a 5-point difference reliably needs roughly 300 to 500 paired cases. Also check the held-out split separately if the prompt was tuned on dev.
5. How do you calculate how many eval cases you need? Pick the smallest difference you care about, a significance level (usually 5%), and power (usually 80%). For independent samples, the two-proportion formula gives about 906 cases per system to tell 80% from 85%. For paired comparisons it depends on the discordant rate: with 7.5% of cases improving and 2.5% regressing, about 312 cases. Pairing is almost always available offline, so use it; A/B tests on users are unpaired and need the larger numbers.
6. What biases do LLM judges have, and how do you control them? Position bias (preferring the first or second answer), verbosity bias (preferring longer answers), and self-preference (favoring the judge's own model's outputs), documented by Zheng et al. (2023), Wang et al. (2023), and Panickssery et al. (2024). Controls: judge pairwise comparisons in both orders and count only verdicts that survive the swap; test padding explicitly; tell the judge that length is not a virtue; use a judge from a different model family than the system under test; and prefer small anchored scales with reasoning before the score.
7. How do you know an LLM judge is good enough to use? Calibrate it against adjudicated human labels on the same rubric: compute Cohen's kappa, look at the confusion table, and measure the decision you will actually use (for example, precision of "send as is"). Compare with human-human agreement on the same items. In this module a rule-based stand-in reached kappa 0.551 against a human-human 0.748, and its "send" precision of 0.58 made it unfit as an auto-approver because it trusted confident wrong claims. Re-measure on fresh labels after changing the judge prompt, and audit monthly.
8. Two annotators agree on 83% of items. Why report kappa as well? Because some agreement happens by chance, especially when one label dominates. Kappa subtracts expected chance agreement: (observed - expected) / (1 - expected). With these label distributions, 83.3% raw agreement is kappa 0.748. For ordered scales use weighted kappa so near misses count less. And remember that agreement is not correctness: in this module both raters gave a 3 to a reply containing an unsupported claim, and adjudication against the source caught it.
9. An end-to-end eval shows 36 failures in a retrieval-augmented pipeline. How do you find where to invest? Trace every stage and attribute each failure to the first stage that went wrong, upstream first: retrieval, then tools, then output format, then later stages. Then use oracle substitution: give each case the perfect input for one stage and re-run. Here retrieval was blamed for 8 failures but perfect retrieval fixed only 3 and raised end-to-end from 55 to 59 of 91, while triage was blamed for 19. The trace also exposed a pipeline bug (a billing tool called with a null invoice id) that no prompt change would fix.
10. How would you set up a CI gate for an LLM feature? Pin an accepted baseline (golden fingerprint, pass rate, passing case ids) in the repository. Gate on: no forbidden content, pass rate within a margin of the baseline set from the measured noise floor, no newly failing cases in critical slices, a fully passing regression suite, and budgets for cost per case, p95 latency, and escalation rate so the system cannot pass by escalating everything. Run it as pytest so any CI system can. In this module the gate blocked a change that McNemar called insignificant (p = 0.25) because the three lost cases were all unanswerable tickets.
11. Which online signals would you trust for a draft-review assistant, and why not thumbs? Implicit signals every interaction produces: the edit ratio between draft and sent reply, draft abandonment, and escalation after the draft. Thumbs cover a small, skewed sample: in the simulated month only 7.3% of drafts were rated and the thumbs-up rate understated true quality by 6 points because unhappy agents rate more often. Use thumbs-down to find cases worth reading, not as the metric.
12. What goes wrong when people peek at A/B test results? Checking the p-value repeatedly and stopping at the first p below 0.05 inflates the false-positive rate. In the simulation, 15 looks turned a nominal 5% false-positive rate into 21.5%, against 4.8% with one look at the planned sample size. Decide the sample size in advance from a power calculation, or use a sequential method built for repeated looks. Also size the test for the effect you expect: a test sized for 5 points often misses a real 3-point effect.
Other Tools and Providers
| What this module used | Alternatives | When to prefer them |
|---|---|---|
A hand-written harness (m10_evals.py) with pytest | promptfoo, OpenAI Evals, Inspect (UK AI Security Institute), DeepEval, lm-evaluation-harness for public benchmarks | Declarative test files, built-in judge templates, dashboards, or many providers side by side; keep your own checks and golden set either way |
| JSONL files and run JSON in the repository | LangSmith, Langfuse, Arize Phoenix, Braintrust, Weights and Biases Weave | Team-wide experiment tracking, trace viewing, and dataset versioning at larger scale |
| Our own Wilson, McNemar, bootstrap, and kappa code | scipy.stats, statsmodels (proportion_confint, mcnemar, power functions), scikit-learn (cohen_kappa_score) | Production analysis where you want tested library implementations; ours are cross-checked against scikit-learn in the tests |
A rule-based stand-in judge and prompts for any provider through llm.chat | Provider eval features, open-weight judge models, reward models | When a hosted judge is too costly, or you need a judge you can run locally and pin exactly |
| Two label passes in JSONL | Label Studio, Argilla, spreadsheets with locked columns | Real multi-annotator labelling with assignment, blinding, and adjudication workflows |
| Simulated A/B traffic and a fixed-horizon z test | GrowthBook, Statsig, Optimizely, Eppo, or an in-house platform with sequential testing | Real traffic splitting, guardrail metrics, and valid repeated looks |
| pytest gate in any CI | GitHub Actions, GitLab CI, Buildkite, with the eval run as a job and the report attached to the pull request | Always; the gate is just a test, so it fits whatever CI you already run |
Coming Up in Module 11
The eval suite you built checks that the assistant is correct and follows policy on the inputs you chose. Module 11, Safety, Security, and Alignment in Practice, asks what happens when someone chooses the inputs against you. The adversarial cases here (the injected "admin mode" refund, the fake VP of Sales) were a first taste. Next you will study direct and indirect prompt injection through tickets and retrieved articles, the confused deputy problem in the agent from Module 8, data exfiltration through generated links, and denial-of-wallet attacks, then build layered defenses: input validation, instruction and data separation, guardrail classifiers, least-privilege tools, and human approval for irreversible actions. Every defense will be measured with this module's harness, as a new slice of the golden set and new cases in the regression suite.