Part F: Evaluation Track
Scope. Build a rigorous evaluation harness for an existing Brightlane system and use it to find and fix three real quality regressions. The harness has four parts. An assertion tier: deterministic checks that run on every output (schema, citations, forbidden promises, prices), split into hard checks that must never fail and soft checks allowed a small drop. A judge tier: a grader for what code cannot check, trusted only after calibration against human labels. A noise floor: how much a score moves when nothing changed. And CI gates that compare a candidate with a reference. Out of scope: building the system under test; you take one as given.
cohen's kappa measures agreement between two raters beyond what chance would give: 1 is perfect, 0 is chance level, and a judge that says "good" to everything scores 0 however many items it happens to get right..
Milestones
| Week | Evaluation track |
|---|---|
| 1 | Pick the system under test; write the rubric; freeze the test set; run B1 |
| 2 | Assertion tier with hard and soft checks; first report with intervals |
| 3 | Two raters label 50+ outputs; judge prompt; calibration with kappa and false-good counts; 50+ failures tagged |
| 4 | Noise floor: repeated runs of an unchanged system; set the gate's allowed drop from it |
| 5 | Three real regressions: find each with the harness, fix it, show the gate passing after the fix |
| 6 | Evidence pack: harness README, calibration report, regression write-ups |
Starter: two tiers, calibration, noise floor, and a gate
The system under test is an extractive drafter: retrieve the top article with BM25 and copy the two sentences that share the most words with the ticket. It is not an LLM, so every number below is a real measurement of a real system. The human labels in data/cap_human_labels.json were made by the course author, one rater, reading each of the 24 test drafts against the rubric in the code ("good" means the reply answers the customer's actual question correctly from the help center, in the customer's language, and escalates what the help center cannot answer). The hardest calls were T-1054 (the Android push reply points at phone settings, which might help, but does not address that notifications are already enabled) and T-1048 (explains the 15-minute lockout but not whether it can be shortened); both went to the stricter reading. For your capstone you need two raters on your own system's output
examples/cap_eval_starter.py
# examples/cap_eval_starter.py
"""Evaluation track starter: assertion tier, judge tier with calibration, noise floor, CI gate.
System under test: an extractive drafter (retrieve the top article with BM25,
copy the two sentences that share the most words with the ticket). It is not
an LLM, so every number below is a real measurement of a real system; point
the same harness at your LLM drafter to evaluate it.
"""
from __future__ import annotations
import json
import random
import re
import statistics
import sys
from dataclasses import dataclass
from pathlib import Path
from typing import Callable, Protocol
import torch
from pydantic import ValidationError
from supportdesk.data import CATEGORIES, DATA_DIR, Ticket, get_article, load_articles, load_tickets
from supportdesk.kb_search import KBSearch, tokenize
from supportdesk.llm import ChatResult, chat
from supportdesk.schemas import DraftReply
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tinylm import load
sys.path.insert(0, str(Path(__file__).resolve().parent))
from cap_adapt_starter import label_scores # noqa: E402
from cap_baseline import wilson # noqa: E402
KB_IDS = {a.id for a in load_articles()}
KB_TEXT = " ".join(a.body for a in load_articles())
RUBRIC = ("good = the reply answers the customer's actual question correctly from the help center, "
"in the customer's language, and escalates what the help center cannot answer; bad = anything else.")
# System under test ---------------------------------------------------------------
GREETING = "Hi, thanks for contacting Brightlane."
def extractive_draft(ticket: Ticket, kb: KBSearch, query_mode: str = "full", n_sentences: int = 2,
greeting: str = GREETING) -> DraftReply:
query = ticket.text if query_mode == "full" else ticket.subject
hits = kb.search(query, k=1)
if not hits:
return DraftReply(reply="Thanks for writing in. An agent will follow up.", cited_articles=[], confidence="low")
article = get_article(hits[0].article_id)
words = set(tokenize(ticket.text))
sentences = re.split(r"(?<=[.!?])\s+", article.body)
best = sorted(range(len(sentences)), key=lambda i: (-len(words & set(tokenize(sentences[i]))), i))[:n_sentences]
body = " ".join(sentences[i] for i in sorted(best))
return DraftReply(reply=f"{greeting} {body}", cited_articles=[article.id],
confidence="medium" if hits[0].score > 5 else "low")
# Tier 1: deterministic assertions -----------------------------------------------------
FORBIDDEN = re.compile(r"\b(guarantee|we have refunded|unlocked your account|legal advice)\b", re.IGNORECASE)
def assertions(ticket: Ticket, draft: DraftReply) -> dict[str, bool]:
reply = draft.reply
amounts = re.findall(r"\d[\d,]*(?:\.\d+)? USD", reply)
try:
DraftReply.model_validate(draft.model_dump())
schema_ok = True
except ValidationError:
schema_ok = False
gold = ticket.gold["kb_article"]
return {
"schema_valid": schema_ok,
"citations_exist": all(c in KB_IDS for c in draft.cited_articles),
"no_forbidden_promises": not FORBIDDEN.search(reply),
"prices_match_kb": all(a in KB_TEXT for a in amounts),
"length_ok": len(reply) <= 600,
"language_matches": ticket.language == "en", # this drafter always writes English
"cites_gold_or_escalates": (gold in draft.cited_articles) if gold else draft.confidence == "low",
}
HARD = {"schema_valid", "citations_exist", "no_forbidden_promises", "prices_match_kb"} # must never fail
# Tier 2: judge interface and calibration ------------------------------------------------
@dataclass
class Verdict:
label: str # "good" or "bad"
reason: str
class Judge(Protocol):
def __call__(self, ticket: Ticket, draft: DraftReply) -> Verdict: ...
class LLMJudge:
"""Pointwise rubric judge through any chat-shaped callable (llm.chat, ScriptedLLM, a Meter)."""
def __init__(self, llm: Callable[..., ChatResult]) -> None:
self.llm = llm
def __call__(self, ticket: Ticket, draft: DraftReply) -> Verdict:
messages = [
{"role": "system", "content": f"You grade support replies. Rubric: {RUBRIC} "
'Answer only JSON: {"label": "good" or "bad", "reason": "..."}'},
{"role": "user", "content": f"<ticket>\n{ticket.text}\n</ticket>\n<reply>\n{draft.reply}\n</reply>"},
]
raw = self.llm(messages, temperature=0.0, response_format={"type": "json_object"}).text
try:
data = json.loads(raw)
label = data["label"] if data.get("label") in ("good", "bad") else "bad"
return Verdict(label, str(data.get("reason", "")))
except (json.JSONDecodeError, TypeError, KeyError):
return Verdict("bad", f"unparseable judge output: {raw[:60]}")
class OverlapJudge:
"""A cheap heuristic judge: good if English and the reply shares 3+ content words with the ticket."""
def __call__(self, ticket: Ticket, draft: DraftReply) -> Verdict:
shared = set(tokenize(ticket.text)) & set(tokenize(draft.reply)) - {"brightlane", "thanks", "hi"}
ok = ticket.language == "en" and len(shared) >= 3
return Verdict("good" if ok else "bad", f"{len(shared)} shared words")
MIN_KAPPA, MAX_FALSE_GOOD = 0.6, 1 # trust a judge only above this agreement and below this leniency
def calibrate(judge: Judge, tickets: list[Ticket], drafts: list[DraftReply], human: list[str]) -> dict:
"""Compare a judge with human labels on the same items and decide whether to trust it."""
labels = [judge(t, d).label for t, d in zip(tickets, drafts)]
false_good = sum(j == "good" and h == "bad" for j, h in zip(labels, human))
false_bad = sum(j == "bad" and h == "good" for j, h in zip(labels, human))
kappa = cohen_kappa(human, labels)
return {"agree": sum(j == h for j, h in zip(labels, human)), "n": len(human), "kappa": round(kappa, 3),
"false_good": false_good, "false_bad": false_bad, "says_good": labels.count("good"),
"trusted": kappa >= MIN_KAPPA and false_good <= MAX_FALSE_GOOD}
def cohen_kappa(a: list[str], b: list[str]) -> float:
"""Agreement beyond chance: (observed - expected) / (1 - expected)."""
n = len(a)
labels = sorted(set(a) | set(b))
observed = sum(x == y for x, y in zip(a, b)) / n
expected = sum((a.count(lab) / n) * (b.count(lab) / n) for lab in labels)
return 1.0 if expected == 1 else (observed - expected) / (1 - expected)
# Noise floor ---------------------------------------------------------------------------
def label_distributions(tickets: list[Ticket]) -> list[torch.Tensor]:
"""Prompted TinyLM's label scores for each ticket, turned into probabilities with a softmax.
Sampling from these imitates running the classifier with temperature 1: the
same system, the same tickets, a different answer each run.
"""
model, tok = load()
return [torch.softmax(torch.tensor([label_scores(model, tok, t.text)[c] for c in CATEGORIES]), dim=0)
for t in tickets]
def noise_floor(tickets: list[Ticket], runs: int = 10) -> list[int]:
"""Accuracy of the SAME system over repeated runs with sampling at temperature 1."""
dists = label_distributions(tickets)
scores = []
for seed in range(runs):
g = torch.Generator().manual_seed(seed)
picks = [CATEGORIES[int(torch.multinomial(d, 1, generator=g))] for d in dists]
scores.append(sum(p == t.gold["category"] for p, t in zip(picks, tickets)))
return scores
def bootstrap_ci(passes: list[bool], reps: int = 2000, seed: int = 0) -> tuple[float, float]:
rng = random.Random(seed)
n = len(passes)
means = sorted(sum(rng.choice(passes) for _ in range(n)) / n for _ in range(reps))
return round(means[int(0.025 * reps)], 3), round(means[int(0.975 * reps)], 3)
# CI gate -------------------------------------------------------------------------------
def gate(current: dict[str, list[bool]], reference: dict[str, list[bool]], max_drop: int) -> list[str]:
"""Fail on any HARD assertion failure, or when a soft check loses more than `max_drop` items."""
problems = []
for name, results in current.items():
if name in HARD and not all(results):
problems.append(f"{name}: {results.count(False)} hard failures")
lost = sum(r and not c for r, c in zip(reference[name], results))
if name not in HARD and lost > max_drop:
problems.append(f"{name}: {lost} items went pass -> fail (allowed {max_drop})")
return problems
def run_assertions(tickets, kb, **variant) -> dict[str, list[bool]]:
table: dict[str, list[bool]] = {}
for t in tickets:
for name, ok in assertions(t, extractive_draft(t, kb, **variant)).items():
table.setdefault(name, []).append(ok)
return table
def main() -> None:
torch.set_num_threads(1)
test, kb = load_tickets("test"), KBSearch()
n = len(test)
print(f"== Tier 1: assertions on {n} test tickets (extractive drafter v1)")
v1 = run_assertions(test, kb)
for name, results in v1.items():
k = sum(results)
print(f" {name:24} {k:2}/{n} {'HARD' if name in HARD else 'soft'} 95% CI {wilson(k, n)}")
print(f"\n== Tier 2: judges vs human labels (n={n}, one human rater); trust needs kappa >= {MIN_KAPPA} "
f"and at most {MAX_FALSE_GOOD} false 'good'")
human = json.loads((DATA_DIR / "cap_human_labels.json").read_text())["labels"]
gold = [human[t.id] for t in test]
drafts = [extractive_draft(t, kb) for t in test]
lenient = ScriptedLLM(responder=lambda m, k: '{"label": "good", "reason": "Polite and relevant."}')
judges: dict[str, Judge] = {"LLMJudge(ScriptedLLM always-good)": LLMJudge(lenient), "OverlapJudge": OverlapJudge()}
if "--real" in sys.argv:
judges["LLMJudge(llm.chat)"] = LLMJudge(chat)
for name, judge in judges.items():
c = calibrate(judge, test, drafts, gold)
print(f" {name:34} agree {c['agree']}/{c['n']} kappa {c['kappa']:+.3f} false good {c['false_good']} "
f"false bad {c['false_bad']} says good {c['says_good']} (humans {gold.count('good')}) "
f"{'TRUSTED' if c['trusted'] else 'not trusted'}")
print("\n== Noise floor: prompted TinyLM category, sampled at temperature 1, 10 runs")
scores = noise_floor(test)
print(f" correct per run: {scores} mean {statistics.mean(scores):.1f} sd {statistics.stdev(scores):.2f} "
f"range {min(scores)} to {max(scores)}")
passes = v1["cites_gold_or_escalates"]
print(f" dataset noise: cites_gold_or_escalates {sum(passes)}/{n}, bootstrap 95% CI {bootstrap_ci(passes)}")
print("\n== CI gate: three candidate changes, each compared with v1 (soft checks may lose at most 1 item)")
candidates = {
"v2a retrieval searches the subject only": {"query_mode": "subject"},
"v2b reply keeps 1 sentence instead of 2": {"n_sentences": 1},
"v2c friendlier greeting from marketing": {"greeting": "Hi! We guarantee a fast fix, thanks for writing."},
}
print(f" v1 vs v1: {'PASS' if not gate(v1, v1, max_drop=1) else 'FAIL'}")
for label, variant in candidates.items():
problems = gate(run_assertions(test, kb, **variant), v1, max_drop=1)
print(f" {label}: {'PASS' if not problems else 'FAIL'} {problems}")
judged = [OverlapJudge()(t, extractive_draft(t, kb, n_sentences=1)).label for t in test]
print(f" v2b seen by OverlapJudge: good {judged.count('good')}/{n} "
f"(v1: {[OverlapJudge()(t, d).label for t, d in zip(test, drafts)].count('good')}/{n}); "
f"assertions cannot see a shorter, less complete answer")
if __name__ == "__main__":
main()
Code explained
In simple words: a quality lab with three instruments: a checklist that never gets tired, a taste tester who must first prove they agree with the head chef, and a scale that tells you how much the readings wobble on their own.
What happens:
extractive_draftis the system under test, with three knobs used later to create candidate changes: which text to search with, how many sentences to keep, and the greeting.assertionsruns seven checks per draft: schema valid, citations exist, no forbidden promises (a regex for "guarantee", "we have refunded", and similar), every USD amount in the reply appears in the help center, length, language match (this drafter always writes English), and cites the gold article or escalates.HARDnames the four that must never fail.Verdictand theJudgeprotocol define the judge interface: any callable taking a ticket and a draft and returning a label and a reason.LLMJudgeimplements it over any chat-shaped callable (realllm.chat, ScriptedLLM, or a Meter), with a strict JSON reply and "bad" on anything unparseable.OverlapJudgeis a cheap heuristic judge (English and 3+ shared content words).calibrateruns a judge over labelled items and reports agreement, kappa, false "good" (judge approves what the human rejected, the dangerous direction), false "bad", and whether it clears the trust rule: kappa at least 0.6 and at most one false good.cohen_kappacomputes (observed agreement minus chance agreement) divided by (1 minus chance agreement); the tests cross-check it against scikit-learn.label_distributionsandnoise_floorsample prompted TinyLM's category answers 10 times at temperature 1 (a softmax over the label scores from Part E) to show run-to-run variation of one unchanged system.bootstrap_ciresamples tickets to show dataset noise.gatefails on any hard failure, or when a soft check loses more thanmax_dropitems that passed in the reference.run_assertionsbuilds the pass or fail table for one variant.mainprints both tiers, the noise floor, and the gate on three candidate changes.--realaddsLLMJudge(llm.chat)to the calibration.
- Comes out: deterministic.
What each block tells you.
Tier 1. The four hard checks pass 24 of 24. Language match fails twice (the German and Spanish tickets got English replies), and cites-gold-or-escalates passes 19 of 24. Even a perfect 24 of 24 has a lower bound of 0.862: 24 items cannot prove a failure rate below about 14%.
Tier 2. The always-good judge agrees with the human 14 of 24 times, which sounds acceptable until you see kappa 0.000 and 10 false goods: it simply says yes. The overlap judge does better (kappa 0.486) but still approves 3 drafts the human rejected, so neither is trusted. This is why calibration comes before use: agreement alone would have shipped the useless judge. Run --real to calibrate an LLM judge the same way; if it fails the rule, improve the rubric and examples (Module 10) rather than lowering the bar.
Noise floor. The same prompted classifier, run 10 times with sampling, scores between 2 and 6 of 24 (standard deviation 1.66). A change that moves one sampled run by 2 tickets has shown nothing. The bootstrap interval for 19 of 24 (0.625 to 0.958) is the dataset-side version of the same warning.
The gate. Two of three candidate changes fail, for the right reasons. Searching with the subject only (v2a) loses 2 correct citations, over the allowed drop of 1. The marketing greeting (v2c) contains "guarantee", a hard failure on all 24 drafts. The third (v2b, one sentence instead of two) passes every assertion, and it is a real regression: the replies are less complete, and the overlap judge rates 12 of 24 good against 14 for v1. That difference is within noise and that judge is not trusted, so the honest statement is "the assertion tier cannot see this; a calibrated judge or human review must". Finding changes like v2b is what the judge tier is for, and an Evaluation-track capstone should include at least one of them among its three regressions.
| Situation | Use this | Why |
|---|---|---|
| Format, citations, forbidden content, numbers | Assertions, hard where a failure is never acceptable | Free, deterministic, and exact |
| Helpfulness, completeness, tone | A judge calibrated against two human raters | Code cannot check these; an uncalibrated judge is a guess |
| Setting the allowed drop in a gate | The noise floor from repeated runs and bootstrap | A threshold inside the noise fails at random |
| A judge that approves bad outputs | Count false goods separately from agreement | High agreement can hide a judge that says yes to everything |
Required evidence
- The harness, runnable by a reviewer, with the rubric and the frozen test set.
- Calibration: two raters' agreement (kappa between the humans too), the judge's kappa and false goods, and the trust decision.
- The noise floor and how it set the gate's thresholds.
- Three real regressions: how the harness found each, the fix, and the gate before and after. Plus B1 to B4.
Grading rubric
| Criterion | Weight | Excellent | Adequate | Missing |
|---|---|---|---|---|
| Assertion tier | 15% | Hard and soft checks mapped to real failure modes, with intervals | Checks without the hard and soft split | None |
| Judge calibration | 25% | Two human raters, kappa between them, judge kappa and false goods, a written trust rule | One rater | Judge used uncalibrated |
| Noise floor and gates | 20% | Thresholds derived from measured noise; gate in CI | Thresholds guessed | No gate |
| Three regressions | 25% | Each found by the harness, fixed, and verified; at least one invisible to assertions | Fewer than three, or planted rather than found | None |
| Error analysis, cost, what did not work | 15% | 50+ failures by cause; the harness's own cost per run measured; honest entries | Partial | Missing |
Module Lab
The lab runs the four shared requirements in one pass and writes the evidence file every track starts from. It is the skeleton of your evidence pack: replace the keyword baseline's records with your system's, the stand-in and TinyLM with your real calls, and the pre-filled "did not work" entries with your own.
examples/cap_lab.py
# examples/cap_lab.py
"""Capstone Module Lab: the four shared requirements in one run.
1. baseline keyword triage + BM25 retrieval on the test split, Wilson intervals
2. error analysis all real failures + a SYNTHETIC set, tagged, clustered, report written
3. measurement Meter around ScriptedLLM (priced; plumbing only) and TinyLM (real local latency)
4. write-up what did not work, pre-filled with the measured rule-fix experiments
Writes runs/error_report.md, runs/failures.jsonl, runs/capstone_evidence.md.
Replace each stand-in with your own system; the tooling stays the same.
"""
from __future__ import annotations
import json
import sys
from pathlib import Path
import torch
from supportdesk.data import load_tickets
from supportdesk.stand_in import ScriptedLLM
sys.path.insert(0, str(Path(__file__).resolve().parent))
from cap_baseline import keyword_triage, run_baseline, score # noqa: E402
from cap_error_analysis import analyze, report # noqa: E402
from cap_measure import Meter, TinyChat, check_budgets # noqa: E402
from cap_rule_fixes import FIXES, compare # noqa: E402
from cap_synthetic_tickets import build # noqa: E402
RUNS = Path("runs")
BUDGET_P95_MS, BUDGET_COST = 250, 0.0005
def main() -> None:
torch.set_num_threads(1)
RUNS.mkdir(exist_ok=True)
lines = ["# Capstone evidence (generated by examples/cap_lab.py)", ""]
# 1. Baseline to beat
real_records = run_baseline(load_tickets())
lines += ["## 1. Baseline to beat (test split, n=24)", "", "| Task | Score | 95% CI |", "|---|---|---|"]
for task in ("category", "priority", "retrieval_top1"):
s = score(real_records, task, "test")
lines.append(f"| {task} | {s['k']}/{s['n']} = {s['acc']:.3f} | {s['ci95']} |")
print(f"[1] baseline {task:15} {s['k']}/{s['n']} {s['ci95']}")
# 2. Error analysis
overrides_path = RUNS / "tag_overrides.json"
overrides = json.loads(overrides_path.read_text()) if overrides_path.exists() else {}
result = analyze(real_records + run_baseline(build()), overrides)
(RUNS / "error_report.md").write_text(report(result) + "\n", encoding="utf-8")
with (RUNS / "failures.jsonl").open("w", encoding="utf-8") as f:
for row in result["failures"]:
f.write(json.dumps(row, ensure_ascii=False) + "\n")
real = [f for f in result["failures"] if f["features"]["source"] == "real"]
real_tags = {}
for f in real:
for tag in f["tags"]:
real_tags[tag] = real_tags.get(tag, 0) + 1
top = sorted(real_tags.items(), key=lambda kv: -kv[1])[:5]
print(f"[2] failures {len(result['failures'])} ({len(real)} real, {len(result['failures']) - len(real)} SYNTHETIC); "
f"top real causes {top}")
lines += ["", "## 2. Error analysis", "",
f"{len(real)} real failures (plus {len(result['failures']) - len(real)} on SYNTHETIC tickets). "
"Top real causes (full report: runs/error_report.md):", ""]
lines += [f"- {tag}: {n}" for tag, n in top]
# 3. Cost and latency, measured
test = load_tickets("test")
responder = lambda m, k: json.dumps(dict(zip(("category", "priority"), keyword_triage(m[-1]["content"])))) # noqa: E731
scripted = Meter(ScriptedLLM(responder=responder), price_as="openai/gpt-oss-120b")
tiny = TinyChat(max_new_tokens=24)
tiny([{"role": "user", "content": "warm-up"}])
local = Meter(tiny)
for t in test:
scripted([{"role": "user", "content": t.text}], task_id=t.id)
local([{"role": "user", "content": t.body}], task_id=t.id)
s1, s2 = scripted.summary(), local.summary()
broken = check_budgets(s2, BUDGET_P95_MS, None)
print(f"[3] ScriptedLLM priced as gpt-oss-120b: {s1['cost_per_task_mean']:.8f} USD/task, "
f"{s1['tokens_per_task_mean']} tokens/task | TinyLM p50 {s2['call_ms_p50']:.0f} ms, "
f"p95 {s2['call_ms_p95']:.0f} ms, budget broken: {broken or 'no'}")
lines += ["", "## 3. Cost and latency (measured)", "",
"| System | Calls | p50 ms | p95 ms | Tokens/task | Cost/task USD |", "|---|---|---|---|---|---|",
f"| ScriptedLLM (plumbing only, priced as gpt-oss-120b) | {s1['calls']} | {s1['call_ms_p50']} | "
f"{s1['call_ms_p95']} | {s1['tokens_per_task_mean']} | {s1['cost_per_task_mean']:.8f} |",
f"| TinyLM local, 24 new tokens | {s2['calls']} | {s2['call_ms_p50']} | {s2['call_ms_p95']} | "
f"{s2['tokens_per_task_mean']} | local compute, not priced |",
"", f"Budgets: task p95 <= {BUDGET_P95_MS} ms, cost <= {BUDGET_COST} USD per task. "
f"TinyLM broken budgets: {broken or 'none'}."]
# 4. What did not work
lines += ["", "## 4. What did not work", ""]
for name in FIXES:
r = compare(name)
d, t = r["dev"], r["test"]
verdict = "within noise, not adopted as a win" if max(d["p"], t["p"]) > 0.05 else "significant"
lines.append(f"- Tried: {name}. Measured: dev {d['before']} -> {d['after']}/{d['n']} "
f"(fixed {len(d['fixed'])}, broke {len(d['broke'])}); test {t['before']} -> {t['after']}/{t['n']} "
f"(fixed {len(t['fixed'])}, broke {len(t['broke'])}). Verdict: {verdict}. "
"Why it failed: WRITE THIS FROM THE FLIPPED TICKETS.")
print(f"[4] {name}: dev {d['before']}->{d['after']}, test {t['before']}->{t['after']}, {verdict}")
(RUNS / "capstone_evidence.md").write_text("\n".join(lines) + "\n", encoding="utf-8")
print("wrote runs/capstone_evidence.md and runs/error_report.md")
if __name__ == "__main__":
main()
Code explained
- In simple words: one button that regenerates every shared piece of capstone evidence from the code, so the numbers in your write-up can never drift from what the code does.
- What happens:
- Section 1 runs the baseline over all 72 real tickets and writes the test-split table with Wilson intervals.
- Section 2 adds the synthetic tickets, applies the reviewer overrides, writes the full report and the failures file, and lists the top five real causes.
- Section 3 meters one ScriptedLLM call per test ticket (priced as
openai/gpt-oss-120b, plumbing only) and one real TinyLM call per ticket after a warm-up, and checks a 250 ms p95 budget. - Section 4 reruns both rule-fix experiments from B4 and writes each as a pre-filled entry. It calls a result significant only if both splits give p at most 0.05, and leaves the "why" for a person to write from the flipped tickets.
- Everything goes to
runs/capstone_evidence.md.
- Comes out: about 3 seconds; everything but the latencies is deterministic.
The evidence file it writes:
# Capstone evidence (generated by examples/cap_lab.py)
## 1. Baseline to beat (test split, n=24)
| Task | Score | 95% CI |
|---|---|---|
| category | 16/24 = 0.667 | (0.467, 0.82) |
| priority | 10/24 = 0.417 | (0.245, 0.612) |
| retrieval_top1 | 17/20 = 0.850 | (0.64, 0.948) |
## 2. Error analysis
69 real failures (plus 170 on SYNTHETIC tickets). Top real causes (full report: runs/error_report.md):
- question_rule: 20
- no_signal: 15
- non_english: 14
- vocab_gap: 9
- urgency_unseen: 8
## 3. Cost and latency (measured)
| System | Calls | p50 ms | p95 ms | Tokens/task | Cost/task USD |
|---|---|---|---|---|---|
| ScriptedLLM (plumbing only, priced as gpt-oss-120b) | 24 | 0.13 | 0.18 | 42.4 | 0.00001208 |
| TinyLM local, 24 new tokens | 24 | 22.83 | 25.26 | 57.0 | local compute, not priced |
Budgets: task p95 <= 250 ms, cost <= 0.0005 USD per task. TinyLM broken budgets: none.
## 4. What did not work
- Tried: A word boundaries (category). Measured: dev 31 -> 31/48 (fixed 1, broke 1); test 16 -> 16/24 (fixed 1, broke 1). Verdict: within noise, not adopted as a win. Why it failed: WRITE THIS FROM THE FLIPPED TICKETS.
- Tried: B no question rule (priority). Measured: dev 28 -> 28/48 (fixed 12, broke 12); test 10 -> 11/24 (fixed 6, broke 5). Verdict: within noise, not adopted as a win. Why it failed: WRITE THIS FROM THE FLIPPED TICKETS.Code explained
- In simple words: the first page of your evidence pack, generated rather than typed.
- What happens: each section comes from the code path above; the only text a person must add is the "why" of each failed attempt, deliberately left as a marker so an unfinished write-up is obvious.
- Comes out: a file you commit next to your code and regenerate after every change.
Finally, the tests pin the tooling's behavior, so a change to a shared tool that breaks a promise is caught before it corrupts your evidence.
PYTHONPATH=. python -m pytest -q tests/test_cap_tooling.pyCode explained
- In simple words: a quick check that the measuring instruments still read true.
- What happens: 13 tests cover known Wilson values, the substring bug the error analysis found, at least 50 real failures, the
question_ruletag on the T-1063 pattern, nearest-rank percentiles, Meter pricing and the cached-price warning, kappa against scikit-learn, the eval gate's hard and soft logic, the sign test, the product gate catching v2's schema break, and the agent blocking the injected refund and escalating under guards v2. No API key and no training are needed. - Comes out: the time varies by machine.
Project Milestone
The capstone repository (work/cap) now contains the shared tooling and one starter per track. The canonical supportdesk/ package is unchanged; everything imports it.
| File | What it adds |
|---|---|
examples/cap_baseline.py | Keyword triage and BM25 baselines, Wilson intervals, the paired sign test, the run-record format |
examples/cap_synthetic_tickets.py | 100 labelled SYNTHETIC tickets in seven variation kinds |
examples/cap_error_analysis.py | Cause taxonomy, auto-tagger, overrides, clusters by tag and feature, Markdown report |
examples/cap_measure.py | Meter, percentiles, budgets, cached-price warning, TinyChat adapter |
examples/cap_measure_demo.py | Measured cost per task (ScriptedLLM, priced) and real TinyLM latency |
examples/cap_rule_fixes.py | The B4 worked example: two fixes measured with flipped tickets |
examples/cap_product_starter.py | Prompt registry with hashes, schema validation with retry, eval and CI gate |
examples/cap_agent_starter.py | Guarded tool loop, approval gate, citation guard, tracing, trajectory eval, attack suite |
examples/cap_adapt_starter.py | TinyLM fine-tune with validation-only recipe selection, regression budget, paired tests |
examples/cap_eval_starter.py | Assertion tier, judge interface and calibration, noise floor, gate on three candidate changes |
examples/cap_lab.py | All four shared requirements in one run |
tests/test_cap_tooling.py | 13 tests pinning the tooling |
data/cap_human_labels.json | One rater's labels for the extractive drafter on the 24 test tickets |
runs/tag_overrides.json | The 11 reviewer corrections to auto tags |
Deliverables checklist
Every track hands in the shared items; each track adds its own.
| Shared (all tracks) | Done when |
|---|---|
| Frozen test split and README | A reviewer can rerun every number with one command per section |
| B1 baseline | Numbers with n and 95% intervals, including a majority floor and the strongest cheap baseline |
| B2 error analysis | 50+ real failures of your system, read by a person, taxonomy of causes, clusters by tag and feature; synthetic data labelled and separate |
| B3 cost and latency | Meter logs from the real system, p50 and p95 per task, cost per task, budgets, machine and load noted |
| B4 what did not work | 3 to 6 entries with measurements and flipped items |
| Final comparison | Your system against the baseline on test, paired wins and losses, sign test p, stated as a win only when it is one |
| Track | Track-specific deliverables |
|---|---|
| Product | Prompt registry with hashes; gate output for 3+ versions including a blocked change; CI configuration running the gate; safety review with tests; fallback path |
| Agent | Tool schemas and guards; approval flow; attack suite of 10+ items with pulled, attempted, executed counts per guard version; clean and attacked traces; trajectory eval on test with the real model |
| Adaptation | Plateau evidence across prompt variants; dataset card; run log with validation-only selection; single test run with paired tests; regression report against a budget set in week 1; serving cost of tuned versus prompted |
| Evaluation | Harness README; two-rater labels with kappa between raters; judge calibration with false goods; noise floor and derived thresholds; three regressions found, fixed, and verified |
The capstone defense
Plan a 15-minute presentation and 15 minutes of questions. A structure that works:
- The claim, in one sentence, and the number that supports it with its interval.
- The baseline and why it is the right one to beat.
- The two biggest failure clusters, with one real example each, and what you did about them.
- Cost and latency per task against budget.
- What did not work, and what you would try next.
- For your track: the gate blocking a bad change, the attacked trace, the regression report, or a regression the harness caught.
Interview Questions
These are the questions a capstone reviewer, or an interviewer reading your project, is most likely to ask.
1. Your system scores 19 of 24 and the baseline 16 of 24. Is it better? Not on that evidence alone. The Wilson intervals are about 0.6 to 0.9 and 0.47 to 0.82, and they overlap heavily. Compare on the same tickets instead: count the tickets your system fixes and the ones it breaks, and run the sign test on those two numbers. Three fixes and zero breaks gives p = 0.125, still not significant; you would need more test items or a larger, consistent gap. The honest claim is "no worse, possibly better", and the next step is a bigger test set.
2. Why 50 failures, and why real ones? With fewer than about 50, the ranking of clusters is mostly chance: a cause with 4 failures and one with 6 are indistinguishable. Real failures matter because they reflect your traffic. In the baseline analysis, synthetic tickets produced 40 misspelling failures and real tickets produced none; ranking by all failures would have sent a week of effort to a problem customers did not have. Synthetic data is for finding hypotheses, not for setting priorities.
3. What makes a good cause tag? It names something you can act on and predicts which fix will help. "Priority too low" is a symptom; "the short-question rule fired on an urgent ticket" is a cause, and it tells you where to look. Keep the taxonomy small (10 to 15 tags), let a person override the auto-tagger, allow two tags when two causes contributed, and include a "label question" tag so doubtful gold labels are counted rather than silently fixed.
4. How did you measure cost, and how is that different from estimating it? Every call went through a wrapper that recorded the provider-reported usage (input, output, cached, and reasoning tokens) and wall-clock time, priced with a single price table, and grouped by task. That captures retries (the v2 prompt doubled the calls and nearly doubled cost per task), reasoning tokens, and multi-call tasks, all of which an estimate misses. Latency is reported as p50 and p95 per task over several runs, with the machine and load noted, because the same TinyLM script measured p50 23 ms on a quiet machine and 106 ms under load.
5. Tell me about something that did not work. Word-boundary matching looked like the fix for keywords matching inside words ("plan" in "plane"). It fixed one ticket per split and broke one per split: "error" stopped matching "errors" and "cancel" stopped matching the Spanish "cancelar", which had been right by accident. Net zero, p = 0.75, dropped. The lesson was to read the flipped tickets, not the net score, and to put effort where the real clusters were largest.
6. (Product) How do you stop a prompt edit from breaking production? Treat prompts as versioned data with a content hash, run every change through the same gate as code: schema validity must be 100%, the paired comparison with the current version or baseline must not lose beyond noise, and cost and p95 latency must stay within budget. The gate exits non-zero so CI blocks the merge. In the starter, a "tidy-up" that removed one line broke validity on every ticket and doubled cost through retries, and the gate caught both.
7. (Agent) How do you know your agent is safe against prompt injection? I assume the model will sometimes follow injected text, so safety lives in code: arguments of irreversible tools must trace to the customer's own ticket, amounts have a hard ceiling, a human approves every refund, drafts may cite only reviewed articles, and loops have step, token, and repeat limits. Then I test it with an attack suite and a worst-case policy that obeys every instruction. Under the injected article, no refund executed, and the trajectory eval found a second problem (the draft cited the attacker's article), which led to the citation guard. I report attacks pulled, attempted, and executed separately.
8. (Adaptation) Your fine-tune did not beat the prompted baseline. Is that a failed project? No, if the experiment was clean. Recipes were chosen on a validation slice, test was scored once, a regression budget was set in advance, and the paired test gave p = 0.91. The model memorized its 32 training tickets while validation stayed at 2 to 7 of 16, and replay kept forgetting to about 5% instead of about 13 times. The conclusion ("with this model and this much data, do not fine-tune; collect labels or use a stronger base") is useful and defensible. An unclean "win" is worth less than a clean "no".
9. (Evaluation) How do you know your LLM judge can be trusted? Calibrate before use: two humans label a sample with a written rubric, I measure their agreement, then the judge's Cohen's kappa against them and, separately, its false goods (outputs it approves that humans rejected). A judge that says "good" to everything agreed with our rater on 14 of 24 but had kappa 0 and 10 false goods. My rule was kappa at least 0.6 and at most one false good on the sample; below that, fix the rubric and examples rather than lower the bar, and recalibrate whenever the judge model or prompt changes.
10. What is a noise floor and how did you use it? It is how much a score moves when nothing changed: rerunning a sampled system, or resampling the test set. The sampled TinyLM classifier ranged from 2 to 6 of 24 across 10 runs, and a bootstrap put 19 of 24 anywhere from 0.63 to 0.96. I set gate thresholds above that wobble (for example "a soft check may lose at most 1 item" only where repeated runs of the unchanged system lost none), so the gate fails for real changes rather than at random.
11. Your gate passed a change that made answers worse. How? Assertions check format, citations, forbidden content, and numbers; they cannot see completeness. Cutting replies from two sentences to one passed every assertion. Catching it needs a calibrated judge or human review on a sample, and a gate that includes the judge tier once it is trusted. The broader lesson: every tier has blind spots, so list what each one cannot see.
12. If you had another month, what would you do? Enlarge the test set, because with 24 items only large gains are provable; collect real labels for non-English tickets, since Japanese and Spanish records failed 5 of 6 and 4 of 6 times; move the urgency decision from punctuation to meaning (an LLM or a classifier trained on real priorities), since that was the largest real cluster; and run the whole evidence pack in CI so every future change regenerates it.
Other Tools and Providers
The capstone used hand-built tools so every mechanism is visible. These are the common alternatives; check each project's current status before adopting it.
| Area | What the capstone used | Alternatives | When to prefer them |
|---|---|---|---|
| Evals and CI gates | cap_eval_starter.py, cap_product_starter.py, pytest | promptfoo (open source; joined OpenAI in March 2026 and states it "will remain open source", announcement), Inspect (UK AI Security Institute), DeepEval, OpenAI Evals, Ragas for retrieval metrics | Declarative test suites, many providers side by side, built-in judge templates; keep your own golden set and thresholds either way |
| Tracing and observability | JSONL traces and Meter logs | Langfuse, Arize Phoenix, LangSmith, Braintrust, Weights and Biases Weave, OpenTelemetry with its generative AI conventions | Several features and teams, trace search, dashboards, and dataset versioning |
| Human labelling | One rater, a JSON file | Label Studio, Argilla | Two or more raters with assignment, blinding, and agreement reports |
| Fine-tuning | Full fine-tuning of TinyLM in plain PyTorch | Hugging Face PEFT and TRL, Unsloth, Axolotl, LLaMA-Factory; managed tuning on Google Vertex AI, Together AI, Fireworks AI (Module 9 has current caveats per provider) | Any real open-weight model; LoRA on a single GPU is the usual starting point |
| Guardrails | Code guards, approval gate, citation guard | NVIDIA NeMo Guardrails, Guardrails AI, Llama Guard 4 and Llama Prompt Guard 2, OpenAI gpt-oss-safeguard, Microsoft Presidio for PII | Broader content and injection classifiers as an extra layer; never a replacement for authorization in code |
| Serving | Hosted APIs through llm.chat, local TinyLM | vLLM, SGLang, llama.cpp, Ollama, NVIDIA TensorRT-LLM | Self-hosting a tuned model: throughput, batching, multi-LoRA serving (Module 13) |
| Statistics | Our Wilson, sign test, bootstrap, kappa | scipy.stats (binomtest), statsmodels (proportion_confint, mcnemar), scikit-learn (cohen_kappa_score) | Production analysis with tested library code; ours is cross-checked against scikit-learn in the tests |
Course Wrap-Up
You started this course with a 1-million-parameter model that answered "The capital of France is" with support-desk nonsense, and with one claim: that a large language model only ever predicts the next token. Everything since has been about living with that fact productively. You learned to read what the model is (Modules 1 to 3), to direct it with prompts, reasoning, structure, and context (Modules 4 to 7), to give it tools and let it act within limits (Module 8), to adapt it (Module 9), to measure it (Module 10), to secure it (Module 11), to extend it to images and documents (Module 12), to run it in production at a known cost (Module 13), and to decide when it should not be used at all (Module 14).
The capstone adds the habit that ties those skills together: every claim comes with a baseline, an interval, real failures, a measured cost, and a record of what did not work. That habit outlasts every model name and price in this course. The models you use next year will be different; Brightlane's 24 test tickets would still say only what 24 tickets can say.
Three suggestions for what comes next:
- Keep the evidence pack alive. Run
cap_lab.py(or your version of it) in CI, so every prompt, model, or data change regenerates the numbers and the gate decides. - Go deeper where your capstone was weakest. If your intervals were wide, study sample size and power (Module 10). If your agent needed many guards, study authorization design (Module 11). If you want to know why attention and context behave as they do, the companion Transformers course covers the internals this course deliberately skipped.
- Read new releases the way Module 14 taught. Run your own eval, through your own gate, at your own cost per task, before believing anyone's benchmark.
Maya's team now has an assistant they can check. That is the real deliverable of this course: not a model that is always right, but a system whose mistakes you can find, count, price, and fix.