Part A: Why evaluation is the bottleneck
By the end of this module, you'll have:
- A working eval harness for the Brightlane assistant (
examples/m10_evals.py): 91 golden cases (the 72 tickets plus 19 curated hard cases in five languages), a system-under-test interface, deterministic checks for schema, category, citation, reply language, and forbidden content (roadmap dates, claimed refunds or unlocks), per-case JSON, and a summary with Wilson confidence intervals. - Real measurements of three systems: a keyword baseline (41.8% of cases pass), the same baseline after error analysis (60.4%), and the noise floor of two stochastic systems (TinyLM sampling and a seeded noisy stand-in, both about 2 points of standard deviation run to run).
- The statistics to read those numbers honestly: a power calculation that says detecting a 5-point difference needs hundreds of cases, and paired comparisons (exact McNemar and a paired bootstrap) that show the v2 gain is real overall but not yet shown on the held-out split.
- LLM-as-judge prompts (pointwise, pairwise, reference-based) that run on any provider through
supportdesk.llm.chat, a swap harness for position and verbosity bias, and a calibration against 30 hand labels with Cohen's kappa. - System-level evals: component metrics against end-to-end, failure attribution that blames retrieval before tools before output, trajectory checks, and a multi-turn eval that catches a truncated-history bug.
- A pytest CI gate on quality, cost, and latency with a pinned baseline file, which you will see block a real change and then pass, plus online signals, a simulated A/B test with real statistics, and a loop that turns a production incident into a regression case.
Prerequisites: Modules 1 to 8. You need supportdesk.llm.chat and ScriptedLLM (Module 1), token counting and pricing.py (Module 2), sampling and seeds (Module 3), the dev/test discipline from Module 4, the Triage and DraftReply schemas (Module 6), KBSearch (Module 7), and the agent loop and trajectory idea (Module 8). Working Python, no statistics background: every formula is explained and computed in code.
Where we are: Module 9 asked when to change the model itself, and its verification step insisted on a held-out evaluation designed before any training. That rule applies to every change, not only fine-tuning. So far the course has measured things one example at a time, with small ad hoc test sets. This module builds the evaluation system the rest of the course runs on: the golden set, the checks, the statistics, the judges, and the CI gate that decides whether a prompt, model, or code change ships.
How this module is organized
| Part | What it covers |
|---|---|
| Part A: Why evaluation is the bottleneck | Non-determinism, subjectivity, moving targets; what vibes-based development costs, measured; writing the eval before the feature |
| Part B: Constructing evals | Sourcing real inputs and curating hard cases, success criteria, deterministic checks, the harness, golden datasets and regression suites |
| Part C: Noise, sample size, and comparing systems | Wilson intervals, the noise floor, power calculations, error analysis before optimization, paired comparisons |
| Part D: Human evaluation | When humans are unavoidable, rating design, agreement (Cohen's kappa), adjudication |
| Part E: LLM-as-judge | Rubrics and scales, pointwise, pairwise, and reference-based judges, bias and the swap test, calibration against human labels, judge cost |
| Part F: System-level evaluation | Component vs end-to-end, failure attribution, agent trajectories, multi-turn conversations |
| Part G: Operating evals | CI gates on quality, cost, and latency; online signals; A/B tests; feeding production failures back |
Part D comes before Part E on purpose: you cannot trust a judge until you have human labels to check it against.
A word on what is real here. No LLM API key was available when this module was built, so every number in "Comes out" was produced by code that runs without one: a keyword baseline, TinyLM (the 1.07M-parameter model from Module 1), KBSearch, seeded simulations, and ScriptedLLM stand-ins. Each is labelled. Every harness also runs against a real model through supportdesk.llm.chat, and the text gives you the exact command. Where a sample of real-model output appears, it is marked as illustrative.
All examples run in order from the root of your supportdesk copy. Setup once:
cd supportdesk
source .venv/bin/activate # or however you activated the course venv in Module 1
pip install scikit-learn==1.9.1 # only used to cross-check our kappa in the tests
export PYTHONPATH=.
mkdir -p evals runs
python -c "import torch; torch.set_num_threads(1); import supportdesk.kb_search, supportdesk.pricing, supportdesk.schemas; print('ok')"
Code explained
- In simple words: step into the project, add one pinned library, make the package importable, and create folders for eval data (
evals/) and run results (runs/). - What happens:
PYTHONPATH=.letsexamples/scripts import bothsupportdesk.*and each other (examples.m10_stats). scikit-learn is optional; the tests use it only to check our own Cohen's kappa. The last line imports the canonical modules this module builds on. Scripts that use TinyLM calltorch.set_num_threads(1)themselves, which keeps runs reproducible and polite on a busy machine. - Comes out:
ok. If you seeModuleNotFoundError: supportdesk, you are not in the repository root or forgot theexport.
The module creates these files. Data files you type in are shown in full where they first appear; the rest are generated by the scripts.
| File | What it is |
|---|---|
examples/m10_stats.py | Wilson interval, exact McNemar, paired bootstrap, power, two-proportion test, Cohen's kappa, percentile |
examples/m10_evals.py | The harness: cases, system-under-test interface, checks, runs, summaries, comparisons, noise floor, error analysis, CLI |
examples/m10_why.py, m10_power.py | Part A and Part C demonstrations |
examples/m10_human.py, m10_judge.py | Human labels and agreement; judge prompts, bias harness, calibration, cost |
examples/m10_system.py, m10_ops.py, m10_lab.py | System-level evals; online signals, A/B, feedback; the lab |
evals/hard_cases.jsonl, labels.jsonl, gate.json | Curated cases, 30 labelled drafts, gate thresholds (typed in) |
evals/golden.jsonl, drafts.jsonl, regressions.jsonl, baseline.json | Generated: the frozen golden set, drafts to label, the regression suite, the pinned baseline |
tests/test_m10_evals.py, tests/test_m10_gate.py | Unit tests for the harness; the CI quality gate |
Part A: Why evaluation is the bottleneck
Three reasons LLM features are hard to test
Ordinary software has a spec and a unit test: add(2, 2) returns 4 or it is broken. An LLM feature breaks that model in three ways.
- Non-determinism. The same input can give different outputs. Module 3 showed why: sampling draws from a probability distribution, and even at temperature 0 hosted providers do not guarantee identical output across runs. One good answer proves little.
- Subjectivity. Many outputs have no single right answer. Two good replies to "can I get a refund?" can use different words, and two careful agents can disagree about whether a reply is good enough to send. Correctness becomes a judgment, and judgments need a written standard.
- Moving targets. The provider updates the model, the help center changes, the product adds a plan, the ticket mix shifts toward a new language. A system that passed last month can fail this month without a single line of your code changing.
Put together: you cannot look at an LLM feature and know whether it works. You have to measure it on many inputs, repeatedly, against written criteria. That measurement, not the prompt or the model, is usually what limits how fast a team can improve an LLM product. Teams with a trusted eval can try ten ideas a day and keep the two that help. Teams without one argue.
Vibes-based development, measured
Vibes-based development means judging a change by trying a few inputs by hand and deciding it "looks good". It is how almost every LLM feature starts, and it fails in a predictable way: you try the inputs you thought of, which are the easy ones. The script below measures both problems on real systems. It samples TinyLM five times on two questions, then scores the course's keyword baseline (built in Part B) on three hand-picked tickets and on the full golden set.
"""Module 10, Part A: why evaluation is the bottleneck, measured.
PYTHONPATH=. python examples/m10_why.py
1. Non-determinism: TinyLM (a real 1M-parameter model) answers the same ticket
differently on every sampled run.
2. Vibes: three tickets you would try by hand look perfect; the full suite does not.
"""
from __future__ import annotations
import torch
from examples.m10_evals import GOLDEN, build_golden, load_cases, make_baseline, run_eval
from supportdesk import tinylm
torch.set_num_threads(1)
model, tokenizer = tinylm.load()
questions = {"seen in training": "can I get a refund on my annual Team plan?",
"a real ticket (T-1004)": "We renewed our annual Business plan 5 days ago by mistake. Can we get our money back?"}
for label, question in questions.items():
print(f"{label}: five sampled runs at temperature 0.7")
for seed in range(5):
params = tinylm.SamplingParams(max_new_tokens=40, temperature=0.7, seed=seed, stop=["\n"])
text = tinylm.generate(model, tokenizer, f"Customer (Ana): Hi, {question}\nAgent (Lena):", params).text
print(f" seed {seed}: {text.strip()[:88]}")
if not GOLDEN.exists():
build_golden() # the golden file is explained in Part B
system = make_baseline(1)
cases = load_cases()
by_id = {c.id: c for c in cases}
hand_picked = [by_id[i] for i in ("T-1004", "T-1011", "T-1019")]
vibes = run_eval(system, hand_picked, "vibes", save=False)["summary"]
full = run_eval(system, cases, "full", save=False)["summary"]
print(f"\nhand-picked tickets: {vibes['passed']}/{vibes['n']} pass")
print(f"full golden set: {full['passed']}/{full['n']} pass ({full['pass_rate']:.1%}, "
f"95% CI {full['ci95'][0]:.1%} to {full['ci95'][1]:.1%})")
Code explained
- In simple words: a two-part experiment: sample the same model repeatedly, then compare a hand-picked check with the full suite.
- What happens: the first loop uses
tinylm.generatewithSamplingParams(max_new_tokens=40, temperature=0.7, seed=seed, stop=["\n"])in the chat format TinyLM was trained on (Customer (Ana): ...thenAgent (Lena):). The second part imports the harness from Part B and runs the same system on 3 cases and on 91. - Comes out: see the run below.
python examples/m10_why.py
Code explained
- In simple words: the same model asked the same thing five times, then the difference between "I tried three tickets and they worked" and "I ran all 91".
- What happens: the script loads TinyLM, samples 40 tokens at temperature 0.7 with seeds 0 to 4 for a question that appears almost verbatim in its training corpus and for real ticket T-1004. Then it builds the golden set if it is missing (Part B explains it) and runs the version 1 keyword baseline on three tickets a developer might try by hand (a refund, an export question, a login question) and on all 91 cases, with
run_evalfrom the harness. - Comes out: (TinyLM text is real model output; the script takes a few seconds)
What does vibes-based development cost? Three things, all visible later in this module. You ship regressions you never saw: the version 3 change in Part G looks harmless and breaks three unanswerable cases. You chase noise: TinyLM's pass rate moves between 11.0% and 17.6% with no change at all (Part C), so a "6-point win" from one run can be nothing. And you cannot answer the question every manager eventually asks, "did the new model make it better?", with anything but an anecdote.
Build the eval before the feature
The discipline this module teaches is simple to state: write the eval before you write the feature. Before changing the prompt, decide which cases the change should fix, which must not break, and what number counts as success. Then make the change and let the eval decide. This is test-driven development with statistics.
| Situation | Use this | Why |
|---|---|---|
| First week of a feature, no data yet | 20 to 50 hand-written cases plus the deterministic checks from Part B | Catches format and policy failures immediately; cheap to write |
| Real traffic exists | Sample real inputs into the golden set, label them, add curated hard cases | Your users' inputs are the distribution that matters |
| Comparing two prompts or models | Paired run on the same frozen cases, with McNemar or a paired bootstrap | Pairing removes case difficulty from the comparison (Part C) |
| Quality is a judgment (tone, helpfulness) | A rubric, human labels, then a calibrated LLM judge | A judge you have not checked against humans is a guess (Parts D and E) |
| Every change after launch | A CI gate plus a regression suite | Nothing that was fixed breaks again silently (Part G) |