Part B: Build the test set before you tune
Why the order matters
If you write a prompt, try it on a few tickets, tweak it, and try again, you are tuning on whatever tickets you happened to look at. The prompt gets better at those tickets and you have no idea whether it got better at tickets in general. The fix is an order of work you do not break:
The Brightlane dataset already has a dev split (48 tickets) for tuning and a test split (24 tickets) for the final check. What it does not have is a gold needs_human label, even though the Triage schema in supportdesk/schemas.py defines one. So step one is labeling.
Step 1: fill in the missing label, from the definition
The schema's definition is: "True if an agent must act (refund, account change, legal, security) or the help center cannot answer it." The dataset marks each ticket answerable (the help center can answer it), which covers the second half. For the first half, the tickets that are answerable but still need an agent to do something are listed by hand, with a reason for each.
{
"definition": "True when an agent must act (refund, account change, legal, security) or the help center cannot answer the ticket. Same wording as supportdesk/schemas.py Triage.needs_human.",
"rule": "needs_human = (not gold.answerable) or (ticket id in agent_action)",
"agent_action": {
"T-1001": "refund of a duplicate charge",
"T-1004": "refund inside the 14-day window",
"T-1011": "all users of an enterprise locked out: security and access incident",
"T-1024": "GDPR erasure request: legal",
"T-1027": "SLA credit: money owed",
"T-1031": "refund of a duplicate charge (Spanish)",
"T-1041": "past invoices are reissued by support on request"
},
"annotator": "one annotator (the module author), labeled before any prompt was written; treat as noisy"
}
Code explained
- In simple words: a small label file that turns a written definition into a rule plus a short list of exceptions, each with its reason.
- What happens:
gold_labelsinm04_prompts.pycomputesneeds_human = (not answerable) or (id in agent_action). Tickets such as a locked account are not listed because the help-center answer is "wait 15 minutes; support cannot unlock it sooner", which needs no agent action. Theannotatorfield records that one person labeled these before writing any prompt. - Comes out: 11 of 48 dev tickets are
needs_human: true(6 unanswerable ones plus 5 answerable ones from the list; the other two listed tickets are in test, where 6 of 24 are true). With one annotator, a few of these calls are arguable; that noise caps how high any prompt can score on this label, which matters for ceiling detection in Part E.
Labeling first also forces you to decide what the label means. If you cannot write down why T-1011 (all users locked out of SSO) needs a human, the model will not guess it either.
Step 2: freeze the sets
A frozen test set is one whose ticket ids and gold labels are recorded with a hash, so any later change (a relabel, a ticket added, a different split) is detected instead of silently changing your scores.
examples/m04_eval.py
"""Module 4: a test-set harness for triage prompts.
Order of work this file enforces:
1. freeze write evals/triage_<split>_manifest.json (ids + hash of gold labels) BEFORE tuning
2. run run one prompt version over the frozen split, save every prediction under runs/
3. compare per-label accuracy with 95% Wilson intervals, paired bootstrap between versions
Backends:
rules keyword baseline wrapped in ScriptedLLM (a real, dumb classifier; not a model)
copy ScriptedLLM that copies the majority label of the few-shot examples in the prompt
mixed ScriptedLLM replaying well-formed and malformed replies (plumbing test only)
llm supportdesk.llm.chat with your provider (needs a key or a local Ollama)
Examples:
PYTHONPATH=. python examples/m04_eval.py freeze --split dev
PYTHONPATH=. python examples/m04_eval.py run --version v2 --backend rules
LLM_PROVIDER=groq PYTHONPATH=. python examples/m04_eval.py run --version v3 --backend llm --shots balanced --k 6
PYTHONPATH=. python examples/m04_eval.py compare runs/a.jsonl runs/b.jsonl
"""
from __future__ import annotations
import argparse
import hashlib
import html
import json
import math
import random
import re
import time
from collections import Counter
from dataclasses import dataclass, field
from pathlib import Path
from supportdesk.data import load_tickets
from supportdesk.stand_in import ScriptedLLM
from examples.m04_prompts import (EVAL_DIR, LABELS, ROOT, ExampleSelector, arrange, gold_labels,
load_prompt, parse_triage)
RUN_DIR = ROOT / "runs"
# Step 1: freeze the test set before tuning -------------------------------------------
def manifest_path(split: str) -> Path:
return EVAL_DIR / f"triage_{split}_manifest.json"
def labels_digest(split: str) -> tuple[list[str], str]:
tickets = load_tickets(split)
payload = json.dumps([[t.id, gold_labels(t)] for t in tickets], sort_keys=True)
return [t.id for t in tickets], hashlib.sha256(payload.encode()).hexdigest()
def freeze_eval_set(split: str) -> dict:
"""Record exactly which tickets and gold labels the prompt will be judged on."""
path = manifest_path(split)
if path.exists():
raise FileExistsError(f"{path.name} already exists; a frozen set is never rewritten silently.")
ids, digest = labels_digest(split)
manifest = {"split": split, "n": len(ids), "ids": ids, "labels_sha256": digest,
"frozen_at": time.strftime("%Y-%m-%d %H:%M:%S")}
path.write_text(json.dumps(manifest, indent=2) + "\n", encoding="utf-8")
return manifest
def verify_manifest(split: str) -> dict:
path = manifest_path(split)
if not path.exists():
raise RuntimeError(f"No frozen eval set for {split!r}. Run: python examples/m04_eval.py freeze --split {split}")
manifest = json.loads(path.read_text(encoding="utf-8"))
ids, digest = labels_digest(split)
if ids != manifest["ids"] or digest != manifest["labels_sha256"]:
raise RuntimeError(f"The {split} set changed after it was frozen. Scores are no longer comparable.")
return manifest
# Statistics ----------------------------------------------------------------------------
def wilson(k: int, n: int, z: float = 1.96) -> tuple[float, float]:
"""95% Wilson score interval for k successes out of n. Behaves well for small n."""
if n == 0:
return 0.0, 0.0
p = k / n
centre = (p + z * z / (2 * n)) / (1 + z * z / n)
half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
return max(0.0, centre - half), min(1.0, centre + half)
def paired_bootstrap(a: list[bool], b: list[bool], iters: int = 5000, seed: int = 0) -> tuple[float, float, float]:
"""Accuracy difference b - a on the SAME tickets, with a 95% bootstrap interval."""
rng = random.Random(seed)
n = len(a)
diffs = []
for _ in range(iters):
idx = [rng.randrange(n) for _ in range(n)]
diffs.append(sum(b[i] - a[i] for i in idx) / n)
diffs.sort()
return (sum(b) - sum(a)) / n, diffs[int(0.025 * iters)], diffs[int(0.975 * iters) - 1]
# Backends: everything below has the same signature as supportdesk.llm.chat ---------------
TICKET_RE = re.compile(r"<ticket[^>]*>\n(.*?)\n</ticket>\s*$", re.S)
LABELS_RE = re.compile(r"<labels>(.*?)</labels>")
def last_ticket_text(messages: list[dict]) -> str:
"""The ticket being classified is the final <ticket> block of the last user message.
Search from the LAST '<ticket' only: a regex run from the start would begin at
the first example's tag and swallow every example (a real bug found in Part E).
"""
content = messages[-1]["content"]
start = content.rfind("<ticket")
match = TICKET_RE.search(content, max(start, 0))
return html.unescape(match.group(1)) if match else ""
def keyword_triage(text: str) -> dict:
"""A deliberately simple English keyword baseline. Rules were written while reading dev tickets."""
t = text.lower()
has = lambda *words: any(w in t for w in words) # noqa: E731
if has("cancel") or (has("refund") and has("annual", "renew", "months")):
category = "cancellation"
elif has("password", "locked", "2fa", "saml", "sso users", "log in", "login", "sign in", "owner",
"gdpr", "region", "stored", "backups", "reset link"):
category = "account_access"
elif has("please add", "integration", "when will", "roadmap", "do you have", "custom domain", "sync"):
category = "feature_request"
elif has("charge", "invoice", "refund", "price", "cost", "pay ", "vat", "tax", "discount", "declined", "credit"):
category = "billing"
elif has("not loading", "stopped", "fails", "error", "500", "isn't firing", "not firing", "lost my changes",
"bug", "no push", "collapse"):
category = "bug"
else:
category = "how_to"
if has("anyone", "all our users", "outage", "attack", "didn't request", "unlock it now"):
priority = "urgent"
elif has("charged", "twice", "lost my phone", "can't access", "compliance", "deal-breaker", "garbage",
"by mistake", "important"):
priority = "high"
elif has("?") and not has("stopped", "fails", "error", "missing", "not ", "isn't"):
priority = "low"
else:
priority = "normal"
needs_human = has("refund", "twice", "gdpr", "legal", "attack", "reissue", "declined", "don't recognize",
"garbage", "left company", "owner left", "sla credit", "owed credits")
return {"category": category, "priority": priority, "needs_human": needs_human}
def rules_backend() -> ScriptedLLM:
return ScriptedLLM(responder=lambda messages, kw: json.dumps(keyword_triage(last_ticket_text(messages))),
model="keyword-rules")
def copy_backend() -> ScriptedLLM:
"""Majority label of the examples shown in the prompt; ties go to the example nearest the ticket."""
def respond(messages: list[dict], kwargs: dict) -> str:
shown = [json.loads(html.unescape(s)) for s in LABELS_RE.findall(messages[-1]["content"])]
if not shown:
return "I need examples to copy from."
out = {}
for label in LABELS:
counts = Counter(ex[label] for ex in shown)
top = max(counts.values())
out[label] = next(ex[label] for ex in reversed(shown) if counts[ex[label]] == top)
return json.dumps(out)
return ScriptedLLM(responder=respond, model="copy-examples")
MIXED_REPLIES = [
'{"category": "billing", "priority": "high", "needs_human": true}',
'```json\n{"category": "how_to", "priority": "low", "needs_human": false}\n```',
'Sure! Here is the triage:\n{"category": "bug", "priority": "normal", "needs_human": "false"}',
'{"category": "Billing Issue", "priority": "high", "needs_human": true}',
'{"category": "account_access", "priority": "high", "needs_human": tr',
"I think this is probably a billing question with high priority.",
]
def mixed_backend() -> ScriptedLLM:
"""Cycles through well-formed and malformed replies to test the parser and failure counting."""
cycle = iter(MIXED_REPLIES * 100)
return ScriptedLLM(responder=lambda messages, kw: next(cycle), model="mixed-replies")
def make_backend(name: str):
if name == "rules":
return rules_backend()
if name == "copy":
return copy_backend()
if name == "mixed":
return mixed_backend()
if name == "llm":
from supportdesk.llm import chat
return chat
raise ValueError(f"Unknown backend {name!r}")
# Step 2: run a prompt version over the frozen split -------------------------------------
@dataclass
class Run:
meta: dict
records: list[dict] = field(default_factory=list)
def correct(self, label: str) -> list[bool]:
"""Per-ticket correctness; a parse failure counts as wrong."""
return [bool(r["pred"]) and r["pred"][label] == r["gold"][label] for r in self.records]
def run_eval(chat_fn, version: str, split: str = "dev", shots: str = "none", k: int = 0,
most_similar: str = "last", backend: str = "custom", allow_test: bool = False,
seed: int = 0, **chat_kwargs) -> Run:
if split == "test" and not allow_test:
raise RuntimeError("The test split is for the final check only. Pass allow_test=True (CLI: --final) once.")
manifest = verify_manifest(split)
template = load_prompt(version)
pool = load_tickets("dev") # few-shot examples only ever come from dev, never from test
selector = ExampleSelector(pool)
by_id = {t.id: t for t in load_tickets(split)}
run = Run(meta={"version": version, "fingerprint": template.fingerprint, "split": split, "n": manifest["n"],
"shots": shots, "k": k, "most_similar": most_similar, "backend": backend, "seed": seed,
"chat_kwargs": chat_kwargs})
for ticket_id in manifest["ids"]:
ticket = by_id[ticket_id]
examples = arrange(selector.select(ticket, shots, k, seed), most_similar)
messages = template.render(ticket, examples)
result = chat_fn(messages, **chat_kwargs)
pred, error = parse_triage(result.text)
run.records.append({"id": ticket.id, "language": ticket.language, "gold": gold_labels(ticket),
"pred": pred, "error": error, "raw": result.text[:300],
"input_tokens": result.usage.input_tokens, "output_tokens": result.usage.output_tokens,
"latency_ms": result.latency_ms, "examples": [e.id for e in examples]})
return run
def save_run(run: Run, name: str | None = None) -> Path:
RUN_DIR.mkdir(exist_ok=True)
m = run.meta
name = name or f"{m['version']}-{m['backend']}-{m['shots']}{m['k']}-{m['split']}"
path = RUN_DIR / f"{name}.jsonl"
with path.open("w", encoding="utf-8") as f:
f.write(json.dumps({"meta": m}) + "\n")
for r in run.records:
f.write(json.dumps(r, ensure_ascii=False) + "\n")
return path
def load_run(path: Path | str) -> Run:
lines = Path(path).read_text(encoding="utf-8").splitlines()
return Run(meta=json.loads(lines[0])["meta"], records=[json.loads(x) for x in lines[1:]])
# Step 3: report and compare ---------------------------------------------------------------
def summarize(run: Run) -> str:
m = run.meta
n = len(run.records)
failures = [r for r in run.records if r["error"]]
tokens = sum(r["input_tokens"] for r in run.records) / max(n, 1)
lines = [f"{m['version']} ({m['fingerprint']}) backend={m['backend']} shots={m['shots']} k={m['k']} "
f"split={m['split']} n={n}",
f" parse failures: {len(failures)}/{n} mean input tokens: {tokens:.0f}"]
for label in LABELS:
ok = run.correct(label)
lo, hi = wilson(sum(ok), n)
lines.append(f" {label:<12} {sum(ok):>3}/{n} acc {sum(ok) / n:.3f} 95% CI [{lo:.3f}, {hi:.3f}]")
return "\n".join(lines)
def compare(a: Run, b: Run) -> str:
if [r["id"] for r in a.records] != [r["id"] for r in b.records]:
raise ValueError("Runs cover different tickets; paired comparison needs the same frozen set.")
lines = [f"{b.meta['version']}/{b.meta['backend']} minus {a.meta['version']}/{a.meta['backend']} "
f"(paired, n={len(a.records)})"]
for label in LABELS:
diff, lo, hi = paired_bootstrap(a.correct(label), b.correct(label))
verdict = "within noise" if lo <= 0 <= hi else ("better" if diff > 0 else "worse")
lines.append(f" {label:<12} {diff:+.3f} 95% CI [{lo:+.3f}, {hi:+.3f}] {verdict}")
return "\n".join(lines)
def plateau(runs: list[Run], label: str = "category", window: int = 3) -> tuple[bool, str]:
"""Ceiling check: are the last `window` versions all within noise of the best one so far?"""
if len(runs) < window:
return False, f"need at least {window} runs"
scores = [sum(r.correct(label)) / len(r.records) for r in runs]
best = max(range(len(runs)), key=lambda i: scores[i])
notes = []
for i in range(len(runs) - window, len(runs)):
if i == best:
continue
diff, lo, hi = paired_bootstrap(runs[i].correct(label), runs[best].correct(label))
notes.append((runs[i].meta["version"], round(diff, 3), round(lo, 3), round(hi, 3)))
if lo > 0:
return False, f"best ({runs[best].meta['version']}) is clearly above {runs[i].meta['version']}: keep iterating"
return True, f"last {window} versions within noise of best ({runs[best].meta['version']}): {notes}"
# Command line -----------------------------------------------------------------------------
def main() -> None:
parser = argparse.ArgumentParser(description="Triage prompt harness (Module 4)")
sub = parser.add_subparsers(dest="cmd", required=True)
f = sub.add_parser("freeze")
f.add_argument("--split", default="dev")
r = sub.add_parser("run")
r.add_argument("--version", required=True)
r.add_argument("--backend", default="rules", choices=["rules", "copy", "mixed", "llm"])
r.add_argument("--split", default="dev")
r.add_argument("--shots", default="none", choices=["none", "random", "similar", "balanced"])
r.add_argument("--k", type=int, default=0)
r.add_argument("--most-similar", default="last", choices=["first", "last"])
r.add_argument("--seed", type=int, default=0, help="seed for --shots random")
r.add_argument("--final", action="store_true", help="allow the test split (use once, at the end)")
r.add_argument("--max-tokens", type=int, default=None)
r.add_argument("--reasoning-effort", default=None, help="for reasoning models, e.g. low (llm backend only)")
c = sub.add_parser("compare")
c.add_argument("runs", nargs="+")
args = parser.parse_args()
if args.cmd == "freeze":
m = freeze_eval_set(args.split)
print(f"froze {m['split']}: n={m['n']} labels_sha256={m['labels_sha256'][:16]}")
elif args.cmd == "run":
kwargs = {"max_tokens": args.max_tokens} if args.max_tokens else {}
if args.reasoning_effort:
kwargs["reasoning_effort"] = args.reasoning_effort
run = run_eval(make_backend(args.backend), args.version, args.split, args.shots, args.k,
args.most_similar, args.backend, allow_test=args.final, seed=args.seed, **kwargs)
path = save_run(run)
print(summarize(run))
print(f" saved {path.relative_to(ROOT)}")
else:
runs = [load_run(p) for p in args.runs]
for run in runs:
print(summarize(run))
for a, b in zip(runs, runs[1:]):
print(compare(a, b))
if __name__ == "__main__":
main()
Code explained
- In simple words: the harness: it freezes the test set, runs a prompt version over every ticket through any chat function, saves every prediction, and reports accuracy with honest error bars.
- What happens:
freeze_eval_set(split)writesevals/triage_<split>_manifest.jsonwith the ticket ids and a SHA-256 hash of all gold labels, and refuses to overwrite an existing manifest.labels_digestcomputes the hash.verify_manifest(split)recomputes it before every run and stops if anything changed.wilson(k, n)is the Wilson score interval, a 95% confidence interval for a proportion that behaves sensibly at small n (it never goes below 0 or above 1, unlike the textbook plus-or-minus formula). With n=48, it is wide: that is the point.paired_bootstrap(a, b)compares two runs on the same tickets. It resamples tickets with replacement 5,000 times and takes the middle 95% of the accuracy differences. Paired means each resample uses the same tickets for both runs, so a hard ticket is hard for both; this makes the comparison much more sensitive than comparing two separate intervals.last_ticket_text(messages)pulls the ticket back out of the rendered prompt and unescapes it. The comment explains a bug the first version had; Part E walks through how it was found.keyword_triage(text)is the keyword baseline: plain English substring rules for all three labels. It is not a model and was written while reading dev tickets, so its dev score is optimistic.rules_backend,copy_backend, andmixed_backendwrap behaviors inScriptedLLMso they have exactly the call signature ofsupportdesk.llm.chat. The harness cannot tell them from a real model, which is what makes them good plumbing tests.copy_backendreads the<labels>of the few-shot examples in the prompt and answers with the majority label (ties go to the example nearest the ticket), so it measures how informative the examples are.mixed_backendcycles through the malformed replies from Part A.make_backend("llm")returns the realchatfunction.Runholds metadata and per-ticket records;correct(label)counts a parse failure as wrong, because a triage system that cannot read its own output has failed that ticket.run_eval(...)verifies the manifest, loads the version, selects and arranges examples for each ticket (from dev only, never from test), renders, calls the chat function, parses, and records gold, prediction, error, raw reply, tokens, latency, and which examples were shown. It refuses the test split unlessallow_test=True.save_runandload_runwrite and read one JSONL file per run.summarizeprints parse failures, mean input tokens, and per-label accuracy with Wilson intervals.compareprints paired differences with a verdict of better, worse, or within noise.plateauis the ceiling check from Part E.main()exposesfreeze,run, andcompareon the command line.
- Comes out: see the runs below.
Now freeze both splits, and watch the two guards work.
PYTHONPATH=. python examples/m04_eval.py freeze --split dev
PYTHONPATH=. python examples/m04_eval.py freeze --split test
PYTHONPATH=. python examples/m04_eval.py freeze --split dev # second time: refused
PYTHONPATH=. python examples/m04_eval.py run --version v1 --backend rules --split test # test: refused
Code explained
- In simple words: record exactly what the prompt will be judged on, then show that the harness will not let you quietly redo it or peek at the test set.
- What happens: the first two commands write the manifests. The third tries to freeze dev again and hits
FileExistsError. The fourth tries to score on test without--finaland hitsRuntimeError. (Only the last line of each traceback is shown.) - Comes out: the dev and test hashes, then the two refusals. If you ever need to relabel, delete the manifest on purpose, relabel, refreeze, and rerun every version you want to compare, because old scores are no longer comparable.
froze dev: n=48 labels_sha256=8174913f3e0aaabc
froze test: n=24 labels_sha256=0ddc5b10e2dd83a1
FileExistsError: triage_dev_manifest.json already exists; a frozen set is never rewritten silently.
RuntimeError: The test split is for the final check only. Pass allow_test=True (CLI: --final) once.
Step 3: a baseline and a control
Before you measure a prompt, measure something dumb. The keyword baseline gives the floor a model has to beat, and because it ignores the instructions, it doubles as a control: any change in its score between prompt versions can only come from a bug in the plumbing.
PYTHONPATH=. python examples/m04_eval.py run --version v1 --backend rules
PYTHONPATH=. python examples/m04_eval.py run --version v2 --backend rules
PYTHONPATH=. python examples/m04_eval.py compare runs/v1-rules-none0-dev.jsonl runs/v2-rules-none0-dev.jsonl
Code explained
- In simple words: score the keyword rules through the full harness with prompts v1 and v2, then compare the two runs ticket by ticket.
- What happens: each
runrenders all 48 dev tickets with the given version, sends them to the rules backend, parses, scores, and saves a run file.compareprints both summaries and then the paired difference (only the difference is shown below the two summaries). - Comes out: 38/48 on category (0.792, CI 0.657 to 0.883), 33/48 on priority, 43/48 on needs_human. The two versions score identically, as a control should. Notice the interval width: with n=48, a true accuracy anywhere from about 66% to 88% is consistent with 38 correct. v2 costs 265 input tokens per ticket on average against 166 for v1; with the rules backend that extra text buys nothing, and whether it buys something from a real model is exactly what the harness is for.
v1 (16af0ee83319) backend=rules shots=none k=0 split=dev n=48
parse failures: 0/48 mean input tokens: 166
category 38/48 acc 0.792 95% CI [0.657, 0.883]
priority 33/48 acc 0.688 95% CI [0.547, 0.801]
needs_human 43/48 acc 0.896 95% CI [0.778, 0.955]
saved runs/v1-rules-none0-dev.jsonl
v2 (6dca0c9eb0b6) backend=rules shots=none k=0 split=dev n=48
parse failures: 0/48 mean input tokens: 265
category 38/48 acc 0.792 95% CI [0.657, 0.883]
priority 33/48 acc 0.688 95% CI [0.547, 0.801]
needs_human 43/48 acc 0.896 95% CI [0.778, 0.955]
saved runs/v2-rules-none0-dev.jsonl
v2/rules minus v1/rules (paired, n=48)
category +0.000 95% CI [+0.000, +0.000] within noise
priority +0.000 95% CI [+0.000, +0.000] within noise
needs_human +0.000 95% CI [+0.000, +0.000] within noise
Read the failures before you trust the number. Every non-English dev ticket (es, de, ja, hi) gets the wrong category from the rules, because the keywords are English. Priority errors are mostly "normal" tickets phrased as questions, which the rules call "low". A model's advantage will likely show up exactly there.
Step 4: test the plumbing with deliberately bad replies
The mixed backend returns the Part A reply corpus in rotation. The goal is not a score; it is to prove that parse failures are counted, labeled, and do not crash anything.
PYTHONPATH=. python examples/m04_eval.py run --version v1 --backend mixed
PYTHONPATH=. python -c "
from collections import Counter
from examples.m04_eval import load_run
run = load_run('runs/v1-mixed-none0-dev.jsonl')
print(Counter(r['error'] for r in run.records if r['error']))"
Code explained
- In simple words: run the harness against a fake model that answers in broken ways half the time, and count why each reply failed.
- What happens: the six scripted replies repeat 8 times over 48 tickets. Three of the six fail to parse, so 24 failures are expected. The one-liner reads the saved run and tallies the error messages.
- Comes out: exactly 24/48 parse failures, split 16 "no JSON object found" (the truncated and prose replies) and 8 "bad category" (the invented label). The accuracies are meaningless, because the replies ignore the tickets. This is
ScriptedLLMoutput, not model output.
v1 (16af0ee83319) backend=mixed shots=none k=0 split=dev n=48
parse failures: 24/48 mean input tokens: 166
category 4/48 acc 0.083 95% CI [0.033, 0.196]
priority 6/48 acc 0.125 95% CI [0.059, 0.247]
needs_human 12/48 acc 0.250 95% CI [0.149, 0.388]
saved runs/v1-mixed-none0-dev.jsonl
Counter({'no JSON object found': 16, "bad category 'billing issue'": 8})
Step 5: the same harness with a real model
With a key (or a local Ollama), the only change is the backend. This is the command to run for your own numbers:
export LLM_PROVIDER=groq GROQ_API_KEY=your-key-here # or: LLM_PROVIDER=gemini GEMINI_API_KEY=...; or LLM_PROVIDER=ollama
PYTHONPATH=. python examples/m04_eval.py run --version v1 --backend llm --reasoning-effort low
PYTHONPATH=. python examples/m04_eval.py run --version v2 --backend llm --reasoning-effort low
PYTHONPATH=. python examples/m04_eval.py run --version v3 --backend llm --shots balanced --k 6 --reasoning-effort low
PYTHONPATH=. python examples/m04_eval.py compare runs/v1-llm-none0-dev.jsonl runs/v2-llm-none0-dev.jsonl runs/v3-llm-balanced6-dev.jsonl
Code explained
- In simple words: the same three versions, now answered by a real model through
supportdesk.llm.chat, then compared pairwise. - What happens:
--backend llmpasses the realchatfunction torun_eval.--reasoning-effort lowis forwarded for reasoning models such as Groq's defaultopenai/gpt-oss-120b(Module 3); leave it out for models that do not accept it. Temperature is 0 by default inchat. Each run makes 48 calls and records real token usage and latency.compareprints all three summaries and the v1-to-v2 and v2-to-v3 paired differences. - Comes out: in this build there is no key, so the first call stops with
RuntimeError: Set GROQ_API_KEY in your environment to use groq.The block below is an illustrative sample run (not captured in this build; produced for teaching). Your output will differ, including the numbers, and you should not quote them. It shows the shape to expect and one pattern worth looking for: a real model usually beats the keyword rules on non-English tickets, and gains between versions often land inside the noise band at n=48.
v2/llm minus v1/llm (paired, n=48)
category +0.021 95% CI [-0.042, +0.083] within noise
priority +0.104 95% CI [+0.021, +0.188] better
needs_human +0.063 95% CI [-0.021, +0.146] within noise