Part B: Shared Requirements
Every track must deliver the same four things. This part builds the tooling for each one, runs it on the course's own simple systems, and shows how to read the output. The demos use the keyword baseline's failures so they run here; your capstone must run the same tools on your own system's failures and calls.
B1. A baseline to beat
A baseline is the simplest reasonable system for the task, measured the same way your final system will be. It turns "our assistant is good" into "our assistant gets 19 of 24 where the baseline gets 16", and it sometimes reveals that the cheap system is already good enough. A majority-class floor (always predict the most common label) is the baseline below the baseline: anything that cannot beat it has learned nothing.
Brightlane gets two baselines. For triage, keyword rules guess a category and a priority. For retrieval, the BM25 search from Module 7 returns help-center articles, scored by whether the gold article comes first (top-1) or in the first three (top-3). Scores come with a Wilson interval, the 95% confidence interval for a proportion that behaves well at small n (Module 10 derived it).
The script writes one record per ticket per task. That record format is the contract for everything after it: the error-analysis tool reads it, so when you replace the baseline with your system, you write the same fields.
examples/cap_baseline.py
# examples/cap_baseline.py
"""Capstone shared requirement 1: a baseline the final system must beat.
Two baselines, both deterministic and free to run:
triage keyword rules that guess category and priority
retrieval BM25 top-1 and top-3 from supportdesk.kb_search
Scores come with Wilson 95% intervals because the test split is only 24 tickets.
Writes runs/baseline_run.jsonl (one record per ticket per task) and runs/baseline.json.
"""
from __future__ import annotations
import json
import math
from collections import Counter
from pathlib import Path
from supportdesk.data import Ticket, load_tickets
from supportdesk.kb_search import KBSearch
RUNS = Path("runs")
# Rules were written by reading the DEV split only. Order matters: first match wins.
CATEGORY_RULES: list[tuple[str, tuple[str, ...]]] = [
("cancellation", ("cancel", "money back", "stop now", "downgrade")),
("feature_request", ("please add", "will you support", "do you have a", "when will", "roadmap")),
("account_access", ("password", "locked", "2fa", "log in", "sign in", "sso", "saml", "gdpr",
"personal data", "region", "owner")),
("bug", ("not working", "stopped", "fails", "error", "broken", "outage", "not loading", "bug")),
("billing", ("charge", "charged", "invoice", "refund", "price", "cost", "pay", "vat", "plan")),
]
URGENT = ("nobody can", "not loading for anyone", "demo in", "outage", "all our users")
HIGH = ("charged twice", "locked", "lost my phone", "deleted", "can't access", "refund", "deal-breaker")
def keyword_triage(text: str) -> tuple[str, str]:
"""Guess (category, priority) from keywords. Falls back to how_to / normal."""
low = text.lower()
category = next((cat for cat, words in CATEGORY_RULES if any(w in low for w in words)), "how_to")
if any(w in low for w in URGENT):
priority = "urgent"
elif any(w in low for w in HIGH):
priority = "high"
elif low.rstrip().endswith("?") and len(low) < 160:
priority = "low"
else:
priority = "normal"
return category, priority
def wilson(k: int, n: int, z: float = 1.96) -> tuple[float, float]:
"""Wilson score interval for k successes in n trials (safer than p +/- 2se at small n)."""
if n == 0:
return (0.0, 0.0)
p = k / n
centre = (p + z * z / (2 * n)) / (1 + z * z / n)
half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
return (round(max(0.0, centre - half), 3), round(min(1.0, centre + half), 3))
def sign_test_p(wins: int, losses: int) -> float:
"""Exact one-sided sign test (McNemar) on paired results.
wins = items the new system gets right and the baseline gets wrong; losses =
the reverse. Returns the chance of at least `wins` wins if both systems were
equally good. Ties (both right or both wrong) carry no information.
"""
n = wins + losses
return 1.0 if n == 0 else sum(math.comb(n, k) for k in range(wins, n + 1)) / 2 ** n
def run_baseline(tickets: list[Ticket], search: KBSearch | None = None) -> list[dict]:
"""One record per ticket per task, in the format the error-analysis tool reads."""
search = search or KBSearch()
records = []
for t in tickets:
category, priority = keyword_triage(t.text)
base = {"id": t.id, "split": t.split, "language": t.language, "tier": t.customer_tier, "text": t.text}
records.append({**base, "task": "category", "expected": t.gold["category"], "predicted": category})
records.append({**base, "task": "priority", "expected": t.gold["priority"], "predicted": priority})
if t.gold["answerable"]:
hits = [h.article_id for h in search.search(t.text, k=3)]
records.append({**base, "task": "retrieval_top1", "expected": t.gold["kb_article"],
"predicted": hits[0] if hits else None, "top3": hits})
for r in records:
r["correct"] = r["predicted"] == r["expected"]
return records
def score(records: list[dict], task: str, split: str | None = None) -> dict:
rows = [r for r in records if r["task"] == task and (split is None or r["split"] == split)]
if task == "retrieval_top1":
top3 = sum(r["expected"] in r["top3"] for r in rows)
k, n = sum(r["correct"] for r in rows), len(rows)
out = {"task": task, "split": split or "all", "k": k, "n": n, "acc": round(k / n, 3), "ci95": wilson(k, n)}
if task == "retrieval_top1":
out["top3_k"] = top3
out["top3_acc"] = round(top3 / n, 3)
out["top3_ci95"] = wilson(top3, n)
return out
def main() -> None:
RUNS.mkdir(exist_ok=True)
records = run_baseline(load_tickets())
with (RUNS / "baseline_run.jsonl").open("w", encoding="utf-8") as f:
for r in records:
f.write(json.dumps(r, ensure_ascii=False) + "\n")
summary = []
for split in ("dev", "test"):
for task in ("category", "priority", "retrieval_top1"):
s = score(records, task, split)
summary.append(s)
extra = f" top3 {s['top3_k']}/{s['n']} = {s['top3_acc']:.3f} {s['top3_ci95']}" if "top3_k" in s else ""
print(f"{split:4} {task:15} {s['k']:2}/{s['n']:2} = {s['acc']:.3f} 95% CI {s['ci95']}{extra}")
dev_counts = Counter(t.gold["category"] for t in load_tickets("dev"))
top_class = dev_counts.most_common(1)[0][0]
test = load_tickets("test")
k = sum(t.gold["category"] == top_class for t in test)
print(f"test majority-class floor (always {top_class!r}, the most common dev label): "
f"{k}/{len(test)} = {k / len(test):.3f} 95% CI {wilson(k, len(test))}")
(RUNS / "baseline.json").write_text(json.dumps(summary, indent=1) + "\n")
print(f"wrote {len(records)} records to runs/baseline_run.jsonl")
if __name__ == "__main__":
main()
Code explained
- In simple words: a deliberately simple triage clerk with a list of keywords, a search box, and a ruler that always prints its error bars.
- What happens:
CATEGORY_RULES,URGENT, andHIGHare phrase lists written by reading the 48 dev tickets only. Order matters: the first category whose phrase appears wins, andhow_tois the fallback.keyword_triagelowercases the text and applies the rules. Priority isurgentorhighif a phrase matches,lowif the ticket is a short question (ends with?, under 160 characters), andnormalotherwise. Substring matching (w in low) is a real weakness on purpose; B2 finds it.wilsonreturns the Wilson score interval for k successes in n trials, clipped to 0 and 1.sign_test_pis the exact one-sided sign test (the exact form of McNemar's test) on paired results: given the tickets a new system fixed (wins) and broke (losses), the chance of at least that many wins if both systems were equally good. Every track uses it to compare against the baseline.run_baselineproduces records withid,split,language,tier,text,task,expected,predicted, andcorrect; retrieval records also keep the top-3 list. Only answerable tickets get a retrieval record.scoreaggregates one task on one split and adds top-3 numbers for retrieval.mainscores dev and test, prints the majority-class floor for test (alwayshow_to, the most common dev label), and writesruns/baseline_run.jsonlandruns/baseline.json.
- Comes out: deterministic, so your numbers will match.
Read it from the bottom up. The majority floor is 6 of 24; keyword category (16 of 24) clears it, and its interval (0.467 to 0.82) barely overlaps the floor's (0.12 to 0.449). Priority is weak at 10 of 24. Retrieval is strong in English: 17 of 20 answerable test tickets get the gold article first, 19 of 20 in the top 3, which matches the course-wide 52 of 62 top-1 (35 dev plus 17 test).
Two things surprise people here. First, the rules score only 31 of 48 on the dev tickets they were written from, so they are not even overfit; they are simply weak. Second, those intervals are wide. A system that scores 19 of 24 on category (0.79) has an interval of about 0.6 to 0.9, which overlaps the baseline's. With 24 test tickets you can only claim a win for a large or consistent improvement, which is why every track compares systems ticket by ticket with
sign_test_prather than by comparing two percentages.Which baseline should your capstone beat? The strongest cheap one you have.
Situation Use this Why Classification with a few dozen labels Keyword rules and a majority floor (this script) Minutes to build, fully explainable, and a real bar: 16 of 24 here Classification with hundreds or thousands of labelled historical tickets TF-IDF plus logistic regression (Module 14 Part C) Stronger than rules once labels are plentiful; the bar a fine-tune or LLM must clear Retrieval or grounded answers BM25 top-1 and top-3 ( KBSearch)Hard to beat in the language it was built for; 17 of 20 on test Agent tasks The same task done by a fixed script (search, read top article, draft) Shows what the agent's autonomy adds, step by step A new LLM feature replacing a human step The human's current accuracy and time, measured on a sample The business decision is against today's process, not against a toy B2. Error analysis on at least 50 failures, clustered by cause
Error analysis means reading failures one by one and writing down why each happened, then counting. A cause taxonomy is the fixed list of reasons you tag with; a good tag names something you could act on ("the short-question rule fired") rather than a symptom ("priority too low"). Clustering groups failures that share a tag or a feature, so you fix the biggest group first. The requirement is at least 50 failures because with fewer, the ranking of clusters is mostly luck.
The baseline fails on 69 real records (a ticket can fail on several tasks), which is enough. Many capstone systems will fail less often on 72 tickets, so the tooling also accepts an expanded synthetic set: tickets generated from templates to probe variations real data has too few of (typos, other languages, negation, two requests in one ticket). They carry a clear label, and the report keeps them apart, because their gold labels come from the template author, not from a support agent.
examples/cap_synthetic_tickets.py
python# examples/cap_synthetic_tickets.py """Generate a SYNTHETIC ticket set to widen error analysis (never to replace real failures). Every ticket here is produced from templates written for this capstone. Gold labels come from the template author, not from a support agent, so treat failure RATES on this set as meaningless and failure KINDS as hints about what to look for in real traffic. Ids start with SYN- and split is "synthetic". Writes data/synthetic_tickets.jsonl. """ from __future__ import annotations import json import random from pathlib import Path from supportdesk.data import DATA_DIR, Ticket OUT = DATA_DIR / "synthetic_tickets.jsonl" # (variation kind, subject, body, language, category, priority, kb_article) TEMPLATES: list[tuple[str, str, str, str, str, str, str | None]] = [ # paraphrase: same intent, none of the obvious keywords ("paraphrase", "Two identical payments", "My bank statement shows the same Brightlane amount taken on the same day, {n} USD each.", "en", "billing", "high", "billing-refunds"), ("paraphrase", "Can't get in", "The app keeps rejecting me at the sign-in screen after I typed the wrong thing a few times.", "en", "account_access", "high", "account-login"), ("paraphrase", "Leaving Brightlane", "We have decided to end our subscription at the end of the month. What are the steps?", "en", "cancellation", "normal", "billing-refunds"), ("paraphrase", "Nothing arrives in our channel", "Updates from our boards no longer show up in our team chat channel since {day}.", "en", "bug", "normal", "integrations-slack"), ("paraphrase", "Getting everything out", "Before we switch vendors we want a complete copy of our boards and files.", "en", "how_to", "normal", "exports-data"), ("paraphrase", "Wish list", "It would be great if timelines could show which tasks block others.", "en", "feature_request", "low", "feature-requests"), # typos and informal writing ("typo", "chargd 2x", "hi i got chargd twise for teem plan pls refnd", "en", "billing", "high", "billing-refunds"), ("typo", "pasword", "forgot pasword and the resett link is expird", "en", "account_access", "normal", "account-login"), ("typo", "cancle", "how do i cancle my subscripton??", "en", "cancellation", "normal", "billing-refunds"), ("typo", "exprt csv", "how to exprt a bord to csv", "en", "how_to", "low", "exports-data"), # languages the rules never saw ("non_english", "Doble cargo", "Me cobraron dos veces este mes. Quiero el reembolso.", "es", "billing", "high", "billing-refunds"), ("non_english", "Konto gesperrt", "Mein Konto ist nach mehreren falschen Passwoertern gesperrt.", "de", "account_access", "high", "account-login"), ("non_english", "Annuler l'abonnement", "Je voudrais annuler notre abonnement Team. Comment faire ?", "fr", "cancellation", "normal", "billing-refunds"), ("non_english", "Exportar quadro", "Como exporto um quadro para CSV?", "pt", "how_to", "low", "exports-data"), ("non_english", "Preis Business", "Was kostet der Business-Plan pro Nutzer im Monat?", "de", "billing", "low", "billing-plans"), # negation and contrast: the keyword is present but denied ("negation", "Not about billing", "This is not a billing question: how do I set up an automation that moves cards to Done?", "en", "how_to", "low", "boards-automations"), ("negation", "Not cancelling", "We are NOT cancelling, we just want to know where invoices are sent.", "en", "billing", "normal", "billing-invoices"), ("negation", "No error, just slow", "There is no error message, I just want to know how many automation runs Team includes.", "en", "how_to", "low", "boards-automations"), # two intents in one ticket ("multi_intent", "Refund and a feature idea", "Please refund the duplicate charge from {day}. Also, will you add dark mode?", "en", "billing", "high", "billing-refunds"), ("multi_intent", "Locked out, then export", "I was locked out yesterday, fine now. How do I export my board to CSV?", "en", "how_to", "low", "exports-data"), ("multi_intent", "SSO and price", "Does Business include SSO, and what does it cost for {n} users?", "en", "billing", "normal", "billing-plans"), # urgency expressed without trigger words ("hidden_urgency", "Whole team blocked", "None of our {n} people can open any board right now, we have a launch today.", "en", "bug", "urgent", "status-incidents"), ("hidden_urgency", "Security worry", "Someone from an unknown country signed in to our admin account an hour ago.", "en", "account_access", "urgent", None), # out of scope for the help center (unanswerable) ("unanswerable", "Custom contract", "Our procurement team needs net-90 payment terms written into the contract.", "en", "billing", "normal", None), ("unanswerable", "On-prem install", "Can we install Brightlane on our own servers?", "en", "feature_request", "normal", None), ] DAYS = ("Monday", "Tuesday", "yesterday", "last week") TIERS = ("free", "team", "business", "enterprise") def build(copies: int = 4, seed: int = 7) -> list[Ticket]: """`copies` filled-in variants of every template, with random numbers, days, and tiers.""" rng = random.Random(seed) tickets = [] for i, (kind, subject, body, lang, cat, prio, kb) in enumerate(TEMPLATES): for c in range(copies): filled = body.format(n=rng.choice((12, 24, 48, 96, 288)), day=rng.choice(DAYS)) tickets.append(Ticket( id=f"SYN-{i:02d}{c}", subject=subject, body=filled, customer_tier=rng.choice(TIERS), language=lang, split="synthetic", gold={"category": cat, "priority": prio, "kb_article": kb, "answerable": kb is not None, "variation": kind, "synthetic": True}, )) return tickets def load_synthetic() -> list[Ticket]: with OUT.open(encoding="utf-8") as f: return [Ticket(**json.loads(line)) for line in f] def main() -> None: tickets = build() with OUT.open("w", encoding="utf-8") as f: for t in tickets: f.write(json.dumps(t.__dict__, ensure_ascii=False) + "\n") kinds = sorted({t.gold["variation"] for t in tickets}) print(f"wrote {len(tickets)} SYNTHETIC tickets to {OUT.relative_to(Path.cwd())}") print("variation kinds:", ", ".join(kinds)) if __name__ == "__main__": main()Code explained
- In simple words: a stencil kit that stamps out variations of known tricky tickets, each stamped "SYNTHETIC".
- What happens:
TEMPLATESholds 25 hand-written templates across seven variation kinds, each with its own gold category, priority, and article (orNonewhen the help center cannot answer).buildfills each template four times with random numbers, days, and customer tiers from a seeded generator, so the set is reproducible. Ids start withSYN-and the split issynthetic; the gold dict records the variation kind andsynthetic: True.load_syntheticreads the saved file back asTicketobjects.mainwritesdata/synthetic_tickets.jsonlin your copy (never in the canonical repository).
- Comes out:
Now the analysis tool. It proposes tags with simple rules, applies your corrections, clusters by tag and by features such as language, length, and customer tier, and writes a Markdown report. The auto-tagger exists to speed up reading, not to replace it.
examples/cap_error_analysis.py
python# examples/cap_error_analysis.py """Capstone shared requirement 2: error analysis on at least 50 failures, clustered by cause. Input: a run file (JSONL), one record per ticket per task with fields id, task, expected, predicted, correct, text, language, tier, split. Steps: collect failures, propose tags with simple rules, apply a human's overrides, cluster by tag and by text features, and write a Markdown report. The auto-tagger only PROPOSES causes to speed up reading. You still read every failure and correct the tags by hand (runs/tag_overrides.json). Use failures from YOUR system: the demo below runs on the keyword baseline. """ from __future__ import annotations import json import re import sys from collections import Counter, defaultdict from pathlib import Path from supportdesk.data import load_articles, load_tickets from supportdesk.kb_search import tokenize sys.path.insert(0, str(Path(__file__).resolve().parent)) from cap_baseline import CATEGORY_RULES, HIGH, URGENT, run_baseline # noqa: E402 (same capstone, same folder) from cap_synthetic_tickets import load_synthetic # noqa: E402 RUNS = Path("runs") PRIORITY_RANK = {"low": 0, "normal": 1, "high": 2, "urgent": 3} # The taxonomy: every tag is a CAUSE you could act on, not a symptom. TAXONOMY: dict[str, str] = { "non_english": "Ticket is not in English; the system was built from English examples.", "no_signal": "Nothing in the text triggered any rule; the system fell back to its default.", "keyword_collision": "Words for two or more categories appear; the first rule to match won.", "substring_match": "A keyword matched inside a longer word ('plan' in 'plane', 'vat' in 'private').", "topic_not_intent": "A keyword names the topic, but the customer's intent is different (human tag).", "distractor_term": "A strong but irrelevant term pulled retrieval to the wrong article (human tag).", "negation": "A keyword appears but is denied or contrasted ('not a billing question').", "multi_intent": "The ticket asks for two things; gold follows the main one.", "misspelling": "Many words are not in the domain vocabulary (typos, slang).", "question_rule": "Priority: the 'short question means low' rule fired, but the customer needs real help.", "urgency_unseen": "Priority: gold is high or urgent, but no urgency phrase matched (stakes stated in other words).", "keyword_overfire": "Priority: an urgency phrase matched, but the situation is calmer than the phrase suggests.", "default_priority": "Priority: no priority rule fired; the default 'normal' was wrong.", "vocab_gap": "Ticket and gold article share almost no words; keyword search cannot bridge it.", "near_miss": "Gold article is in the top 3 but not ranked first.", "label_question": "A reviewer thinks the gold label itself is debatable.", "other": "No rule fired; a human must name the cause.", } VOCAB = {w for a in load_articles() for w in tokenize(a.title + " " + a.body)} | { w for t in load_tickets() for w in tokenize(t.text)} NEGATION = re.compile(r"\b(not|no|never|isn't|aren't|don't|NOT)\b", re.IGNORECASE) def rule_hits(text: str) -> list[str]: low = text.lower() return [cat for cat, words in CATEGORY_RULES if any(w in low for w in words)] def substring_only(text: str) -> bool: """True if the first rule that fired matched only inside longer words.""" low = text.lower() for _, words in CATEGORY_RULES: found = [w for w in words if w in low] if found: return not any(re.search(rf"\b{re.escape(w)}\b", low) for w in found) return False def features(r: dict) -> dict: """Simple, explainable features of a record, used for clustering beyond tags.""" text = r["text"] body_len = len(text) return { "task": r["task"], "source": "synthetic" if r["split"] == "synthetic" else "real", "language": r["language"], "length": "short (<90 chars)" if body_len < 90 else ("long (>180)" if body_len > 180 else "medium"), "has_number": str(bool(re.search(r"\d", text))), "tier": r["tier"], "confusion": f"{r['expected']} -> {r['predicted']}", } def propose_tags(r: dict) -> list[str]: """Heuristic first pass at WHY a record failed. A human confirms or overrides.""" text, tags = r["text"], [] words = [w for w in tokenize(text) if w.isalpha()] if r["language"] != "en": tags.append("non_english") if r["task"] in ("category", "priority"): hits = rule_hits(text) if r["task"] == "category": if not hits: tags.append("no_signal") elif len(hits) > 1: tags.append("keyword_collision") if hits and substring_only(text): tags.append("substring_match") if NEGATION.search(text) and hits: tags.append("negation") body = text.split("\n\n", 1)[-1] # the subject usually restates the question; count the body only if body.count("?") >= 2 or re.search(r"\balso\b|, then |\band what\b", body, re.IGNORECASE): tags.append("multi_intent") if r["language"] == "en" and words and sum(w not in VOCAB for w in words) / len(words) > 0.3: tags.append("misspelling") if r["task"] == "priority": low = text.lower() exp, got = PRIORITY_RANK[r["expected"]], PRIORITY_RANK[r["predicted"]] fired = any(w in low for w in URGENT + HIGH) if got > exp and fired: tags.append("keyword_overfire") elif r["predicted"] == "low": tags.append("question_rule") elif exp >= PRIORITY_RANK["high"] and not fired: tags.append("urgency_unseen") else: tags.append("default_priority") if r["task"] == "retrieval_top1": if r["expected"] in r.get("top3", []): tags.append("near_miss") gold = next(a for a in load_articles() if a.id == r["expected"]) if len(set(tokenize(text)) & set(tokenize(gold.title + " " + gold.body))) <= 2: tags.append("vocab_gap") return tags or ["other"] def failure_key(r: dict) -> str: return f"{r['id']}:{r['task']}" def analyze(records: list[dict], overrides: dict[str, list[str]] | None = None) -> dict: overrides = overrides or {} failures = [dict(r) for r in records if not r["correct"]] for f in failures: f["tags"] = overrides.get(failure_key(f), propose_tags(f)) f["reviewed"] = failure_key(f) in overrides f["features"] = features(f) by_tag = Counter(tag for f in failures for tag in f["tags"]) # Failure RATE per feature value needs the successes too: a cluster is only # interesting when the rate there is higher than the rate overall. # Rates use REAL records only: synthetic gold labels come from a template author, # so a synthetic failure rate says nothing about production. real_records = [r for r in records if r["split"] != "synthetic"] rates: dict[str, dict[str, tuple[int, int]]] = defaultdict(dict) for name in ("task", "language", "length", "has_number", "tier"): totals, fails = Counter(), Counter() for r in real_records: value = features(r)[name] totals[value] += 1 fails[value] += not r["correct"] rates[name] = {v: (fails[v], totals[v]) for v in totals} confusion = Counter(f["features"]["confusion"] for f in failures if f["task"] == "category" and f["split"] != "synthetic") direction = Counter("under (too calm)" if PRIORITY_RANK[f["predicted"]] < PRIORITY_RANK[f["expected"]] else "over (false alarm)" for f in failures if f["task"] == "priority" and f["split"] != "synthetic") # Near-duplicate failures (synthetic copies of one template) inflate a cluster. distinct = {(f["id"][:6] if f["id"].startswith("SYN-") else f["id"], f["task"]) for f in failures} return {"failures": failures, "by_tag": by_tag, "rates": rates, "confusion": confusion, "direction": direction, "n_records": len(records), "n_distinct": len(distinct)} def report(result: dict) -> str: fails = result["failures"] real = [f for f in fails if f["features"]["source"] == "real"] lines = ["# Error analysis report", "", f"Records scored: {result['n_records']}. Failures: {len(fails)} " f"({len(real)} real, {len(fails) - len(real)} synthetic). " f"Distinct after merging synthetic copies of one template: {result['n_distinct']}.", f"Real failures read by a reviewer: {len(real)}; auto tags overridden: {sum(f['reviewed'] for f in fails)}.", "Clusters are sorted by REAL count. Synthetic counts only hint at what to look for.", "", "## Clusters by cause tag", "", "| Tag | All failures | Real | Share of real | Meaning |", "|---|---|---|---|---|"] real_count = Counter(tag for f in real for tag in f["tags"]) for tag, n in sorted(result["by_tag"].items(), key=lambda kv: (-real_count[kv[0]], -kv[1], kv[0])): n_real = real_count[tag] lines.append(f"| {tag} | {n} | {n_real} | {n_real / max(len(real), 1):.0%} | {TAXONOMY[tag]} |") lines += ["", "## Failure rate by feature (real records only)", "", "| Feature | Value | Failed / n | Rate |", "|---|---|---|---|"] for name, table in result["rates"].items(): for value, (k, n) in sorted(table.items(), key=lambda kv: -kv[1][0] / kv[1][1]): if n >= 5: lines.append(f"| {name} | {value} | {k}/{n} | {k / n:.0%} |") lines += ["", "## Top category confusions (real)", "", "| Expected -> predicted | Count |", "|---|---|"] lines += [f"| {pair} | {n} |" for pair, n in result["confusion"].most_common(6)] lines += ["", "## Priority errors by direction (real)", ""] lines += [f"- {d}: {n}" for d, n in result["direction"].most_common()] lines += ["", "## Examples from the three largest real clusters", ""] for tag, _ in real_count.most_common(3): lines.append(f"### {tag}") for f in [f for f in real if tag in f["tags"]][:3]: snippet = f["text"].replace("\n", " ")[:90] lines.append(f"- {f['id']} [{f['task']}] expected {f['expected']}, got {f['predicted']}: {snippet}") lines.append("") return "\n".join(lines) def main() -> None: RUNS.mkdir(exist_ok=True) records = run_baseline(load_tickets()) + run_baseline(load_synthetic()) overrides_path = RUNS / "tag_overrides.json" overrides = json.loads(overrides_path.read_text()) if overrides_path.exists() else {} result = analyze(records, overrides) text = report(result) (RUNS / "error_report.md").write_text(text + "\n", encoding="utf-8") with (RUNS / "failures.jsonl").open("w", encoding="utf-8") as f: for row in result["failures"]: f.write(json.dumps(row, ensure_ascii=False) + "\n") print(text) if __name__ == "__main__": main()Code explained
- In simple words: a triage nurse for failures: it takes a first guess at what went wrong, a human confirms or corrects it, and then it counts the patients by diagnosis.
- What happens:
TAXONOMYlists every cause tag with a one-line meaning that appears in the report. Four tags are specific to priority (question_rule,urgency_unseen,keyword_overfire,default_priority): together they say which part of the priority logic produced each wrong answer. Two tags (topic_not_intent,distractor_term) are never proposed automatically; only a reader can see them.VOCABis every word in the help center and the real tickets, used to detect misspellings.NEGATIONfinds words such as "not" and "never".rule_hitslists which category rules fire on a text;substring_onlychecks whether the first rule to fire matched only inside a longer word.featurescomputes explainable features per record: task, source (real or synthetic), language, length bucket, whether the text contains a number, customer tier, and the confusion pair.propose_tagsis the rule-based first pass. For the multi-intent check it counts question marks in the body only, because the subject usually restates the question (an earlier version counted both and tagged "How much is Business? ... What would Business cost?" as two intents).failure_keynames a failure asid:task, the key used in the overrides file.analyzecollects failures, applies overrides, counts tags, computes failure rates per feature value on real records only (a synthetic failure rate says nothing about production), counts category confusions and priority error directions on real records, and counts distinct failures after merging synthetic copies of one template.reportwrites the Markdown report, sorted by the real count of each cause, with examples from the three largest real clusters.mainruns the baseline over all 72 real tickets plus the 100 synthetic ones, appliesruns/tag_overrides.json, and writesruns/error_report.mdandruns/failures.jsonl.
- Comes out: the whole report, deterministic.text
# Error analysis report Records scored: 494. Failures: 239 (69 real, 170 synthetic). Distinct after merging synthetic copies of one template: 112. Real failures read by a reviewer: 69; auto tags overridden: 11. Clusters are sorted by REAL count. Synthetic counts only hint at what to look for. ## Clusters by cause tag | Tag | All failures | Real | Share of real | Meaning | |---|---|---|---|---| | question_rule | 40 | 20 | 29% | Priority: the 'short question means low' rule fired, but the customer needs real help. | | no_signal | 59 | 15 | 22% | Nothing in the text triggered any rule; the system fell back to its default. | | non_english | 54 | 14 | 20% | Ticket is not in English; the system was built from English examples. | | vocab_gap | 45 | 9 | 13% | Ticket and gold article share almost no words; keyword search cannot bridge it. | | urgency_unseen | 32 | 8 | 12% | Priority: gold is high or urgent, but no urgency phrase matched (stakes stated in other words). | | default_priority | 21 | 5 | 7% | Priority: no priority rule fired; the default 'normal' was wrong. | | near_miss | 15 | 5 | 7% | Gold article is in the top 3 but not ranked first. | | topic_not_intent | 5 | 5 | 7% | A keyword names the topic, but the customer's intent is different (human tag). | | keyword_overfire | 7 | 3 | 4% | Priority: an urgency phrase matched, but the situation is calmer than the phrase suggests. | | keyword_collision | 10 | 2 | 3% | Words for two or more categories appear; the first rule to match won. | | substring_match | 10 | 2 | 3% | A keyword matched inside a longer word ('plan' in 'plane', 'vat' in 'private'). | | distractor_term | 1 | 1 | 1% | A strong but irrelevant term pulled retrieval to the wrong article (human tag). | | label_question | 1 | 1 | 1% | A reviewer thinks the gold label itself is debatable. | | misspelling | 40 | 0 | 0% | Many words are not in the domain vocabulary (typos, slang). | | multi_intent | 16 | 0 | 0% | The ticket asks for two things; gold follows the main one. | | negation | 12 | 0 | 0% | A keyword appears but is denied or contrasted ('not a billing question'). | | other | 4 | 0 | 0% | No rule fired; a human must name the cause. | ## Failure rate by feature (real records only) | Feature | Value | Failed / n | Rate | |---|---|---|---| | task | priority | 34/72 | 47% | | task | category | 25/72 | 35% | | task | retrieval_top1 | 10/62 | 16% | | language | ja | 5/6 | 83% | | language | es | 4/6 | 67% | | language | de | 3/9 | 33% | | language | en | 55/182 | 30% | | length | short (<90 chars) | 18/53 | 34% | | length | medium | 51/153 | 33% | | has_number | False | 49/131 | 37% | | has_number | True | 20/75 | 27% | | tier | enterprise | 13/27 | 48% | | tier | business | 24/72 | 33% | | tier | team | 24/79 | 30% | | tier | free | 8/28 | 29% | ## Top category confusions (real) | Expected -> predicted | Count | |---|---| | account_access -> how_to | 5 | | billing -> how_to | 4 | | how_to -> billing | 3 | | how_to -> bug | 3 | | bug -> how_to | 3 | | feature_request -> how_to | 2 | ## Priority errors by direction (real) - under (too calm): 26 - over (false alarm): 8 ## Examples from the three largest real clusters ### question_rule - T-1002 [priority] expected normal, got low: Subject: How much is Business? We are 14 people. What would Business cost per month if we - T-1003 [priority] expected normal, got low: Subject: Cancel my subscription Please cancel our plan. We moved to another tool. How do - T-1006 [priority] expected normal, got low: Subject: VAT number on invoice Our accountant needs our VAT number DE811234567 on the inv ### no_signal - T-1027 [category] expected billing, got how_to: Subject: SLA credit Last Tuesday you were down for 3 hours. We are Enterprise. Are we owe - T-1032 [category] expected account_access, got how_to: Subject: Passwort vergessen Ich habe mein Passwort vergessen und der Link zum Zuruecksetz - T-1034 [category] expected cancellation, got how_to: Subject: サブスクリプションの解約 Teamプランを解約したいです。どこから手続きできますか? ### non_english - T-1031 [priority] expected high, got normal: Subject: Cobro duplicado Hola, me cobraron dos veces el plan Team este mes (factura INV-2 - T-1031 [retrieval_top1] expected billing-refunds, got billing-invoices: Subject: Cobro duplicado Hola, me cobraron dos veces el plan Team este mes (factura INV-2 - T-1032 [category] expected account_access, got how_to: Subject: Passwort vergessen Ich habe mein Passwort vergessen und der Link zum Zuruecksetz
Reading this report is the skill, so go slowly.
The top real cause is one rule.
question_ruleaccounts for 20 of 69 real failures (29%): the "short question means low priority" rule marked real requests as low. Two of those are serious. T-1063 ("I got 20 password reset emails I didn't request in the last hour. Is my account being attacked?") is goldurgent, and the rule called itlow. The direction table says the same thing from another angle: 26 of 34 real priority errors are too calm, the expensive direction for a support desk.The next two causes are about coverage, not bugs.
no_signal(15) means no category rule fired and the ticket fell tohow_to;non_english(14) overlaps it heavily. The feature table confirms it: 5 of 6 Japanese records and 4 of 6 Spanish records fail, against 55 of 182 English ones. More English keywords cannot fix that.Shares add up to more than 100% because a failure can carry two tags (T-1031 is
non_englishandurgency_unseen). Say so when you present the table.The synthetic set found causes that real data does not show.
misspelling(40),multi_intent(16), andnegation(12) have zero real failures. That is the correct use of synthetic data: they are hypotheses to watch for in production logs, not priorities for this week. If you had ranked by the "All failures" column,no_signalandnon_englishwould lead andmisspellingwould sit near the top, and you would spend a week on typos your real customers do not make.Rates need their denominators. Enterprise records fail 13 of 27 times (48%) against 24 of 79 for Team (30%). With 27 records that difference is suggestive, not proven; look at which Enterprise tickets failed before claiming the system is worse for big customers.
The overrides file is where your reading shows up. After reading all 69 real failures, 11 auto tags were corrected. Each override replaces the proposed tags for one failure.
jsonCopy
json{ "T-1001:retrieval_top1": ["distractor_term"], "T-1013:category": ["topic_not_intent"], "T-1014:category": ["topic_not_intent"], "T-1027:priority": ["question_rule", "urgency_unseen"], "T-1040:category": ["topic_not_intent"], "T-1044:category": ["label_question"], "T-1048:priority": ["keyword_overfire"], "T-1052:category": ["topic_not_intent"], "T-1057:category": ["topic_not_intent"], "T-1063:priority": ["question_rule", "urgency_unseen"], "T-1066:priority": ["urgency_unseen"] }Code explained
- In simple words: the reviewer's red pen: where the auto-tagger guessed wrong, this file says what really happened.
- What happens: each key is
ticket:taskand each value is the full list of tags that replaces the proposal. Some examples of why. T-1014 ("Our automations stopped running and there's a banner about a limit") was classed asbugbecause "stopped" is a bug keyword, but the customer is asking a how-to question about plan limits:topic_not_intent. T-1066 ("a charge of 48 USD ... but I only use the free plan") was auto-taggednegationbecause of "don't", but the real cause is that money-is-wrong urgency was stated without any trigger phrase:urgency_unseen. T-1063 and T-1027 keepquestion_ruleand gainurgency_unseen, because both the rule and the missing cue contributed. T-1044 (downgrade from Business to Team, goldbilling, predictedcancellation) is taggedlabel_question: a reasonable reviewer could call a downgrade a partial cancellation, and that belongs in your write-up, not silently in your score. - Comes out: the report's header line "auto tags overridden: 11" and the changed cluster counts.
To run this on your own system, write your system's predictions in the same JSONL format as
runs/baseline_run.jsonl(the fields listed in B1), pointanalyzeat it, read every failure, and fill in overrides. For an LLM system, add araw_outputfield so you can see what the model actually said; many "wrong category" failures turn out to be parse failures or refusals. The requirement is 50 real failures of your system. The synthetic set is optional, labelled, and never counts toward the 50.Situation Use this Why Your system fails 50 or more times on real tickets Real failures only, all read by a person The ranking of causes then reflects your traffic Your system fails fewer than 50 times on the 72 tickets Add real tickets first (anonymized history, new labelled samples), then a labelled synthetic set to probe variations Real data decides priorities; synthetic data suggests what to look for Hundreds of failures Read a random sample of 100 or more, tag them, and estimate each cluster's share with an interval Reading every one is not required once the cluster ranking is stable A cluster appears only in synthetic data Log it as a hypothesis and add a production check for it Fixing it now spends effort on a problem customers may not have B3. Cost and latency measured, not estimated
An estimate multiplies a guessed token count by a price. A measurement records what each real call reported (input, output, cached, and reasoning tokens) and how long it took by the wall clock, then aggregates. They differ for reasons you only see in measurements: retries double some calls, reasoning models emit hidden output tokens, and one task often takes several calls.
Report latency as p50 (the median call) and p95 (the value 95% of calls are at or below), because users feel the slow tail. Report cost per task (one ticket, however many calls it took), because that is what the business pays for. The
Meterwraps any callable with the signature ofsupportdesk.llm.chat, so the real helper,ScriptedLLM, and a local model all go through the same code.examples/cap_measure.py
python# examples/cap_measure.py """Capstone shared requirement 3: cost and latency measured, not estimated. Meter wraps any callable with the signature of supportdesk.llm.chat (the real helper, ScriptedLLM, or a local model adapter). For every call it records the usage the provider reported, the wall-clock latency the wrapper measured, and the dollar cost from supportdesk.pricing. It then aggregates per call and per TASK (one ticket may take several calls), with p50 and p95. """ from __future__ import annotations import json import math import time from collections import defaultdict from dataclasses import asdict, dataclass from pathlib import Path from typing import Any, Callable from supportdesk.llm import ChatResult, Usage from supportdesk.pricing import PRICES, cost_usd from supportdesk.tinylm import SamplingParams, generate, load def percentile(values: list[float], q: float) -> float: """Nearest-rank percentile: the smallest value with at least q% of values at or below it.""" if not values: return 0.0 ordered = sorted(values) rank = max(1, math.ceil(q / 100 * len(ordered))) return ordered[rank - 1] @dataclass class CallRecord: task_id: str step: str model: str provider: str input_tokens: int output_tokens: int cached_tokens: int reasoning_tokens: int wall_ms: float # measured here, around the whole call (includes client overhead and retries) reported_ms: float # ChatResult.latency_ms (what the helper measured) cost_usd: float | None # None when the model has no price and no price_as was given priced_as: str | None ok: bool error: str = "" class Meter: """Wrap an LLM callable and record usage, latency, and cost for every call.""" def __init__(self, llm: Callable[..., ChatResult], price_as: str | None = None, log_path: Path | str | None = None) -> None: self.llm = llm self.price_as = price_as self.log_path = Path(log_path) if log_path else None self.records: list[CallRecord] = [] def __call__(self, messages: list[dict[str, Any]], *, task_id: str, step: str = "call", **kwargs: Any) -> ChatResult: started = time.perf_counter() try: result = self.llm(messages, **kwargs) except Exception as exc: # record the failure, then let the caller decide self._log(CallRecord(task_id, step, "", "", 0, 0, 0, 0, round((time.perf_counter() - started) * 1000, 2), 0.0, None, None, False, repr(exc))) raise wall_ms = round((time.perf_counter() - started) * 1000, 2) priced_as = result.model if result.model in PRICES else self.price_as cost = cost_usd(result.usage, priced_as) if priced_as else None u = result.usage self._log(CallRecord(task_id, step, result.model, result.provider, u.input_tokens, u.output_tokens, u.cached_tokens, u.reasoning_tokens, wall_ms, result.latency_ms, cost, priced_as, True)) return result def _log(self, record: CallRecord) -> None: self.records.append(record) if self.log_path: with self.log_path.open("a", encoding="utf-8") as f: f.write(json.dumps(asdict(record)) + "\n") def summary(self) -> dict[str, Any]: ok = [r for r in self.records if r.ok] tasks: dict[str, list[CallRecord]] = defaultdict(list) for r in ok: tasks[r.task_id].append(r) task_ms = [sum(r.wall_ms for r in rs) for rs in tasks.values()] task_tokens = [sum(r.input_tokens + r.output_tokens for r in rs) for rs in tasks.values()] priced = all(r.cost_usd is not None for r in ok) task_cost = [sum(r.cost_usd or 0.0 for r in rs) for rs in tasks.values()] return { "calls": len(self.records), "errors": len(self.records) - len(ok), "tasks": len(tasks), "call_ms_p50": round(percentile([r.wall_ms for r in ok], 50), 2), "call_ms_p95": round(percentile([r.wall_ms for r in ok], 95), 2), "task_ms_p50": round(percentile(task_ms, 50), 2), "task_ms_p95": round(percentile(task_ms, 95), 2), "tokens_per_task_mean": round(sum(task_tokens) / max(len(tasks), 1), 1), "output_tokens_share": round(sum(r.output_tokens for r in ok) / max(sum(r.input_tokens + r.output_tokens for r in ok), 1), 3), "cost_per_task_mean": round(sum(task_cost) / max(len(tasks), 1), 8) if priced else None, "cost_per_task_p95": round(percentile(task_cost, 95), 8) if priced else None, "priced_as": sorted({r.priced_as for r in ok if r.priced_as}), "cached_tokens": sum(r.cached_tokens for r in ok), "warnings": self._warnings(ok), } @staticmethod def _warnings(ok: list[CallRecord]) -> list[str]: """Flag price-table entries that would overstate the cost of cached input.""" flagged = sorted({r.priced_as for r in ok if r.cached_tokens and r.priced_as and PRICES[r.priced_as].cached_input >= PRICES[r.priced_as].input}) return [f"{m}: pricing.py bills cached input at the full input rate; check the provider's cached price" for m in flagged] def check_budgets(summary: dict[str, Any], max_task_ms_p95: float, max_cost_per_task: float | None) -> list[str]: """Return the list of broken budgets (empty means within budget).""" broken = [] if summary["task_ms_p95"] > max_task_ms_p95: broken.append(f"task p95 {summary['task_ms_p95']} ms > budget {max_task_ms_p95} ms") if max_cost_per_task is not None and summary["cost_per_task_mean"] is not None \ and summary["cost_per_task_mean"] > max_cost_per_task: broken.append(f"mean cost/task {summary['cost_per_task_mean']} > budget {max_cost_per_task}") return broken class TinyChat: """Adapter that makes local TinyLM look like llm.chat, so Meter can measure it.""" def __init__(self, model_dir: str | None = None, max_new_tokens: int = 40) -> None: self.model, self.tok = load(model_dir) if model_dir else load() self.params = SamplingParams(max_new_tokens=max_new_tokens, temperature=0.0, stop=["\nCustomer"]) def __call__(self, messages: list[dict[str, Any]], **kwargs: Any) -> ChatResult: user = next(m["content"] for m in reversed(messages) if m["role"] == "user") prompt = f"Customer (Ana): {user}\nAgent (Dara):" started = time.perf_counter() gen = generate(self.model, self.tok, prompt, self.params) latency = (time.perf_counter() - started) * 1000 n_in = min(len(self.tok.encode(prompt).ids), self.model.cfg.context - self.params.max_new_tokens) return ChatResult(text=gen.text.strip(), usage=Usage(input_tokens=n_in, output_tokens=len(gen.token_ids)), latency_ms=round(latency, 1), finish_reason=gen.stop_reason, model="tinylm-base", provider="local")Code explained
- In simple words: a taxi meter you clip onto any model call: it notes distance (tokens), time, and fare for every trip, then summarizes by journey.
- What happens:
percentileuses the nearest-rank method: sort, then take the value at rank ceil(q/100 x n). With 24 tasks, p95 is the 23rd value, so one slow ticket moves it.CallRecordis one row per call: which task and step, model and provider, the four usage counts, the wall time measured by the Meter, the latency the helper reported, the cost, which price was used, and whether the call succeeded.Meter.__call__times the call, prices it withsupportdesk.pricing.cost_usd(by the model the result names if it is inPRICES, else byprice_as), and logs it. A failed call is recorded with its error and re-raised, so failures show up in the error count instead of vanishing.Meter.summarygroups calls by task and reports call and task p50 and p95, tokens per task, the output share of tokens, mean and p95 cost per task, the cached token total, and warnings.Meter._warningsflags a priced model whose cached-input price inpricing.pyis not lower than its input price while cached tokens were used. For Groq's gpt-oss modelspricing.pyhascached_inputequal toinput, which overstates cached cost by about 2 times given Groq's documented 50% discount. The Meter does not change the canonical price table; it tells you to check it.check_budgetsreturns the list of broken budgets, empty when everything is within limits, so a CI job can fail on it.TinyChatadapts TinyLM to thechatsignature: it builds a customer and agent prompt, generates greedily up to 24 new tokens (safely below the 128-token limit ofgenerate()), and returns aChatResultwith real token counts and latency.
- Comes out: nothing on its own; the demo below uses it.
The demo drives the Meter two ways. ScriptedLLM produces triage JSON from the keyword rules and a fixed draft, priced as if a provider billed its token counts (o200k estimates, since the stand-in has no provider tokenizer). TinyLM runs for real on this CPU, so its latency is a genuine measurement of local inference.
python# examples/cap_measure_demo.py """Run Meter for real: ScriptedLLM (plumbing, priced as if billed) and TinyLM (real local latency).""" from __future__ import annotations import json import sys from pathlib import Path import torch from supportdesk.data import load_tickets from supportdesk.stand_in import ScriptedLLM sys.path.insert(0, str(Path(__file__).resolve().parent)) from cap_baseline import keyword_triage # noqa: E402 from cap_measure import Meter, TinyChat, check_budgets # noqa: E402 torch.set_num_threads(1) tickets = load_tickets("test") SYSTEM = "You triage Brightlane support tickets. Reply with JSON: category, priority." def responder(messages, kwargs): """Stand-in: answers with the keyword rules. NOT a model; it only exercises the plumbing.""" text = messages[-1]["content"] if kwargs.get("max_tokens") == 120: # the drafting step return "Thanks for writing in. A support agent will review your request shortly." category, priority = keyword_triage(text) return json.dumps({"category": category, "priority": priority}) print("== ScriptedLLM: 2 calls per task (triage, then draft); token counts are o200k estimates") for price_as in ("openai/gpt-oss-120b", "gemini-3.5-flash"): meter = Meter(ScriptedLLM(responder=responder), price_as=price_as) for t in tickets: msgs = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": t.text}] meter(msgs, task_id=t.id, step="triage") meter(msgs, task_id=t.id, step="draft", max_tokens=120) s = meter.summary() print(f"priced as {price_as:20} tasks={s['tasks']} calls={s['calls']} tokens/task={s['tokens_per_task_mean']} " f"cost/task mean={s['cost_per_task_mean']:.8f} p95={s['cost_per_task_p95']:.8f} " f"task p95={s['task_ms_p95']} ms") print("\n== TinyLM (local, real latency on this CPU): 1 call per task, greedy, up to 24 new tokens") tiny = TinyChat(max_new_tokens=24) tiny([{"role": "user", "content": "warm-up call, not measured"}]) Path("runs/tinylm_calls.jsonl").unlink(missing_ok=True) meter = Meter(tiny, log_path="runs/tinylm_calls.jsonl") for t in tickets: meter([{"role": "user", "content": t.body}], task_id=t.id, step="reply") s = meter.summary() print({k: s[k] for k in ("tasks", "call_ms_p50", "call_ms_p95", "tokens_per_task_mean", "output_tokens_share", "cost_per_task_mean")}) print("broken budgets at p95 <= 250 ms:", check_budgets(s, max_task_ms_p95=250, max_cost_per_task=None) or "none") print("broken budgets at p95 <= 10 ms:", check_budgets(s, max_task_ms_p95=10, max_cost_per_task=None) or "none") first = meter.records[0] print("one call record:", {k: getattr(first, k) for k in ("task_id", "input_tokens", "output_tokens", "wall_ms", "reported_ms", "cost_usd")})Code explained
- In simple words: run the meter on a pretend two-step pipeline for cost arithmetic, then on a real small model for latency.
- What happens:
responderis the ScriptedLLM rule: keyword triage for the triage step, a fixed sentence for the draft step (recognized bymax_tokens=120). It is not a model.- For each of two price tables, every test ticket makes two metered calls under one task id, so the summary shows cost per task, not per call.
TinyChatmakes one warm-up call first (the first call pays one-time setup costs), then 24 measured calls logged toruns/tinylm_calls.jsonl.- The budgets show one met (p95 at most 250 ms) and one broken (p95 at most 10 ms) so you can see what a failure looks like.
- Comes out: token counts and costs are deterministic; latencies vary by machine and load.text
== ScriptedLLM: 2 calls per task (triage, then draft); token counts are o200k estimates priced as openai/gpt-oss-120b tasks=24 calls=48 tokens/task=126.1 cost/task mean=0.00003094 p95=0.00003240 task p95=0.17 ms priced as gemini-3.5-flash tasks=24 calls=48 tokens/task=126.1 cost/task mean=0.00038950 p95=0.00040500 task p95=0.06 ms == TinyLM (local, real latency on this CPU): 1 call per task, greedy, up to 24 new tokens {'tasks': 24, 'call_ms_p50': 24.23, 'call_ms_p95': 26.89, 'tokens_per_task_mean': 57.0, 'output_tokens_share': 0.398, 'cost_per_task_mean': None} broken budgets at p95 <= 250 ms: none broken budgets at p95 <= 10 ms: ['task p95 26.89 ms > budget 10 ms'] one call record: {'task_id': 'T-1003', 'input_tokens': 34, 'output_tokens': 24,
Three readings. The same token counts cost 12.6 times more priced as gemini-3.5-flash (0.00038950 USD per task) than as openai/gpt-oss-120b (0.00003094), which is the model-choice lever from Module 2 in one line; at 100,000 tickets a month that is about 39 versus 3.1 USD, before any real reasoning tokens. TinyLM's p50 and p95 sit close together here (p50 about 23 to 24 ms and p95 about 24 to 27 ms across runs) because every call generates exactly 24 tokens. And timings move with load: when this same script ran earlier on the same machine while other jobs were busy, it measured p50 106 ms and p95 129 ms. That is a factor of four with no code change, so record the machine and load with every latency number, and set budgets from several runs.
For your capstone, run the Meter around your real system with supportdesk.llm.chat. The usage fields then come from the provider, including reasoning tokens, and latency_ms includes the network. The output below shows the shape to expect.
Illustrative sample run (not captured in this build; produced for teaching). Your output will differ.
priced as openai/gpt-oss-120b tasks=24 calls=26 tokens/task=903.4 cost/task mean=0.00031520 p95=0.00058800 task p95=2210.0 ms
Code explained
- In simple words: what a real provider run tends to look like next to the stand-in: more calls, more tokens, a much longer tail.
- What happens: with a reasoning model, 26 calls for 24 tasks would mean two retries after invalid output, and tokens per task grow several-fold because reasoning tokens bill as output. These numbers are invented for illustration only.
- Comes out: use your own run's line in the evidence pack, never this one.
| Situation | Use this | Why |
|---|---|---|
| Deciding between models before building | Estimate: token counts from supportdesk.tokens times pricing.py | Cheap and good enough to rule options in or out |
| Reporting your capstone's cost | Meter around the real system, cost per task, p50 and p95 | Includes retries, reasoning tokens, and multi-call tasks, which estimates miss |
| Latency for a user-facing feature | p95 per task across several runs at realistic load, plus time to first token when streaming (Module 13) | Users feel the tail; one quiet run understates it |
| Local or self-hosted models | Wall time and throughput per task, and hardware cost per hour | No per-token bill; the cost is the machine |
B4. A written account of what did not work
Reviewers trust a capstone more when it shows its dead ends. The account is short: each entry says what you tried, why you expected it to help, what you measured, why it failed, and what you decided. A negative result with a clear measurement is evidence; a list of things that "didn't pan out" is not.
## What did not work
### <Attempt name>
- Tried: <the change, precisely enough to reproduce: file, commit, prompt version>
- Expected: <which failure cluster it targeted and the gain you predicted>
- Measured: <before -> after on dev and test, n, tickets fixed and broken, sign test p>
- Why it failed: <what the flipped tickets show, in one or two sentences>
- Decision: <kept, dropped, or kept for a different reason, and what you did next>Code explained
- In simple words: a lab-notebook entry with five fixed fields, so every dead end is recorded the same way.
- What happens: "Tried" makes it reproducible; "Expected" ties it to a failure cluster from B2, so you cannot quietly change the goal; "Measured" forces before and after numbers with the paired counts; "Why it failed" comes from reading the tickets that flipped; "Decision" closes the loop.
- Comes out: one entry per attempt in your evidence pack. Aim for three to six entries.
Here is a short worked example on the baseline. B2 pointed at two causes with obvious-looking fixes: word-boundary matching for substring_match ("plan" inside "plane"), and dropping the short-question rule for question_rule. The script measures both before anyone believes them.
# examples/cap_rule_fixes.py
"""Two obvious fixes for the two clusters error analysis pointed at, measured before believing them.
fix A word-boundary matching, aimed at the substring_match cluster ('plan' in 'plane')
fix B drop the 'short question means low priority' rule, aimed at the question_rule cluster
Each is scored on dev and test against the unchanged keyword baseline, with the
tickets that flipped listed, because a net score hides fixes that cancel out.
"""
from __future__ import annotations
import re
import sys
from pathlib import Path
from supportdesk.data import load_tickets
sys.path.insert(0, str(Path(__file__).resolve().parent))
from cap_baseline import CATEGORY_RULES, keyword_triage, sign_test_p # noqa: E402
def category_word_boundary(text: str) -> str:
low = text.lower()
return next((cat for cat, words in CATEGORY_RULES
if any(re.search(rf"\b{re.escape(w)}\b", low) for w in words)), "how_to")
def priority_no_question_rule(text: str) -> str:
priority = keyword_triage(text)[1]
return "normal" if priority == "low" else priority
FIXES = {
"A word boundaries (category)": ("category", lambda t: keyword_triage(t)[0], category_word_boundary),
"B no question rule (priority)": ("priority", lambda t: keyword_triage(t)[1], priority_no_question_rule),
}
def compare(name: str) -> dict:
field, before, after = FIXES[name]
out = {}
for split in ("dev", "test"):
tickets = load_tickets(split)
old = [before(t.text) == t.gold[field] for t in tickets]
new = [after(t.text) == t.gold[field] for t in tickets]
fixed = [t.id for t, o, n in zip(tickets, old, new) if n and not o]
broke = [t.id for t, o, n in zip(tickets, old, new) if o and not n]
out[split] = {"n": len(tickets), "before": sum(old), "after": sum(new), "fixed": fixed, "broke": broke,
"p": sign_test_p(len(fixed), len(broke))}
return out
def main() -> None:
for name in FIXES:
print(name)
for split, r in compare(name).items():
print(f" {split:4} {r['before']:2}/{r['n']} -> {r['after']:2}/{r['n']} fixed {r['fixed']} "
f"broke {r['broke']} sign test p = {r['p']:.2f}")
moved = [(t.id, t.gold["priority"]) for t in load_tickets()
if keyword_triage(t.text)[1] == "low" and t.gold["priority"] in ("high", "urgent")]
print("fix B also turns these high/urgent tickets from 'low' into 'normal' (less wrong, still wrong):", moved)
if __name__ == "__main__":
main()Code explained
- In simple words: try the two obvious repairs and count exactly which tickets each one fixes and breaks.
- What happens:
category_word_boundarymatches every keyword only as a whole word (\bon both sides).priority_no_question_rulereplaceslowwithnormal.comparescores each fix against the unchanged baseline on dev and test and lists the tickets that flipped, with the sign test on fixed versus broken. The last line lists high or urgent tickets the question rule had calledlow. - Comes out:
A word boundaries (category)
dev 31/48 -> 31/48 fixed ['T-1022'] broke ['T-1068'] sign test p = 0.75
test 16/24 -> 16/24 fixed ['T-1018'] broke ['T-1060'] sign test p = 0.75
B no question rule (priority)
dev 28/48 -> 28/48 fixed ['T-1002', 'T-1013', 'T-1014', 'T-1019', 'T-1020', 'T-1032', 'T-1041', 'T-1043', 'T-1044', 'T-1047', 'T-1065', 'T-1068'] broke ['T-1016', 'T-1022', 'T-1029', 'T-1035', 'T-1046', 'T-1052', 'T-1053', 'T-1056', 'T-1059', 'T-1064', 'T-1067', 'T-1070'] sign test p = 0.58
test 10/24 -> 11/24 fixed ['T-1003', 'T-1006', 'T-1036', 'T-1060', 'T-1069', 'T-1072'] broke ['T-1012', 'T-1015', 'T-1033', 'T-1042', 'T-1057'] sign test p = 0.50
fix B also turns these high/urgent tickets from 'low' into 'normal' (less wrong, still wrong): [('T-1027', 'high'), ('T-1063', 'urgent')]And the write-up those numbers support:
## Word-boundary matching for category keywords
- Tried: match each keyword as a whole word (examples/cap_rule_fixes.py, fix A).
- Expected: fix the substring_match cluster (2 real failures: "plan" in "plane", "vat" in "private").
- Measured: dev 31 -> 31/48 (fixed T-1022, broke T-1068); test 16 -> 16/24 (fixed T-1018, broke T-1060). p = 0.75.
- Why it failed: the substring matches were also doing useful work. "error" no longer matches "errors" (T-1068),
and "cancel" no longer matches the Spanish "cancelar" (T-1060), which the rules had been getting right by accident.
- Decision: dropped. Stemming would patch these two and break others; the cluster is small, so effort goes to
question_rule and non_english instead.
### Removing the short-question rule for priority
- Tried: never predict low from the question-mark rule (fix B).
- Expected: fix most of the 20 real question_rule failures.
- Measured: dev 28 -> 28/48 (12 fixed, 12 broken); test 10 -> 11/24 (6 fixed, 5 broken). p = 0.58 and 0.50.
- Why it failed: the rule is right about half the time. Short questions are often genuinely low ("Is it free for
a team of 3?"), and often not ("Is my account being attacked?"). Question form does not carry urgency.
- Decision: kept the rule for now, because removing it only trades errors. It does move T-1063 (urgent) and
T-1027 (high) from low to normal: less wrong, still wrong. Urgency needs meaning, not punctuation, which is
the case for an LLM or a trained classifier in the capstoneCode explained
- In simple words: the two dead ends above, written the way a reviewer wants to read them.
- What happens: every claim cites a measured number or a ticket id from the script's output; the "why" comes from reading the flipped tickets; each decision says where the effort went instead.
- Comes out: two entries ready to paste into an evidence pack. Your entries will be about your own system's attempts.
The lesson of the worked example is the lesson of the whole requirement: both fixes looked right from the cluster names, both did nothing measurable, and only the flipped tickets explained why.