CourseLarge Language Models · Capstone Project: Four Tracks · part 76 of 80
Part 76 · Capstone Project: Four Tracks

Part B: Shared Requirements

43 min read·22 Sept 2026

Every track must deliver the same four things. This part builds the tooling for each one, runs it on the course's own simple systems, and shows how to read the output. The demos use the keyword baseline's failures so they run here; your capstone must run the same tools on your own system's failures and calls.

B1. A baseline to beat

A baseline is the simplest reasonable system for the task, measured the same way your final system will be. It turns "our assistant is good" into "our assistant gets 19 of 24 where the baseline gets 16", and it sometimes reveals that the cheap system is already good enough. A majority-class floor (always predict the most common label) is the baseline below the baseline: anything that cannot beat it has learned nothing.

Brightlane gets two baselines. For triage, keyword rules guess a category and a priority. For retrieval, the BM25 search from Module 7 returns help-center articles, scored by whether the gold article comes first (top-1) or in the first three (top-3). Scores come with a Wilson interval, the 95% confidence interval for a proportion that behaves well at small n (Module 10 derived it).

The script writes one record per ticket per task. That record format is the contract for everything after it: the error-analysis tool reads it, so when you replace the baseline with your system, you write the same fields.

examples/cap_baseline.py

python
# examples/cap_baseline.py
"""Capstone shared requirement 1: a baseline the final system must beat.

Two baselines, both deterministic and free to run:
  triage     keyword rules that guess category and priority
  retrieval  BM25 top-1 and top-3 from supportdesk.kb_search
Scores come with Wilson 95% intervals because the test split is only 24 tickets.
Writes runs/baseline_run.jsonl (one record per ticket per task) and runs/baseline.json.
"""
from __future__ import annotations

import json
import math
from collections import Counter
from pathlib import Path

from supportdesk.data import Ticket, load_tickets
from supportdesk.kb_search import KBSearch

RUNS = Path("runs")

# Rules were written by reading the DEV split only. Order matters: first match wins.
CATEGORY_RULES: list[tuple[str, tuple[str, ...]]] = [
    ("cancellation", ("cancel", "money back", "stop now", "downgrade")),
    ("feature_request", ("please add", "will you support", "do you have a", "when will", "roadmap")),
    ("account_access", ("password", "locked", "2fa", "log in", "sign in", "sso", "saml", "gdpr",
                        "personal data", "region", "owner")),
    ("bug", ("not working", "stopped", "fails", "error", "broken", "outage", "not loading", "bug")),
    ("billing", ("charge", "charged", "invoice", "refund", "price", "cost", "pay", "vat", "plan")),
]
URGENT = ("nobody can", "not loading for anyone", "demo in", "outage", "all our users")
HIGH = ("charged twice", "locked", "lost my phone", "deleted", "can't access", "refund", "deal-breaker")


def keyword_triage(text: str) -> tuple[str, str]:
    """Guess (category, priority) from keywords. Falls back to how_to / normal."""
    low = text.lower()
    category = next((cat for cat, words in CATEGORY_RULES if any(w in low for w in words)), "how_to")
    if any(w in low for w in URGENT):
        priority = "urgent"
    elif any(w in low for w in HIGH):
        priority = "high"
    elif low.rstrip().endswith("?") and len(low) < 160:
        priority = "low"
    else:
        priority = "normal"
    return category, priority


def wilson(k: int, n: int, z: float = 1.96) -> tuple[float, float]:
    """Wilson score interval for k successes in n trials (safer than p +/- 2se at small n)."""
    if n == 0:
        return (0.0, 0.0)
    p = k / n
    centre = (p + z * z / (2 * n)) / (1 + z * z / n)
    half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
    return (round(max(0.0, centre - half), 3), round(min(1.0, centre + half), 3))


def sign_test_p(wins: int, losses: int) -> float:
    """Exact one-sided sign test (McNemar) on paired results.

    wins = items the new system gets right and the baseline gets wrong; losses =
    the reverse. Returns the chance of at least `wins` wins if both systems were
    equally good. Ties (both right or both wrong) carry no information.
    """
    n = wins + losses
    return 1.0 if n == 0 else sum(math.comb(n, k) for k in range(wins, n + 1)) / 2 ** n


def run_baseline(tickets: list[Ticket], search: KBSearch | None = None) -> list[dict]:
    """One record per ticket per task, in the format the error-analysis tool reads."""
    search = search or KBSearch()
    records = []
    for t in tickets:
        category, priority = keyword_triage(t.text)
        base = {"id": t.id, "split": t.split, "language": t.language, "tier": t.customer_tier, "text": t.text}
        records.append({**base, "task": "category", "expected": t.gold["category"], "predicted": category})
        records.append({**base, "task": "priority", "expected": t.gold["priority"], "predicted": priority})
        if t.gold["answerable"]:
            hits = [h.article_id for h in search.search(t.text, k=3)]
            records.append({**base, "task": "retrieval_top1", "expected": t.gold["kb_article"],
                            "predicted": hits[0] if hits else None, "top3": hits})
    for r in records:
        r["correct"] = r["predicted"] == r["expected"]
    return records


def score(records: list[dict], task: str, split: str | None = None) -> dict:
    rows = [r for r in records if r["task"] == task and (split is None or r["split"] == split)]
    if task == "retrieval_top1":
        top3 = sum(r["expected"] in r["top3"] for r in rows)
    k, n = sum(r["correct"] for r in rows), len(rows)
    out = {"task": task, "split": split or "all", "k": k, "n": n, "acc": round(k / n, 3), "ci95": wilson(k, n)}
    if task == "retrieval_top1":
        out["top3_k"] = top3
        out["top3_acc"] = round(top3 / n, 3)
        out["top3_ci95"] = wilson(top3, n)
    return out


def main() -> None:
    RUNS.mkdir(exist_ok=True)
    records = run_baseline(load_tickets())
    with (RUNS / "baseline_run.jsonl").open("w", encoding="utf-8") as f:
        for r in records:
            f.write(json.dumps(r, ensure_ascii=False) + "\n")
    summary = []
    for split in ("dev", "test"):
        for task in ("category", "priority", "retrieval_top1"):
            s = score(records, task, split)
            summary.append(s)
            extra = f"  top3 {s['top3_k']}/{s['n']} = {s['top3_acc']:.3f} {s['top3_ci95']}" if "top3_k" in s else ""
            print(f"{split:4} {task:15} {s['k']:2}/{s['n']:2} = {s['acc']:.3f}  95% CI {s['ci95']}{extra}")
    dev_counts = Counter(t.gold["category"] for t in load_tickets("dev"))
    top_class = dev_counts.most_common(1)[0][0]
    test = load_tickets("test")
    k = sum(t.gold["category"] == top_class for t in test)
    print(f"test majority-class floor (always {top_class!r}, the most common dev label): "
          f"{k}/{len(test)} = {k / len(test):.3f}  95% CI {wilson(k, len(test))}")
    (RUNS / "baseline.json").write_text(json.dumps(summary, indent=1) + "\n")
    print(f"wrote {len(records)} records to runs/baseline_run.jsonl")


if __name__ == "__main__":
    main()

Code explained

  • In simple words: a deliberately simple triage clerk with a list of keywords, a search box, and a ruler that always prints its error bars.
  • What happens:
    1. CATEGORY_RULES, URGENT, and HIGH are phrase lists written by reading the 48 dev tickets only. Order matters: the first category whose phrase appears wins, and how_to is the fallback.
    2. keyword_triage lowercases the text and applies the rules. Priority is urgent or high if a phrase matches, low if the ticket is a short question (ends with ?, under 160 characters), and normal otherwise. Substring matching (w in low) is a real weakness on purpose; B2 finds it.
    3. wilson returns the Wilson score interval for k successes in n trials, clipped to 0 and 1.
    4. sign_test_p is the exact one-sided sign test (the exact form of McNemar's test) on paired results: given the tickets a new system fixed (wins) and broke (losses), the chance of at least that many wins if both systems were equally good. Every track uses it to compare against the baseline.
    5. run_baseline produces records with id, split, language, tier, text, task, expected, predicted, and correct; retrieval records also keep the top-3 list. Only answerable tickets get a retrieval record.
    6. score aggregates one task on one split and adds top-3 numbers for retrieval.
    7. main scores dev and test, prints the majority-class floor for test (always how_to, the most common dev label), and writes runs/baseline_run.jsonl and runs/baseline.json.
    • Comes out: deterministic, so your numbers will match.

    Read it from the bottom up. The majority floor is 6 of 24; keyword category (16 of 24) clears it, and its interval (0.467 to 0.82) barely overlaps the floor's (0.12 to 0.449). Priority is weak at 10 of 24. Retrieval is strong in English: 17 of 20 answerable test tickets get the gold article first, 19 of 20 in the top 3, which matches the course-wide 52 of 62 top-1 (35 dev plus 17 test).

    Two things surprise people here. First, the rules score only 31 of 48 on the dev tickets they were written from, so they are not even overfit; they are simply weak. Second, those intervals are wide. A system that scores 19 of 24 on category (0.79) has an interval of about 0.6 to 0.9, which overlaps the baseline's. With 24 test tickets you can only claim a win for a large or consistent improvement, which is why every track compares systems ticket by ticket with sign_test_p rather than by comparing two percentages.

    Which baseline should your capstone beat? The strongest cheap one you have.

    SituationUse thisWhy
    Classification with a few dozen labelsKeyword rules and a majority floor (this script)Minutes to build, fully explainable, and a real bar: 16 of 24 here
    Classification with hundreds or thousands of labelled historical ticketsTF-IDF plus logistic regression (Module 14 Part C)Stronger than rules once labels are plentiful; the bar a fine-tune or LLM must clear
    Retrieval or grounded answersBM25 top-1 and top-3 (KBSearch)Hard to beat in the language it was built for; 17 of 20 on test
    Agent tasksThe same task done by a fixed script (search, read top article, draft)Shows what the agent's autonomy adds, step by step
    A new LLM feature replacing a human stepThe human's current accuracy and time, measured on a sampleThe business decision is against today's process, not against a toy

    B2. Error analysis on at least 50 failures, clustered by cause

    Error analysis means reading failures one by one and writing down why each happened, then counting. A cause taxonomy is the fixed list of reasons you tag with; a good tag names something you could act on ("the short-question rule fired") rather than a symptom ("priority too low"). Clustering groups failures that share a tag or a feature, so you fix the biggest group first. The requirement is at least 50 failures because with fewer, the ranking of clusters is mostly luck.

    The baseline fails on 69 real records (a ticket can fail on several tasks), which is enough. Many capstone systems will fail less often on 72 tickets, so the tooling also accepts an expanded synthetic set: tickets generated from templates to probe variations real data has too few of (typos, other languages, negation, two requests in one ticket). They carry a clear label, and the report keeps them apart, because their gold labels come from the template author, not from a support agent.

    examples/cap_synthetic_tickets.py

    python
    # examples/cap_synthetic_tickets.py
    """Generate a SYNTHETIC ticket set to widen error analysis (never to replace real failures).
    
    Every ticket here is produced from templates written for this capstone. Gold
    labels come from the template author, not from a support agent, so treat
    failure RATES on this set as meaningless and failure KINDS as hints about what
    to look for in real traffic. Ids start with SYN- and split is "synthetic".
    Writes data/synthetic_tickets.jsonl.
    """
    from __future__ import annotations
    
    import json
    import random
    from pathlib import Path
    
    from supportdesk.data import DATA_DIR, Ticket
    
    OUT = DATA_DIR / "synthetic_tickets.jsonl"
    
    # (variation kind, subject, body, language, category, priority, kb_article)
    TEMPLATES: list[tuple[str, str, str, str, str, str, str | None]] = [
        # paraphrase: same intent, none of the obvious keywords
        ("paraphrase", "Two identical payments", "My bank statement shows the same Brightlane amount taken on the same day, {n} USD each.", "en", "billing", "high", "billing-refunds"),
        ("paraphrase", "Can't get in", "The app keeps rejecting me at the sign-in screen after I typed the wrong thing a few times.", "en", "account_access", "high", "account-login"),
        ("paraphrase", "Leaving Brightlane", "We have decided to end our subscription at the end of the month. What are the steps?", "en", "cancellation", "normal", "billing-refunds"),
        ("paraphrase", "Nothing arrives in our channel", "Updates from our boards no longer show up in our team chat channel since {day}.", "en", "bug", "normal", "integrations-slack"),
        ("paraphrase", "Getting everything out", "Before we switch vendors we want a complete copy of our boards and files.", "en", "how_to", "normal", "exports-data"),
        ("paraphrase", "Wish list", "It would be great if timelines could show which tasks block others.", "en", "feature_request", "low", "feature-requests"),
        # typos and informal writing
        ("typo", "chargd 2x", "hi i got chargd twise for teem plan pls refnd", "en", "billing", "high", "billing-refunds"),
        ("typo", "pasword", "forgot pasword and the resett link is expird", "en", "account_access", "normal", "account-login"),
        ("typo", "cancle", "how do i cancle my subscripton??", "en", "cancellation", "normal", "billing-refunds"),
        ("typo", "exprt csv", "how to exprt a bord to csv", "en", "how_to", "low", "exports-data"),
        # languages the rules never saw
        ("non_english", "Doble cargo", "Me cobraron dos veces este mes. Quiero el reembolso.", "es", "billing", "high", "billing-refunds"),
        ("non_english", "Konto gesperrt", "Mein Konto ist nach mehreren falschen Passwoertern gesperrt.", "de", "account_access", "high", "account-login"),
        ("non_english", "Annuler l'abonnement", "Je voudrais annuler notre abonnement Team. Comment faire ?", "fr", "cancellation", "normal", "billing-refunds"),
        ("non_english", "Exportar quadro", "Como exporto um quadro para CSV?", "pt", "how_to", "low", "exports-data"),
        ("non_english", "Preis Business", "Was kostet der Business-Plan pro Nutzer im Monat?", "de", "billing", "low", "billing-plans"),
        # negation and contrast: the keyword is present but denied
        ("negation", "Not about billing", "This is not a billing question: how do I set up an automation that moves cards to Done?", "en", "how_to", "low", "boards-automations"),
        ("negation", "Not cancelling", "We are NOT cancelling, we just want to know where invoices are sent.", "en", "billing", "normal", "billing-invoices"),
        ("negation", "No error, just slow", "There is no error message, I just want to know how many automation runs Team includes.", "en", "how_to", "low", "boards-automations"),
        # two intents in one ticket
        ("multi_intent", "Refund and a feature idea", "Please refund the duplicate charge from {day}. Also, will you add dark mode?", "en", "billing", "high", "billing-refunds"),
        ("multi_intent", "Locked out, then export", "I was locked out yesterday, fine now. How do I export my board to CSV?", "en", "how_to", "low", "exports-data"),
        ("multi_intent", "SSO and price", "Does Business include SSO, and what does it cost for {n} users?", "en", "billing", "normal", "billing-plans"),
        # urgency expressed without trigger words
        ("hidden_urgency", "Whole team blocked", "None of our {n} people can open any board right now, we have a launch today.", "en", "bug", "urgent", "status-incidents"),
        ("hidden_urgency", "Security worry", "Someone from an unknown country signed in to our admin account an hour ago.", "en", "account_access", "urgent", None),
        # out of scope for the help center (unanswerable)
        ("unanswerable", "Custom contract", "Our procurement team needs net-90 payment terms written into the contract.", "en", "billing", "normal", None),
        ("unanswerable", "On-prem install", "Can we install Brightlane on our own servers?", "en", "feature_request", "normal", None),
    ]
    DAYS = ("Monday", "Tuesday", "yesterday", "last week")
    TIERS = ("free", "team", "business", "enterprise")
    
    
    def build(copies: int = 4, seed: int = 7) -> list[Ticket]:
        """`copies` filled-in variants of every template, with random numbers, days, and tiers."""
        rng = random.Random(seed)
        tickets = []
        for i, (kind, subject, body, lang, cat, prio, kb) in enumerate(TEMPLATES):
            for c in range(copies):
                filled = body.format(n=rng.choice((12, 24, 48, 96, 288)), day=rng.choice(DAYS))
                tickets.append(Ticket(
                    id=f"SYN-{i:02d}{c}", subject=subject, body=filled, customer_tier=rng.choice(TIERS),
                    language=lang, split="synthetic",
                    gold={"category": cat, "priority": prio, "kb_article": kb, "answerable": kb is not None,
                          "variation": kind, "synthetic": True},
                ))
        return tickets
    
    
    def load_synthetic() -> list[Ticket]:
        with OUT.open(encoding="utf-8") as f:
            return [Ticket(**json.loads(line)) for line in f]
    
    
    def main() -> None:
        tickets = build()
        with OUT.open("w", encoding="utf-8") as f:
            for t in tickets:
                f.write(json.dumps(t.__dict__, ensure_ascii=False) + "\n")
        kinds = sorted({t.gold["variation"] for t in tickets})
        print(f"wrote {len(tickets)} SYNTHETIC tickets to {OUT.relative_to(Path.cwd())}")
        print("variation kinds:", ", ".join(kinds))
    
    
    if __name__ == "__main__":
        main()
    

    Code explained

    • In simple words: a stencil kit that stamps out variations of known tricky tickets, each stamped "SYNTHETIC".
    • What happens:
      1. TEMPLATES holds 25 hand-written templates across seven variation kinds, each with its own gold category, priority, and article (or None when the help center cannot answer).
      2. build fills each template four times with random numbers, days, and customer tiers from a seeded generator, so the set is reproducible. Ids start with SYN- and the split is synthetic; the gold dict records the variation kind and synthetic: True.
      3. load_synthetic reads the saved file back as Ticket objects.
      4. main writes data/synthetic_tickets.jsonl in your copy (never in the canonical repository).
    • Comes out:

    Now the analysis tool. It proposes tags with simple rules, applies your corrections, clusters by tag and by features such as language, length, and customer tier, and writes a Markdown report. The auto-tagger exists to speed up reading, not to replace it.

    examples/cap_error_analysis.py

    python
    # examples/cap_error_analysis.py
    """Capstone shared requirement 2: error analysis on at least 50 failures, clustered by cause.
    
    Input: a run file (JSONL), one record per ticket per task with fields
    id, task, expected, predicted, correct, text, language, tier, split.
    Steps: collect failures, propose tags with simple rules, apply a human's
    overrides, cluster by tag and by text features, and write a Markdown report.
    
    The auto-tagger only PROPOSES causes to speed up reading. You still read every
    failure and correct the tags by hand (runs/tag_overrides.json). Use failures
    from YOUR system: the demo below runs on the keyword baseline.
    """
    from __future__ import annotations
    
    import json
    import re
    import sys
    from collections import Counter, defaultdict
    from pathlib import Path
    
    from supportdesk.data import load_articles, load_tickets
    from supportdesk.kb_search import tokenize
    
    sys.path.insert(0, str(Path(__file__).resolve().parent))
    from cap_baseline import CATEGORY_RULES, HIGH, URGENT, run_baseline  # noqa: E402  (same capstone, same folder)
    from cap_synthetic_tickets import load_synthetic  # noqa: E402
    
    RUNS = Path("runs")
    PRIORITY_RANK = {"low": 0, "normal": 1, "high": 2, "urgent": 3}
    
    # The taxonomy: every tag is a CAUSE you could act on, not a symptom.
    TAXONOMY: dict[str, str] = {
        "non_english": "Ticket is not in English; the system was built from English examples.",
        "no_signal": "Nothing in the text triggered any rule; the system fell back to its default.",
        "keyword_collision": "Words for two or more categories appear; the first rule to match won.",
        "substring_match": "A keyword matched inside a longer word ('plan' in 'plane', 'vat' in 'private').",
        "topic_not_intent": "A keyword names the topic, but the customer's intent is different (human tag).",
        "distractor_term": "A strong but irrelevant term pulled retrieval to the wrong article (human tag).",
        "negation": "A keyword appears but is denied or contrasted ('not a billing question').",
        "multi_intent": "The ticket asks for two things; gold follows the main one.",
        "misspelling": "Many words are not in the domain vocabulary (typos, slang).",
        "question_rule": "Priority: the 'short question means low' rule fired, but the customer needs real help.",
        "urgency_unseen": "Priority: gold is high or urgent, but no urgency phrase matched (stakes stated in other words).",
        "keyword_overfire": "Priority: an urgency phrase matched, but the situation is calmer than the phrase suggests.",
        "default_priority": "Priority: no priority rule fired; the default 'normal' was wrong.",
        "vocab_gap": "Ticket and gold article share almost no words; keyword search cannot bridge it.",
        "near_miss": "Gold article is in the top 3 but not ranked first.",
        "label_question": "A reviewer thinks the gold label itself is debatable.",
        "other": "No rule fired; a human must name the cause.",
    }
    
    VOCAB = {w for a in load_articles() for w in tokenize(a.title + " " + a.body)} | {
        w for t in load_tickets() for w in tokenize(t.text)}
    NEGATION = re.compile(r"\b(not|no|never|isn't|aren't|don't|NOT)\b", re.IGNORECASE)
    
    
    def rule_hits(text: str) -> list[str]:
        low = text.lower()
        return [cat for cat, words in CATEGORY_RULES if any(w in low for w in words)]
    
    
    def substring_only(text: str) -> bool:
        """True if the first rule that fired matched only inside longer words."""
        low = text.lower()
        for _, words in CATEGORY_RULES:
            found = [w for w in words if w in low]
            if found:
                return not any(re.search(rf"\b{re.escape(w)}\b", low) for w in found)
        return False
    
    
    def features(r: dict) -> dict:
        """Simple, explainable features of a record, used for clustering beyond tags."""
        text = r["text"]
        body_len = len(text)
        return {
            "task": r["task"],
            "source": "synthetic" if r["split"] == "synthetic" else "real",
            "language": r["language"],
            "length": "short (<90 chars)" if body_len < 90 else ("long (>180)" if body_len > 180 else "medium"),
            "has_number": str(bool(re.search(r"\d", text))),
            "tier": r["tier"],
            "confusion": f"{r['expected']} -> {r['predicted']}",
        }
    
    
    def propose_tags(r: dict) -> list[str]:
        """Heuristic first pass at WHY a record failed. A human confirms or overrides."""
        text, tags = r["text"], []
        words = [w for w in tokenize(text) if w.isalpha()]
        if r["language"] != "en":
            tags.append("non_english")
        if r["task"] in ("category", "priority"):
            hits = rule_hits(text)
            if r["task"] == "category":
                if not hits:
                    tags.append("no_signal")
                elif len(hits) > 1:
                    tags.append("keyword_collision")
                if hits and substring_only(text):
                    tags.append("substring_match")
            if NEGATION.search(text) and hits:
                tags.append("negation")
            body = text.split("\n\n", 1)[-1]  # the subject usually restates the question; count the body only
            if body.count("?") >= 2 or re.search(r"\balso\b|, then |\band what\b", body, re.IGNORECASE):
                tags.append("multi_intent")
            if r["language"] == "en" and words and sum(w not in VOCAB for w in words) / len(words) > 0.3:
                tags.append("misspelling")
        if r["task"] == "priority":
            low = text.lower()
            exp, got = PRIORITY_RANK[r["expected"]], PRIORITY_RANK[r["predicted"]]
            fired = any(w in low for w in URGENT + HIGH)
            if got > exp and fired:
                tags.append("keyword_overfire")
            elif r["predicted"] == "low":
                tags.append("question_rule")
            elif exp >= PRIORITY_RANK["high"] and not fired:
                tags.append("urgency_unseen")
            else:
                tags.append("default_priority")
        if r["task"] == "retrieval_top1":
            if r["expected"] in r.get("top3", []):
                tags.append("near_miss")
            gold = next(a for a in load_articles() if a.id == r["expected"])
            if len(set(tokenize(text)) & set(tokenize(gold.title + " " + gold.body))) <= 2:
                tags.append("vocab_gap")
        return tags or ["other"]
    
    
    def failure_key(r: dict) -> str:
        return f"{r['id']}:{r['task']}"
    
    
    def analyze(records: list[dict], overrides: dict[str, list[str]] | None = None) -> dict:
        overrides = overrides or {}
        failures = [dict(r) for r in records if not r["correct"]]
        for f in failures:
            f["tags"] = overrides.get(failure_key(f), propose_tags(f))
            f["reviewed"] = failure_key(f) in overrides
            f["features"] = features(f)
        by_tag = Counter(tag for f in failures for tag in f["tags"])
        # Failure RATE per feature value needs the successes too: a cluster is only
        # interesting when the rate there is higher than the rate overall.
        # Rates use REAL records only: synthetic gold labels come from a template author,
        # so a synthetic failure rate says nothing about production.
        real_records = [r for r in records if r["split"] != "synthetic"]
        rates: dict[str, dict[str, tuple[int, int]]] = defaultdict(dict)
        for name in ("task", "language", "length", "has_number", "tier"):
            totals, fails = Counter(), Counter()
            for r in real_records:
                value = features(r)[name]
                totals[value] += 1
                fails[value] += not r["correct"]
            rates[name] = {v: (fails[v], totals[v]) for v in totals}
        confusion = Counter(f["features"]["confusion"] for f in failures
                            if f["task"] == "category" and f["split"] != "synthetic")
        direction = Counter("under (too calm)" if PRIORITY_RANK[f["predicted"]] < PRIORITY_RANK[f["expected"]]
                            else "over (false alarm)" for f in failures
                            if f["task"] == "priority" and f["split"] != "synthetic")
        # Near-duplicate failures (synthetic copies of one template) inflate a cluster.
        distinct = {(f["id"][:6] if f["id"].startswith("SYN-") else f["id"], f["task"]) for f in failures}
        return {"failures": failures, "by_tag": by_tag, "rates": rates, "confusion": confusion, "direction": direction,
                "n_records": len(records), "n_distinct": len(distinct)}
    
    
    def report(result: dict) -> str:
        fails = result["failures"]
        real = [f for f in fails if f["features"]["source"] == "real"]
        lines = ["# Error analysis report", "",
                 f"Records scored: {result['n_records']}. Failures: {len(fails)} "
                 f"({len(real)} real, {len(fails) - len(real)} synthetic). "
                 f"Distinct after merging synthetic copies of one template: {result['n_distinct']}.",
                 f"Real failures read by a reviewer: {len(real)}; auto tags overridden: {sum(f['reviewed'] for f in fails)}.",
                 "Clusters are sorted by REAL count. Synthetic counts only hint at what to look for.", "",
                 "## Clusters by cause tag", "", "| Tag | All failures | Real | Share of real | Meaning |", "|---|---|---|---|---|"]
        real_count = Counter(tag for f in real for tag in f["tags"])
        for tag, n in sorted(result["by_tag"].items(), key=lambda kv: (-real_count[kv[0]], -kv[1], kv[0])):
            n_real = real_count[tag]
            lines.append(f"| {tag} | {n} | {n_real} | {n_real / max(len(real), 1):.0%} | {TAXONOMY[tag]} |")
        lines += ["", "## Failure rate by feature (real records only)", "",
                  "| Feature | Value | Failed / n | Rate |", "|---|---|---|---|"]
        for name, table in result["rates"].items():
            for value, (k, n) in sorted(table.items(), key=lambda kv: -kv[1][0] / kv[1][1]):
                if n >= 5:
                    lines.append(f"| {name} | {value} | {k}/{n} | {k / n:.0%} |")
        lines += ["", "## Top category confusions (real)", "", "| Expected -> predicted | Count |", "|---|---|"]
        lines += [f"| {pair} | {n} |" for pair, n in result["confusion"].most_common(6)]
        lines += ["", "## Priority errors by direction (real)", ""]
        lines += [f"- {d}: {n}" for d, n in result["direction"].most_common()]
        lines += ["", "## Examples from the three largest real clusters", ""]
        for tag, _ in real_count.most_common(3):
            lines.append(f"### {tag}")
            for f in [f for f in real if tag in f["tags"]][:3]:
                snippet = f["text"].replace("\n", " ")[:90]
                lines.append(f"- {f['id']} [{f['task']}] expected {f['expected']}, got {f['predicted']}: {snippet}")
            lines.append("")
        return "\n".join(lines)
    
    
    def main() -> None:
        RUNS.mkdir(exist_ok=True)
        records = run_baseline(load_tickets()) + run_baseline(load_synthetic())
        overrides_path = RUNS / "tag_overrides.json"
        overrides = json.loads(overrides_path.read_text()) if overrides_path.exists() else {}
        result = analyze(records, overrides)
        text = report(result)
        (RUNS / "error_report.md").write_text(text + "\n", encoding="utf-8")
        with (RUNS / "failures.jsonl").open("w", encoding="utf-8") as f:
            for row in result["failures"]:
                f.write(json.dumps(row, ensure_ascii=False) + "\n")
        print(text)
    
    
    if __name__ == "__main__":
        main()
    

    Code explained

    • In simple words: a triage nurse for failures: it takes a first guess at what went wrong, a human confirms or corrects it, and then it counts the patients by diagnosis.
    • What happens:
      1. TAXONOMY lists every cause tag with a one-line meaning that appears in the report. Four tags are specific to priority (question_rule, urgency_unseen, keyword_overfire, default_priority): together they say which part of the priority logic produced each wrong answer. Two tags (topic_not_intent, distractor_term) are never proposed automatically; only a reader can see them.
      2. VOCAB is every word in the help center and the real tickets, used to detect misspellings. NEGATION finds words such as "not" and "never".
      3. rule_hits lists which category rules fire on a text; substring_only checks whether the first rule to fire matched only inside a longer word.
      4. features computes explainable features per record: task, source (real or synthetic), language, length bucket, whether the text contains a number, customer tier, and the confusion pair.
      5. propose_tags is the rule-based first pass. For the multi-intent check it counts question marks in the body only, because the subject usually restates the question (an earlier version counted both and tagged "How much is Business? ... What would Business cost?" as two intents).
      6. failure_key names a failure as id:task, the key used in the overrides file.
      7. analyze collects failures, applies overrides, counts tags, computes failure rates per feature value on real records only (a synthetic failure rate says nothing about production), counts category confusions and priority error directions on real records, and counts distinct failures after merging synthetic copies of one template.
      8. report writes the Markdown report, sorted by the real count of each cause, with examples from the three largest real clusters.
      9. main runs the baseline over all 72 real tickets plus the 100 synthetic ones, applies runs/tag_overrides.json, and writes runs/error_report.md and runs/failures.jsonl.
    • Comes out: the whole report, deterministic.

      text
      # Error analysis report
      
      Records scored: 494. Failures: 239 (69 real, 170 synthetic). Distinct after merging synthetic copies of one template: 112.
      Real failures read by a reviewer: 69; auto tags overridden: 11.
      Clusters are sorted by REAL count. Synthetic counts only hint at what to look for.
      
      ## Clusters by cause tag
      
      | Tag | All failures | Real | Share of real | Meaning |
      |---|---|---|---|---|
      | question_rule | 40 | 20 | 29% | Priority: the 'short question means low' rule fired, but the customer needs real help. |
      | no_signal | 59 | 15 | 22% | Nothing in the text triggered any rule; the system fell back to its default. |
      | non_english | 54 | 14 | 20% | Ticket is not in English; the system was built from English examples. |
      | vocab_gap | 45 | 9 | 13% | Ticket and gold article share almost no words; keyword search cannot bridge it. |
      | urgency_unseen | 32 | 8 | 12% | Priority: gold is high or urgent, but no urgency phrase matched (stakes stated in other words). |
      | default_priority | 21 | 5 | 7% | Priority: no priority rule fired; the default 'normal' was wrong. |
      | near_miss | 15 | 5 | 7% | Gold article is in the top 3 but not ranked first. |
      | topic_not_intent | 5 | 5 | 7% | A keyword names the topic, but the customer's intent is different (human tag). |
      | keyword_overfire | 7 | 3 | 4% | Priority: an urgency phrase matched, but the situation is calmer than the phrase suggests. |
      | keyword_collision | 10 | 2 | 3% | Words for two or more categories appear; the first rule to match won. |
      | substring_match | 10 | 2 | 3% | A keyword matched inside a longer word ('plan' in 'plane', 'vat' in 'private'). |
      | distractor_term | 1 | 1 | 1% | A strong but irrelevant term pulled retrieval to the wrong article (human tag). |
      | label_question | 1 | 1 | 1% | A reviewer thinks the gold label itself is debatable. |
      | misspelling | 40 | 0 | 0% | Many words are not in the domain vocabulary (typos, slang). |
      | multi_intent | 16 | 0 | 0% | The ticket asks for two things; gold follows the main one. |
      | negation | 12 | 0 | 0% | A keyword appears but is denied or contrasted ('not a billing question'). |
      | other | 4 | 0 | 0% | No rule fired; a human must name the cause. |
      
      ## Failure rate by feature (real records only)
      
      | Feature | Value | Failed / n | Rate |
      |---|---|---|---|
      | task | priority | 34/72 | 47% |
      | task | category | 25/72 | 35% |
      | task | retrieval_top1 | 10/62 | 16% |
      | language | ja | 5/6 | 83% |
      | language | es | 4/6 | 67% |
      | language | de | 3/9 | 33% |
      | language | en | 55/182 | 30% |
      | length | short (<90 chars) | 18/53 | 34% |
      | length | medium | 51/153 | 33% |
      | has_number | False | 49/131 | 37% |
      | has_number | True | 20/75 | 27% |
      | tier | enterprise | 13/27 | 48% |
      | tier | business | 24/72 | 33% |
      | tier | team | 24/79 | 30% |
      | tier | free | 8/28 | 29% |
      
      ## Top category confusions (real)
      
      | Expected -> predicted | Count |
      |---|---|
      | account_access -> how_to | 5 |
      | billing -> how_to | 4 |
      | how_to -> billing | 3 |
      | how_to -> bug | 3 |
      | bug -> how_to | 3 |
      | feature_request -> how_to | 2 |
      
      ## Priority errors by direction (real)
      
      - under (too calm): 26
      - over (false alarm): 8
      
      ## Examples from the three largest real clusters
      
      ### question_rule
      - T-1002 [priority] expected normal, got low: Subject: How much is Business?  We are 14 people. What would Business cost per month if we
      - T-1003 [priority] expected normal, got low: Subject: Cancel my subscription  Please cancel our plan. We moved to another tool. How do 
      - T-1006 [priority] expected normal, got low: Subject: VAT number on invoice  Our accountant needs our VAT number DE811234567 on the inv
      
      ### no_signal
      - T-1027 [category] expected billing, got how_to: Subject: SLA credit  Last Tuesday you were down for 3 hours. We are Enterprise. Are we owe
      - T-1032 [category] expected account_access, got how_to: Subject: Passwort vergessen  Ich habe mein Passwort vergessen und der Link zum Zuruecksetz
      - T-1034 [category] expected cancellation, got how_to: Subject: サブスクリプションの解約  Teamプランを解約したいです。どこから手続きできますか?
      
      ### non_english
      - T-1031 [priority] expected high, got normal: Subject: Cobro duplicado  Hola, me cobraron dos veces el plan Team este mes (factura INV-2
      - T-1031 [retrieval_top1] expected billing-refunds, got billing-invoices: Subject: Cobro duplicado  Hola, me cobraron dos veces el plan Team este mes (factura INV-2
      - T-1032 [category] expected account_access, got how_to: Subject: Passwort vergessen  Ich habe mein Passwort vergessen und der Link zum Zuruecksetz
      

    Reading this report is the skill, so go slowly.

    The top real cause is one rule. question_rule accounts for 20 of 69 real failures (29%): the "short question means low priority" rule marked real requests as low. Two of those are serious. T-1063 ("I got 20 password reset emails I didn't request in the last hour. Is my account being attacked?") is gold urgent, and the rule called it low. The direction table says the same thing from another angle: 26 of 34 real priority errors are too calm, the expensive direction for a support desk.

    The next two causes are about coverage, not bugs. no_signal (15) means no category rule fired and the ticket fell to how_to; non_english (14) overlaps it heavily. The feature table confirms it: 5 of 6 Japanese records and 4 of 6 Spanish records fail, against 55 of 182 English ones. More English keywords cannot fix that.

    Shares add up to more than 100% because a failure can carry two tags (T-1031 is non_english and urgency_unseen). Say so when you present the table.

    The synthetic set found causes that real data does not show. misspelling (40), multi_intent (16), and negation (12) have zero real failures. That is the correct use of synthetic data: they are hypotheses to watch for in production logs, not priorities for this week. If you had ranked by the "All failures" column, no_signal and non_english would lead and misspelling would sit near the top, and you would spend a week on typos your real customers do not make.

    Rates need their denominators. Enterprise records fail 13 of 27 times (48%) against 24 of 79 for Team (30%). With 27 records that difference is suggestive, not proven; look at which Enterprise tickets failed before claiming the system is worse for big customers.

    The overrides file is where your reading shows up. After reading all 69 real failures, 11 auto tags were corrected. Each override replaces the proposed tags for one failure.

    jsonCopy

    json
    {
      "T-1001:retrieval_top1": ["distractor_term"],
      "T-1013:category": ["topic_not_intent"],
      "T-1014:category": ["topic_not_intent"],
      "T-1027:priority": ["question_rule", "urgency_unseen"],
      "T-1040:category": ["topic_not_intent"],
      "T-1044:category": ["label_question"],
      "T-1048:priority": ["keyword_overfire"],
      "T-1052:category": ["topic_not_intent"],
      "T-1057:category": ["topic_not_intent"],
      "T-1063:priority": ["question_rule", "urgency_unseen"],
      "T-1066:priority": ["urgency_unseen"]
    }
    

    Code explained

    • In simple words: the reviewer's red pen: where the auto-tagger guessed wrong, this file says what really happened.
    • What happens: each key is ticket:task and each value is the full list of tags that replaces the proposal. Some examples of why. T-1014 ("Our automations stopped running and there's a banner about a limit") was classed as bug because "stopped" is a bug keyword, but the customer is asking a how-to question about plan limits: topic_not_intent. T-1066 ("a charge of 48 USD ... but I only use the free plan") was auto-tagged negation because of "don't", but the real cause is that money-is-wrong urgency was stated without any trigger phrase: urgency_unseen. T-1063 and T-1027 keep question_rule and gain urgency_unseen, because both the rule and the missing cue contributed. T-1044 (downgrade from Business to Team, gold billing, predicted cancellation) is tagged label_question: a reasonable reviewer could call a downgrade a partial cancellation, and that belongs in your write-up, not silently in your score.
    • Comes out: the report's header line "auto tags overridden: 11" and the changed cluster counts.

    To run this on your own system, write your system's predictions in the same JSONL format as runs/baseline_run.jsonl (the fields listed in B1), point analyze at it, read every failure, and fill in overrides. For an LLM system, add a raw_output field so you can see what the model actually said; many "wrong category" failures turn out to be parse failures or refusals. The requirement is 50 real failures of your system. The synthetic set is optional, labelled, and never counts toward the 50.

    SituationUse thisWhy
    Your system fails 50 or more times on real ticketsReal failures only, all read by a personThe ranking of causes then reflects your traffic
    Your system fails fewer than 50 times on the 72 ticketsAdd real tickets first (anonymized history, new labelled samples), then a labelled synthetic set to probe variationsReal data decides priorities; synthetic data suggests what to look for
    Hundreds of failuresRead a random sample of 100 or more, tag them, and estimate each cluster's share with an intervalReading every one is not required once the cluster ranking is stable
    A cluster appears only in synthetic dataLog it as a hypothesis and add a production check for itFixing it now spends effort on a problem customers may not have

    B3. Cost and latency measured, not estimated

    An estimate multiplies a guessed token count by a price. A measurement records what each real call reported (input, output, cached, and reasoning tokens) and how long it took by the wall clock, then aggregates. They differ for reasons you only see in measurements: retries double some calls, reasoning models emit hidden output tokens, and one task often takes several calls.

    Report latency as p50 (the median call) and p95 (the value 95% of calls are at or below), because users feel the slow tail. Report cost per task (one ticket, however many calls it took), because that is what the business pays for. The Meter wraps any callable with the signature of supportdesk.llm.chat, so the real helper, ScriptedLLM, and a local model all go through the same code.

    examples/cap_measure.py

    python
    # examples/cap_measure.py
    """Capstone shared requirement 3: cost and latency measured, not estimated.
    
    Meter wraps any callable with the signature of supportdesk.llm.chat (the real
    helper, ScriptedLLM, or a local model adapter). For every call it records the
    usage the provider reported, the wall-clock latency the wrapper measured, and
    the dollar cost from supportdesk.pricing. It then aggregates per call and per
    TASK (one ticket may take several calls), with p50 and p95.
    """
    from __future__ import annotations
    
    import json
    import math
    import time
    from collections import defaultdict
    from dataclasses import asdict, dataclass
    from pathlib import Path
    from typing import Any, Callable
    
    from supportdesk.llm import ChatResult, Usage
    from supportdesk.pricing import PRICES, cost_usd
    from supportdesk.tinylm import SamplingParams, generate, load
    
    
    def percentile(values: list[float], q: float) -> float:
        """Nearest-rank percentile: the smallest value with at least q% of values at or below it."""
        if not values:
            return 0.0
        ordered = sorted(values)
        rank = max(1, math.ceil(q / 100 * len(ordered)))
        return ordered[rank - 1]
    
    
    @dataclass
    class CallRecord:
        task_id: str
        step: str
        model: str
        provider: str
        input_tokens: int
        output_tokens: int
        cached_tokens: int
        reasoning_tokens: int
        wall_ms: float          # measured here, around the whole call (includes client overhead and retries)
        reported_ms: float      # ChatResult.latency_ms (what the helper measured)
        cost_usd: float | None  # None when the model has no price and no price_as was given
        priced_as: str | None
        ok: bool
        error: str = ""
    
    
    class Meter:
        """Wrap an LLM callable and record usage, latency, and cost for every call."""
    
        def __init__(self, llm: Callable[..., ChatResult], price_as: str | None = None,
                     log_path: Path | str | None = None) -> None:
            self.llm = llm
            self.price_as = price_as
            self.log_path = Path(log_path) if log_path else None
            self.records: list[CallRecord] = []
    
        def __call__(self, messages: list[dict[str, Any]], *, task_id: str, step: str = "call", **kwargs: Any) -> ChatResult:
            started = time.perf_counter()
            try:
                result = self.llm(messages, **kwargs)
            except Exception as exc:  # record the failure, then let the caller decide
                self._log(CallRecord(task_id, step, "", "", 0, 0, 0, 0,
                                     round((time.perf_counter() - started) * 1000, 2), 0.0, None, None, False, repr(exc)))
                raise
            wall_ms = round((time.perf_counter() - started) * 1000, 2)
            priced_as = result.model if result.model in PRICES else self.price_as
            cost = cost_usd(result.usage, priced_as) if priced_as else None
            u = result.usage
            self._log(CallRecord(task_id, step, result.model, result.provider, u.input_tokens, u.output_tokens,
                                 u.cached_tokens, u.reasoning_tokens, wall_ms, result.latency_ms, cost, priced_as, True))
            return result
    
        def _log(self, record: CallRecord) -> None:
            self.records.append(record)
            if self.log_path:
                with self.log_path.open("a", encoding="utf-8") as f:
                    f.write(json.dumps(asdict(record)) + "\n")
    
        def summary(self) -> dict[str, Any]:
            ok = [r for r in self.records if r.ok]
            tasks: dict[str, list[CallRecord]] = defaultdict(list)
            for r in ok:
                tasks[r.task_id].append(r)
            task_ms = [sum(r.wall_ms for r in rs) for rs in tasks.values()]
            task_tokens = [sum(r.input_tokens + r.output_tokens for r in rs) for rs in tasks.values()]
            priced = all(r.cost_usd is not None for r in ok)
            task_cost = [sum(r.cost_usd or 0.0 for r in rs) for rs in tasks.values()]
            return {
                "calls": len(self.records), "errors": len(self.records) - len(ok), "tasks": len(tasks),
                "call_ms_p50": round(percentile([r.wall_ms for r in ok], 50), 2),
                "call_ms_p95": round(percentile([r.wall_ms for r in ok], 95), 2),
                "task_ms_p50": round(percentile(task_ms, 50), 2),
                "task_ms_p95": round(percentile(task_ms, 95), 2),
                "tokens_per_task_mean": round(sum(task_tokens) / max(len(tasks), 1), 1),
                "output_tokens_share": round(sum(r.output_tokens for r in ok) / max(sum(r.input_tokens + r.output_tokens for r in ok), 1), 3),
                "cost_per_task_mean": round(sum(task_cost) / max(len(tasks), 1), 8) if priced else None,
                "cost_per_task_p95": round(percentile(task_cost, 95), 8) if priced else None,
                "priced_as": sorted({r.priced_as for r in ok if r.priced_as}),
                "cached_tokens": sum(r.cached_tokens for r in ok),
                "warnings": self._warnings(ok),
            }
    
        @staticmethod
        def _warnings(ok: list[CallRecord]) -> list[str]:
            """Flag price-table entries that would overstate the cost of cached input."""
            flagged = sorted({r.priced_as for r in ok if r.cached_tokens and r.priced_as
                              and PRICES[r.priced_as].cached_input >= PRICES[r.priced_as].input})
            return [f"{m}: pricing.py bills cached input at the full input rate; check the provider's cached price"
                    for m in flagged]
    
    
    def check_budgets(summary: dict[str, Any], max_task_ms_p95: float, max_cost_per_task: float | None) -> list[str]:
        """Return the list of broken budgets (empty means within budget)."""
        broken = []
        if summary["task_ms_p95"] > max_task_ms_p95:
            broken.append(f"task p95 {summary['task_ms_p95']} ms > budget {max_task_ms_p95} ms")
        if max_cost_per_task is not None and summary["cost_per_task_mean"] is not None \
                and summary["cost_per_task_mean"] > max_cost_per_task:
            broken.append(f"mean cost/task {summary['cost_per_task_mean']} > budget {max_cost_per_task}")
        return broken
    
    
    class TinyChat:
        """Adapter that makes local TinyLM look like llm.chat, so Meter can measure it."""
    
        def __init__(self, model_dir: str | None = None, max_new_tokens: int = 40) -> None:
            self.model, self.tok = load(model_dir) if model_dir else load()
            self.params = SamplingParams(max_new_tokens=max_new_tokens, temperature=0.0, stop=["\nCustomer"])
    
        def __call__(self, messages: list[dict[str, Any]], **kwargs: Any) -> ChatResult:
            user = next(m["content"] for m in reversed(messages) if m["role"] == "user")
            prompt = f"Customer (Ana): {user}\nAgent (Dara):"
            started = time.perf_counter()
            gen = generate(self.model, self.tok, prompt, self.params)
            latency = (time.perf_counter() - started) * 1000
            n_in = min(len(self.tok.encode(prompt).ids), self.model.cfg.context - self.params.max_new_tokens)
            return ChatResult(text=gen.text.strip(), usage=Usage(input_tokens=n_in, output_tokens=len(gen.token_ids)),
                              latency_ms=round(latency, 1), finish_reason=gen.stop_reason, model="tinylm-base", provider="local")
    

    Code explained

    • In simple words: a taxi meter you clip onto any model call: it notes distance (tokens), time, and fare for every trip, then summarizes by journey.
    • What happens:
      1. percentile uses the nearest-rank method: sort, then take the value at rank ceil(q/100 x n). With 24 tasks, p95 is the 23rd value, so one slow ticket moves it.
      2. CallRecord is one row per call: which task and step, model and provider, the four usage counts, the wall time measured by the Meter, the latency the helper reported, the cost, which price was used, and whether the call succeeded.
      3. Meter.__call__ times the call, prices it with supportdesk.pricing.cost_usd (by the model the result names if it is in PRICES, else by price_as), and logs it. A failed call is recorded with its error and re-raised, so failures show up in the error count instead of vanishing.
      4. Meter.summary groups calls by task and reports call and task p50 and p95, tokens per task, the output share of tokens, mean and p95 cost per task, the cached token total, and warnings.
      5. Meter._warnings flags a priced model whose cached-input price in pricing.py is not lower than its input price while cached tokens were used. For Groq's gpt-oss models pricing.py has cached_input equal to input, which overstates cached cost by about 2 times given Groq's documented 50% discount. The Meter does not change the canonical price table; it tells you to check it.
      6. check_budgets returns the list of broken budgets, empty when everything is within limits, so a CI job can fail on it.
      7. TinyChat adapts TinyLM to the chat signature: it builds a customer and agent prompt, generates greedily up to 24 new tokens (safely below the 128-token limit of generate()), and returns a ChatResult with real token counts and latency.
      • Comes out: nothing on its own; the demo below uses it.

      The demo drives the Meter two ways. ScriptedLLM produces triage JSON from the keyword rules and a fixed draft, priced as if a provider billed its token counts (o200k estimates, since the stand-in has no provider tokenizer). TinyLM runs for real on this CPU, so its latency is a genuine measurement of local inference.

      python
      # examples/cap_measure_demo.py
      """Run Meter for real: ScriptedLLM (plumbing, priced as if billed) and TinyLM (real local latency)."""
      from __future__ import annotations
      
      import json
      import sys
      from pathlib import Path
      
      import torch
      
      from supportdesk.data import load_tickets
      from supportdesk.stand_in import ScriptedLLM
      
      sys.path.insert(0, str(Path(__file__).resolve().parent))
      from cap_baseline import keyword_triage  # noqa: E402
      from cap_measure import Meter, TinyChat, check_budgets  # noqa: E402
      
      torch.set_num_threads(1)
      tickets = load_tickets("test")
      SYSTEM = "You triage Brightlane support tickets. Reply with JSON: category, priority."
      
      
      def responder(messages, kwargs):
          """Stand-in: answers with the keyword rules. NOT a model; it only exercises the plumbing."""
          text = messages[-1]["content"]
          if kwargs.get("max_tokens") == 120:  # the drafting step
              return "Thanks for writing in. A support agent will review your request shortly."
          category, priority = keyword_triage(text)
          return json.dumps({"category": category, "priority": priority})
      
      
      print("== ScriptedLLM: 2 calls per task (triage, then draft); token counts are o200k estimates")
      for price_as in ("openai/gpt-oss-120b", "gemini-3.5-flash"):
          meter = Meter(ScriptedLLM(responder=responder), price_as=price_as)
          for t in tickets:
              msgs = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": t.text}]
              meter(msgs, task_id=t.id, step="triage")
              meter(msgs, task_id=t.id, step="draft", max_tokens=120)
          s = meter.summary()
          print(f"priced as {price_as:20} tasks={s['tasks']} calls={s['calls']} tokens/task={s['tokens_per_task_mean']} "
                f"cost/task mean={s['cost_per_task_mean']:.8f} p95={s['cost_per_task_p95']:.8f} "
                f"task p95={s['task_ms_p95']} ms")
      
      print("\n== TinyLM (local, real latency on this CPU): 1 call per task, greedy, up to 24 new tokens")
      tiny = TinyChat(max_new_tokens=24)
      tiny([{"role": "user", "content": "warm-up call, not measured"}])
      Path("runs/tinylm_calls.jsonl").unlink(missing_ok=True)
      meter = Meter(tiny, log_path="runs/tinylm_calls.jsonl")
      for t in tickets:
          meter([{"role": "user", "content": t.body}], task_id=t.id, step="reply")
      s = meter.summary()
      print({k: s[k] for k in ("tasks", "call_ms_p50", "call_ms_p95", "tokens_per_task_mean", "output_tokens_share", "cost_per_task_mean")})
      print("broken budgets at p95 <= 250 ms:", check_budgets(s, max_task_ms_p95=250, max_cost_per_task=None) or "none")
      print("broken budgets at p95 <= 10 ms:", check_budgets(s, max_task_ms_p95=10, max_cost_per_task=None) or "none")
      first = meter.records[0]
      print("one call record:", {k: getattr(first, k) for k in ("task_id", "input_tokens", "output_tokens", "wall_ms", "reported_ms", "cost_usd")})
      

      Code explained

      • In simple words: run the meter on a pretend two-step pipeline for cost arithmetic, then on a real small model for latency.
      • What happens:
        1. responder is the ScriptedLLM rule: keyword triage for the triage step, a fixed sentence for the draft step (recognized by max_tokens=120). It is not a model.
        2. For each of two price tables, every test ticket makes two metered calls under one task id, so the summary shows cost per task, not per call.
        3. TinyChat makes one warm-up call first (the first call pays one-time setup costs), then 24 measured calls logged to runs/tinylm_calls.jsonl.
        4. The budgets show one met (p95 at most 250 ms) and one broken (p95 at most 10 ms) so you can see what a failure looks like.
      • Comes out: token counts and costs are deterministic; latencies vary by machine and load.

        text
        == ScriptedLLM: 2 calls per task (triage, then draft); token counts are o200k estimates
        priced as openai/gpt-oss-120b  tasks=24 calls=48 tokens/task=126.1 cost/task mean=0.00003094 p95=0.00003240 task p95=0.17 ms
        priced as gemini-3.5-flash     tasks=24 calls=48 tokens/task=126.1 cost/task mean=0.00038950 p95=0.00040500 task p95=0.06 ms
        
        == TinyLM (local, real latency on this CPU): 1 call per task, greedy, up to 24 new tokens
        {'tasks': 24, 'call_ms_p50': 24.23, 'call_ms_p95': 26.89, 'tokens_per_task_mean': 57.0, 'output_tokens_share': 0.398, 'cost_per_task_mean': None}
        broken budgets at p95 <= 250 ms: none
        broken budgets at p95 <= 10 ms: ['task p95 26.89 ms > budget 10 ms']
        one call record: {'task_id': 'T-1003', 'input_tokens': 34, 'output_tokens': 24,

Three readings. The same token counts cost 12.6 times more priced as gemini-3.5-flash (0.00038950 USD per task) than as openai/gpt-oss-120b (0.00003094), which is the model-choice lever from Module 2 in one line; at 100,000 tickets a month that is about 39 versus 3.1 USD, before any real reasoning tokens. TinyLM's p50 and p95 sit close together here (p50 about 23 to 24 ms and p95 about 24 to 27 ms across runs) because every call generates exactly 24 tokens. And timings move with load: when this same script ran earlier on the same machine while other jobs were busy, it measured p50 106 ms and p95 129 ms. That is a factor of four with no code change, so record the machine and load with every latency number, and set budgets from several runs.

For your capstone, run the Meter around your real system with supportdesk.llm.chat. The usage fields then come from the provider, including reasoning tokens, and latency_ms includes the network. The output below shows the shape to expect.

Illustrative sample run (not captured in this build; produced for teaching). Your output will differ.

text
priced as openai/gpt-oss-120b  tasks=24 calls=26 tokens/task=903.4 cost/task mean=0.00031520 p95=0.00058800 task p95=2210.0 ms

Code explained

  • In simple words: what a real provider run tends to look like next to the stand-in: more calls, more tokens, a much longer tail.
  • What happens: with a reasoning model, 26 calls for 24 tasks would mean two retries after invalid output, and tokens per task grow several-fold because reasoning tokens bill as output. These numbers are invented for illustration only.
  • Comes out: use your own run's line in the evidence pack, never this one.
SituationUse thisWhy
Deciding between models before buildingEstimate: token counts from supportdesk.tokens times pricing.pyCheap and good enough to rule options in or out
Reporting your capstone's costMeter around the real system, cost per task, p50 and p95Includes retries, reasoning tokens, and multi-call tasks, which estimates miss
Latency for a user-facing featurep95 per task across several runs at realistic load, plus time to first token when streaming (Module 13)Users feel the tail; one quiet run understates it
Local or self-hosted modelsWall time and throughput per task, and hardware cost per hourNo per-token bill; the cost is the machine

B4. A written account of what did not work

Reviewers trust a capstone more when it shows its dead ends. The account is short: each entry says what you tried, why you expected it to help, what you measured, why it failed, and what you decided. A negative result with a clear measurement is evidence; a list of things that "didn't pan out" is not.


text
## What did not work

### <Attempt name>
- Tried: <the change, precisely enough to reproduce: file, commit, prompt version>
- Expected: <which failure cluster it targeted and the gain you predicted>
- Measured: <before -> after on dev and test, n, tickets fixed and broken, sign test p>
- Why it failed: <what the flipped tickets show, in one or two sentences>
- Decision: <kept, dropped, or kept for a different reason, and what you did next>

Code explained

  • In simple words: a lab-notebook entry with five fixed fields, so every dead end is recorded the same way.
  • What happens: "Tried" makes it reproducible; "Expected" ties it to a failure cluster from B2, so you cannot quietly change the goal; "Measured" forces before and after numbers with the paired counts; "Why it failed" comes from reading the tickets that flipped; "Decision" closes the loop.
  • Comes out: one entry per attempt in your evidence pack. Aim for three to six entries.

Here is a short worked example on the baseline. B2 pointed at two causes with obvious-looking fixes: word-boundary matching for substring_match ("plan" inside "plane"), and dropping the short-question rule for question_rule. The script measures both before anyone believes them.

python
# examples/cap_rule_fixes.py
"""Two obvious fixes for the two clusters error analysis pointed at, measured before believing them.

  fix A  word-boundary matching, aimed at the substring_match cluster ('plan' in 'plane')
  fix B  drop the 'short question means low priority' rule, aimed at the question_rule cluster
Each is scored on dev and test against the unchanged keyword baseline, with the
tickets that flipped listed, because a net score hides fixes that cancel out.
"""
from __future__ import annotations

import re
import sys
from pathlib import Path

from supportdesk.data import load_tickets

sys.path.insert(0, str(Path(__file__).resolve().parent))
from cap_baseline import CATEGORY_RULES, keyword_triage, sign_test_p  # noqa: E402


def category_word_boundary(text: str) -> str:
    low = text.lower()
    return next((cat for cat, words in CATEGORY_RULES
                 if any(re.search(rf"\b{re.escape(w)}\b", low) for w in words)), "how_to")


def priority_no_question_rule(text: str) -> str:
    priority = keyword_triage(text)[1]
    return "normal" if priority == "low" else priority


FIXES = {
    "A word boundaries (category)": ("category", lambda t: keyword_triage(t)[0], category_word_boundary),
    "B no question rule (priority)": ("priority", lambda t: keyword_triage(t)[1], priority_no_question_rule),
}


def compare(name: str) -> dict:
    field, before, after = FIXES[name]
    out = {}
    for split in ("dev", "test"):
        tickets = load_tickets(split)
        old = [before(t.text) == t.gold[field] for t in tickets]
        new = [after(t.text) == t.gold[field] for t in tickets]
        fixed = [t.id for t, o, n in zip(tickets, old, new) if n and not o]
        broke = [t.id for t, o, n in zip(tickets, old, new) if o and not n]
        out[split] = {"n": len(tickets), "before": sum(old), "after": sum(new), "fixed": fixed, "broke": broke,
                      "p": sign_test_p(len(fixed), len(broke))}
    return out


def main() -> None:
    for name in FIXES:
        print(name)
        for split, r in compare(name).items():
            print(f"   {split:4} {r['before']:2}/{r['n']} -> {r['after']:2}/{r['n']}  fixed {r['fixed']}  "
                  f"broke {r['broke']}  sign test p = {r['p']:.2f}")
    moved = [(t.id, t.gold["priority"]) for t in load_tickets()
             if keyword_triage(t.text)[1] == "low" and t.gold["priority"] in ("high", "urgent")]
    print("fix B also turns these high/urgent tickets from 'low' into 'normal' (less wrong, still wrong):", moved)


if __name__ == "__main__":
    main()

Code explained

  • In simple words: try the two obvious repairs and count exactly which tickets each one fixes and breaks.
  • What happens: category_word_boundary matches every keyword only as a whole word (\b on both sides). priority_no_question_rule replaces low with normal. compare scores each fix against the unchanged baseline on dev and test and lists the tickets that flipped, with the sign test on fixed versus broken. The last line lists high or urgent tickets the question rule had called low.
  • Comes out:
text
A word boundaries (category)
   dev  31/48 -> 31/48  fixed ['T-1022']  broke ['T-1068']  sign test p = 0.75
   test 16/24 -> 16/24  fixed ['T-1018']  broke ['T-1060']  sign test p = 0.75
B no question rule (priority)
   dev  28/48 -> 28/48  fixed ['T-1002', 'T-1013', 'T-1014', 'T-1019', 'T-1020', 'T-1032', 'T-1041', 'T-1043', 'T-1044', 'T-1047', 'T-1065', 'T-1068']  broke ['T-1016', 'T-1022', 'T-1029', 'T-1035', 'T-1046', 'T-1052', 'T-1053', 'T-1056', 'T-1059', 'T-1064', 'T-1067', 'T-1070']  sign test p = 0.58
   test 10/24 -> 11/24  fixed ['T-1003', 'T-1006', 'T-1036', 'T-1060', 'T-1069', 'T-1072']  broke ['T-1012', 'T-1015', 'T-1033', 'T-1042', 'T-1057']  sign test p = 0.50
fix B also turns these high/urgent tickets from 'low' into 'normal' (less wrong, still wrong): [('T-1027', 'high'), ('T-1063', 'urgent')]

And the write-up those numbers support:

text
## Word-boundary matching for category keywords
- Tried: match each keyword as a whole word (examples/cap_rule_fixes.py, fix A).
- Expected: fix the substring_match cluster (2 real failures: "plan" in "plane", "vat" in "private").
- Measured: dev 31 -> 31/48 (fixed T-1022, broke T-1068); test 16 -> 16/24 (fixed T-1018, broke T-1060). p = 0.75.
- Why it failed: the substring matches were also doing useful work. "error" no longer matches "errors" (T-1068),
  and "cancel" no longer matches the Spanish "cancelar" (T-1060), which the rules had been getting right by accident.
- Decision: dropped. Stemming would patch these two and break others; the cluster is small, so effort goes to
  question_rule and non_english instead.

### Removing the short-question rule for priority
- Tried: never predict low from the question-mark rule (fix B).
- Expected: fix most of the 20 real question_rule failures.
- Measured: dev 28 -> 28/48 (12 fixed, 12 broken); test 10 -> 11/24 (6 fixed, 5 broken). p = 0.58 and 0.50.
- Why it failed: the rule is right about half the time. Short questions are often genuinely low ("Is it free for
  a team of 3?"), and often not ("Is my account being attacked?"). Question form does not carry urgency.
- Decision: kept the rule for now, because removing it only trades errors. It does move T-1063 (urgent) and
  T-1027 (high) from low to normal: less wrong, still wrong. Urgency needs meaning, not punctuation, which is
  the case for an LLM or a trained classifier in the capstone

Code explained

  • In simple words: the two dead ends above, written the way a reviewer wants to read them.
  • What happens: every claim cites a measured number or a ticket id from the script's output; the "why" comes from reading the flipped tickets; each decision says where the effort went instead.
  • Comes out: two entries ready to paste into an evidence pack. Your entries will be about your own system's attempts.

The lesson of the worked example is the lesson of the whole requirement: both fixes looked right from the cluster names, both did nothing measurable, and only the flipped tickets explained why.