Part C: Judgement
The most valuable skill in applied LLM work is knowing when not to use one. This part gives you three tools for that judgement: a list of problem shapes that LLMs handle badly, a real head-to-head against classical methods on the Brightlane data, and an expected-cost calculation that says whether a use case clears the bar.
C1. Problems LLMs should not be used for
Some problems fit the model's strengths (language, ambiguity, open-ended phrasing). Others fight its nature as a next-token predictor that is usually right and occasionally, fluently wrong (Module 1).
| Situation | Use this | Why |
|---|---|---|
| Exact arithmetic, totals, dates, counts | Code | Models make arithmetic and counting errors (Modules 1 and 2); code is exact and free |
| A fixed pattern (ids, emails, dates in one format) | Regex or parser | Deterministic, auditable, microseconds (C2 measures it) |
| Decisions that need a written, stable, explainable rule (refund eligibility, access rights) | Rules in code, reviewed by the policy owner | The same input must give the same answer, and auditors must be able to read why |
| High-volume classification with plenty of labelled history | A classical classifier, LLM only for the leftovers | Cheaper by orders of magnitude, and often as accurate once labels are plentiful |
| Facts that change daily (prices, stock, account state) | A database lookup or tool call | Model knowledge is frozen at training time |
| Questions where no answer is better than a wrong one, with no way to verify | Do not automate; route to a person | Without a verifier, you cannot tell good output from plausible output |
| Tasks with a verifier, ambiguity, or open language (drafting, summarizing, messy extraction, code with tests) | An LLM | This is where it earns its cost |
C2. When a rule, a regex, or a classical model is the correct answer
Here is the comparison you should run before any LLM feature ships: the simplest methods first, on the same data, with intervals. Three baselines, all measured:
- (a) A regex for invoice ids, scored with precision (of the ids we extracted, how many were right) and recall (of the ids that were there, how many we found).
- (b) Keyword rules for ticket category, written by reading the dev tickets only.
- (c) A TF-IDF plus logistic regression classifier: TF-IDF turns each ticket into a vector of weighted character sequences, and logistic regression learns one weight per sequence and category. Trained on dev, tested on test. Its settings were fixed before looking at the test split.
# examples/m14_classical_baselines.py
"""Judgement: when a regex, a rule, or a classical model is the right answer.
Real measurements on the Brightlane data:
(a) a regex for invoice ids: precision and recall on tickets plus hand-labelled hard cases
(b) keyword rules for ticket category: accuracy on dev and test
(c) TF-IDF + logistic regression (scikit-learn 1.9.1) trained on dev, tested on test
and the cost and latency of each next to an LLM.
"""
from __future__ import annotations
import os
os.environ.setdefault("OMP_NUM_THREADS", "1") # 2-core machine: BLAS thread oversubscription made fits 4x slower
os.environ.setdefault("OPENBLAS_NUM_THREADS", "1")
import re # noqa: E402
import time # noqa: E402
from collections import Counter # noqa: E402
from sklearn.feature_extraction.text import TfidfVectorizer # noqa: E402
from sklearn.linear_model import LogisticRegression # noqa: E402
from sklearn.model_selection import RepeatedStratifiedKFold, cross_val_score # noqa: E402
from sklearn.pipeline import make_pipeline # noqa: E402
from examples.m14_common import fmt_rate # noqa: E402
from supportdesk.data import load_tickets # noqa: E402
# (a) Invoice ids ---------------------------------------------------------------------------
NAIVE = re.compile(r"INV-\d+")
STRICT = re.compile(r"\bINV-\d{4}-\d{6}\b")
ROBUST = re.compile(r"(?<![A-Za-z0-9])INV[- ]?(\d{4})[- ]?(\d{6})(?![0-9])", re.I)
HARD_CASES = [ # (text, invoice ids a careful human would extract), labelled by hand
("Please check INV-2026-004512.", ["INV-2026-004512"]),
("invoice inv-2026-004513 is wrong", ["INV-2026-004513"]),
("Ref INV 2026 004514 from March", ["INV-2026-004514"]),
("INV-2026-004515,INV-2026-004516 both duplicated", ["INV-2026-004515", "INV-2026-004516"]),
("請求書INV-2026-004517について", ["INV-2026-004517"]),
("Rechnung (INV-2026-004518) doppelt", ["INV-2026-004518"]),
("Our VAT id is DE811234567", []),
("Order 2026-004519 shipped", []),
("SINV-2026-004520 is from our supplier", []),
("INV-2026-00452 looks short", []),
("INV-2026-0045211 has too many digits", []),
("See INV-99 in the old system", []),
]
def normalise(match: re.Match) -> str:
return f"INV-{match.group(1)}-{match.group(2)}" if match.groups() else match.group(0).upper()
def extract(pattern: re.Pattern, text: str) -> list[str]:
return [normalise(m) for m in pattern.finditer(text)]
def precision_recall(pattern: re.Pattern, cases) -> tuple[int, int, int]:
tp = fp = fn = 0
for text, gold in cases:
found = Counter(extract(pattern, text))
want = Counter(gold)
tp += sum((found & want).values())
fp += sum((found - want).values())
fn += sum((want - found).values())
return tp, fp, fn
# (b) Keyword rules (written by reading dev tickets only) ---------------------------------------
RULES = [ # every phrase below appears in at least one DEV ticket; none was taken from test
("feature_request", r"when will|do you have|roadmap|promise|own domain"),
("cancellation", r"cancel|money back|refund for the remaining|stop now|解約"),
("account_access", r"password|locked|log ?in|sign(ing)? in|2fa|saml|sso users|ownership|region|migrate|"
r"passwort|ロック|data processing|data stored"),
("bug", r"not loading|stopped posting|firing|500 errors|lost my changes|keine"),
("billing", r"charge|invoice|price|cost|pay|tax|cobr|downgrade|plan\b|कीमत"),
]
def rule_category(text: str) -> str:
low = text.lower()
for category, pattern in RULES:
if re.search(pattern, low):
return category
return "how_to" # the most common category is the fallback
# (c) TF-IDF + logistic regression ------------------------------------------------------------
def make_classifier():
"""Chosen BEFORE looking at test: character n-grams cope with typos and non-English words."""
return make_pipeline(TfidfVectorizer(analyzer="char_wb", ngram_range=(2, 5), sublinear_tf=True),
LogisticRegression(C=10, max_iter=2000))
def accuracy(pred: list[str], gold: list[str]) -> tuple[int, int]:
return sum(p == g for p, g in zip(pred, gold)), len(gold)
if __name__ == "__main__":
tickets = load_tickets()
ticket_cases = [(t.text, re.findall(r"INV-\d{4}-\d{6}", t.text)) for t in tickets] # gold: 2 ids in 72 tickets
print("(a) invoice ids: 72 tickets (2 ids) + 12 hard cases (7 ids)")
for name, pattern in (("naive", NAIVE), ("strict", STRICT), ("robust", ROBUST)):
tp, fp, fn = precision_recall(pattern, ticket_cases + HARD_CASES)
print(f" {name:6} precision {fmt_rate(tp, tp + fp)} recall {fmt_rate(tp, tp + fn)}")
dev, test = load_tickets("dev"), load_tickets("test")
y_dev, y_test = [t.gold["category"] for t in dev], [t.gold["category"] for t in test]
majority = Counter(y_dev).most_common(1)[0][0]
print(f"\n(b, c) category accuracy (dev n={len(dev)}, test n={len(test)})")
print(f" majority '{majority}' test {fmt_rate(*accuracy([majority] * len(test), y_test))}")
for split, ts, ys in (("dev", dev, y_dev), ("test", test, y_test)):
print(f" keyword rules {split:4} {fmt_rate(*accuracy([rule_category(t.text) for t in ts], ys))}")
clf = make_classifier()
cv = cross_val_score(make_classifier(), [t.text for t in dev], y_dev,
cv=RepeatedStratifiedKFold(n_splits=4, n_repeats=5, random_state=14))
started = time.perf_counter()
clf.fit([t.text for t in dev], y_dev)
fit_ms = (time.perf_counter() - started) * 1000
pred = [str(p) for p in clf.predict([t.text for t in test])]
print(f" tfidf+logreg dev 4-fold CV x5: mean {cv.mean():.1%}, fold std {cv.std():.1%}")
print(f" tfidf+logreg test {fmt_rate(*accuracy(pred, y_test))}")
wrong = [(t.id, g, p) for t, g, p in zip(test, y_test, pred) if g != p]
print(f" tfidf+logreg errors (id, gold, predicted): {wrong}")
texts = [t.text for t in tickets] * 20
for name, fn in (("regex", lambda xs: [extract(ROBUST, x) for x in xs]),
("rules", lambda xs: [rule_category(x) for x in xs]),
("tfidf+logreg", clf.predict)):
started = time.perf_counter()
fn(texts)
per = (time.perf_counter() - started) / len(texts) * 1000
print(f" latency {name:13} {per:.3f} ms/ticket, CPU for 100k tickets {per * 100_000 / 1000:.1f} s")
print(f" tfidf+logreg training time: {fit_ms:.0f} ms on {len(dev)} tickets")
# (d) The LLM on the same test split: only with a real model (M14_LIVE=1), never with the stand-in.
if os.environ.get("M14_LIVE") == "1":
from examples.m14_triage_at_scale import triage_many
from supportdesk.llm import chat
outcomes, seconds = triage_many(test, chat, workers=4)
llm_pred = [o.triage.category if o.triage else "invalid" for o in outcomes]
print(f" LLM test {fmt_rate(*accuracy(llm_pred, y_test))}, wall time {seconds:.1f} s, "
f"median latency {sorted(o.latency_ms for o in outcomes)[len(outcomes) // 2]:.0f} ms")
else:
print(" LLM skipped (set M14_LIVE=1 and a provider key to measure it on the same 24 tickets)")
Code explained
- In simple words: before hiring an expensive expert, check whether a stencil, a checklist, or a trained intern already does the job, and measure each one honestly.
- What happens:
- Three regexes of increasing care.
NAIVEmatchesINV-and any digits.STRICTwants the exact format with word boundaries.ROBUSTaccepts lower case and spaces, normalizes to the canonical form, and uses lookarounds instead of\b. HARD_CASESare 12 strings labelled by hand, holding 7 ids, including traps (a VAT id, a supplier'sSINV-number, wrong digit counts). The 2 ids in the 72 tickets were checked by hand.precision_recallcompares extracted and gold ids as multisets, so a duplicate extraction counts as a false positive.RULESare regexes tried in order, withhow_toas the fallback. Every phrase appears in at least one dev ticket; none was copied from test.make_classifieruses character n-grams of 2 to 5 inside word boundaries (char_wb), which tolerate typos and share pieces across languages. Dev accuracy is estimated with 4-fold cross-validation repeated 5 times, then the model is trained on all of dev and scored once on test.- Latency is measured over 1,440 tickets (the 72 repeated 20 times) and converted to CPU seconds per 100,000 tickets.
- The top of the file pins BLAS to one thread. On this 2-core machine the default threading made each fit about 4 times slower (0.84 s versus 0.21 s per fit measured), and the first version of this experiment timed out. If a small scikit-learn job is inexplicably slow, check thread oversubscription first.
- With
M14_LIVE=1, the last block runs the A1 triage prompt on the same 24 test tickets with a real model.
- Three regexes of increasing care.
- Comes out: accuracies are deterministic; latencies vary by machine and load.text
(a) invoice ids: 72 tickets (2 ids) + 12 hard cases (7 ids) naive precision 0/11 = 0.0% (95% CI 0.0% to 25.9%) recall 0/9 = 0.0% (95% CI 0.0% to 29.9%) strict precision 6/6 = 100.0% (95% CI 61.0% to 100.0%) recall 6/9 = 66.7% (95% CI 35.4% to 87.9%) robust precision 9/9 = 100.0% (95% CI 70.1% to 100.0%) recall 9/9 = 100.0% (95% CI 70.1% to 100.0%) (b, c) category accuracy (dev n=48, test n=24) majority 'how_to' test 6/24 = 25.0% (95% CI 12.0% to 44.9%) keyword rules dev 45/48 = 93.8% (95% CI 83.2% to 97.9%) keyword rules test 14/24 = 58.3% (95% CI 38.8% to 75.5%) tfidf+logreg dev 4-fold CV x5: mean 51.7%, fold std 12.0% tfidf+logreg test 12/24 = 50.0% (95% CI 31.4% to 68.6%) tfidf+logreg errors (id, gold, predicted): [('T-1012', 'how_to', 'billing'), ('T-1018', 'how_to', 'bug'), ('T-1021', 'bug', 'billing'), ('T-1024', 'account_access', 'how_to'), ('T-1030', 'feature_request', 'how_to'), ('T-1036', 'feature_request', 'how_to'), ('T-1039', 'bug', 'how_to'), ('T-1051', 'how_to', 'bug'), ('T-1057', 'how_to', 'account_access'), ('T-1060', 'cancellation', 'billing'), ('T-1066', 'billing', 'how_to'), ('T-1072', 'account_access', 'how_to')] latency regex 0.019 ms/ticket, CPU for 100k tickets 1.9 s latency rules 0.026 ms/ticket, CPU for 100k tickets 2.6 s latency tfidf+logreg 0.513 ms/ticket, CPU for 100k tickets 51.3 s tfidf+logreg training time: 311 ms on 48 tickets LLM skipped (set M14_LIVE=1 and a provider key to measure it on the same 24 tickets)
Read the results one baseline at a time.
The regex. The naive pattern is useless: it extracts INV-2026 fragments, so precision and recall are both zero. The strict pattern is precise but misses 3 of 9 ids: the lowercase one, the spaced one, and the id right after Japanese text. That last one is instructive: \b needs a change between a word character and a non-word character, and in Python's Unicode regexes the kanji 書 is a word character, so there is no boundary before INV. The robust pattern gets 9 of 9, with a caveat you must state: I wrote it after writing the hard cases, so 100% is optimistic, and the intervals (70% to 100%) reflect only 9 ids. The honest next step is a fresh set of real invoice mentions from production logs. Even so, the conclusion is not in doubt: extraction of a fixed-format id is a regex job at 0.02 ms per ticket.
The rules. Keyword rules score 93.8% on dev and 58.3% on test. That gap is the most important number in this section. Rules written by reading the dev tickets overfit to them exactly the way a model overfits its training data (Module 1's TinyLM did the same thing with its corpus). If you had only measured on the tickets you read while writing the rules, you would have shipped a 58% classifier believing it was 94%.
The classifier. TF-IDF plus logistic regression gets 50% on test, with an interval of 31% to 69%. Rules (14 of 24) and classifier (12 of 24) are within noise of each other; both are clearly better than always guessing how_to (25%). With 48 training tickets the classifier has seen only 4 examples of cancellation and feature_request, and its errors show it. This is a data problem, not a method problem: Brightlane's agents have labelled thousands of historical tickets, and a classifier trained on those is the baseline to beat.
Where the LLM stands. No model was called for this module, so there is no measured LLM accuracy to put here. The harness is ready: M14_LIVE=1 LLM_PROVIDER=groq GROQ_API_KEY=... PYTHONPATH=. python examples/m14_classical_baselines.py adds a line with accuracy, wall time, and median latency on the same 24 tickets. As an illustration only (not measured in this build; your numbers will differ), a capable instruction-tuned model given the category definitions often lands somewhere around 80% to 90% on a set like this, and with 24 tickets its interval will still be about 25 points wide. The cost side is measured: from A1, the LLM costs about 1 to 195 USD per 100,000 tickets depending on the model, plus network latency per call, versus about a minute of CPU time for the classifier and a few seconds for the rules.
| Situation | Use this | Why |
|---|---|---|
| Fixed-format fields (ids, VAT numbers, dates in one format) | Regex, tested on hard cases | Measured 100% on the cases here at 0.02 ms, and fully explainable |
| Categories with thousands of labelled historical tickets | TF-IDF plus logistic regression, or a small fine-tuned model (Module 9) | Cheap, fast, and strong once labels are plentiful |
| Few labels, many categories, multilingual | LLM with a schema, then collect labels from its reviewed output | Gets you started; the reviewed labels train the cheaper model later |
| Most tickets are easy, some are not | Cascade: rules or classifier when confident, LLM otherwise | The Module Lab shows the rules are right 9 of 11 times when they fire |
C3. Estimating whether a use case clears the reliability bar
"Is 90% accurate good enough?" has no answer until you price the errors. The reliability bar is the accuracy at which automating a task costs less, in expectation, than the next best way of doing it. Expected cost is each outcome's cost times its probability, summed.
For a Brightlane how-to ticket there are three ways to handle it: an agent writes the answer, the model drafts and an agent reviews, or the model sends the answer itself. The script writes every business input as a named assumption so Maya can change the ones she disagrees with.
# examples/m14_reliability_bar.py
"""Judgement: does a use case clear the reliability bar? An expected-cost calculation.
Compares three ways to handle a Brightlane how-to ticket (manual, AI draft with
human review, AI auto-send) as a function of the model's accuracy. All the
business inputs are ASSUMPTIONS written down so Maya can argue with them.
"""
from __future__ import annotations
from examples.m14_common import wilson
from supportdesk.llm import Usage
from supportdesk.pricing import cost_usd
AGENT_PER_MIN = 0.75 # ASSUMED loaded cost of an agent minute (45 USD/hour)
MANUAL_MIN = 6.0 # ASSUMED minutes to answer a how-to ticket by hand
REVIEW_MIN = 1.5 # ASSUMED minutes to review a correct draft
FIX_MIN = 5.0 # ASSUMED minutes to fix a wrong draft that the reviewer catches
REVIEW_CATCH = 0.85 # ASSUMED share of wrong drafts the reviewer catches
ESCAPE_COST = 12.0 # ASSUMED cost of a wrong answer reaching a customer (follow-up ticket + churn risk)
LLM_COST = cost_usd(Usage(input_tokens=1200, output_tokens=350), "openai/gpt-oss-120b") # retrieval prompt + reply
def cost_manual(_: float) -> float:
return MANUAL_MIN * AGENT_PER_MIN
def cost_draft_review(acc: float, escape_cost: float = ESCAPE_COST) -> float:
wrong = 1 - acc
return (LLM_COST + REVIEW_MIN * AGENT_PER_MIN
+ wrong * REVIEW_CATCH * FIX_MIN * AGENT_PER_MIN
+ wrong * (1 - REVIEW_CATCH) * escape_cost)
def cost_auto(acc: float, escape_cost: float = ESCAPE_COST) -> float:
return LLM_COST + (1 - acc) * escape_cost
def break_even(f, g, lo: float = 0.0, hi: float = 1.0) -> float | None:
"""Accuracy at which f and g cost the same (bisection); None if one always wins."""
if (f(lo) - g(lo)) * (f(hi) - g(hi)) > 0:
return None
for _ in range(60):
mid = (lo + hi) / 2
if (f(lo) - g(lo)) * (f(mid) - g(mid)) <= 0:
hi = mid
else:
lo = mid
return (lo + hi) / 2
if __name__ == "__main__":
print(f"LLM cost per ticket: {LLM_COST:.5f} USD (it rounds to zero next to agent time)")
print(f"{'accuracy':>8} {'manual':>8} {'draft+review':>13} {'auto-send':>10} cheapest")
for acc in (0.70, 0.80, 0.90, 0.95, 0.98, 0.99):
costs = {"manual": cost_manual(acc), "draft+review": cost_draft_review(acc), "auto-send": cost_auto(acc)}
print(f"{acc:>8.0%} {costs['manual']:>8.2f} {costs['draft+review']:>13.2f} {costs['auto-send']:>10.2f} "
f"{min(costs, key=costs.get)}")
for escape in (12.0, 60.0, 300.0):
be = break_even(lambda a: cost_auto(a, escape), lambda a: cost_draft_review(a, escape))
be_manual = break_even(lambda a: cost_draft_review(a, escape), cost_manual)
print(f"escape cost {escape:>5.0f} USD: auto-send beats draft+review above {be:.1%} accuracy; "
f"draft+review beats manual above {be_manual:.1%}" if be_manual else "draft+review always beats manual")
for right, n in ((23, 24), (230, 240)):
lo, _ = wilson(right, n)
print(f"measured {right}/{n} correct: plan with the lower 95% bound, {lo:.1%}, not {right / n:.1%}")
Code explained
- In simple words: a spreadsheet with the assumptions at the top, answering "at what accuracy does each option become the cheapest?".
- What happens:
- The constants are assumptions: agent time at 0.75 USD a minute, 6 minutes to answer by hand, 1.5 minutes to review a good draft, 5 minutes to fix a caught bad one, an 85% reviewer catch rate (Part B), and 12 USD for a wrong answer that reaches a customer.
- The LLM cost is real arithmetic from
cost_usdfor a 1,200-token grounded prompt and a 350-token reply on gpt-oss-120b. cost_draft_reviewadds the review time, the fix time for caught errors, and the escape cost for missed errors.cost_autopays only the model and the escape cost.break_evenfinds, by bisection, the accuracy where two options cost the same.- The escape cost is varied from 12 USD (a wrong how-to answer) to 60 and 300 USD (a wrong answer about refunds or data deletion).
- The last lines turn a measured accuracy into the number you should plan with: the lower end of its interval.
- Comes out:text
LLM cost per ticket: 0.00039 USD (it rounds to zero next to agent time) accuracy manual draft+review auto-send cheapest 70% 4.50 2.62 3.60 draft+review 80% 4.50 2.12 2.40 draft+review 90% 4.50 1.62 1.20 auto-send 95% 4.50 1.37 0.60 auto-send 98% 4.50 1.23 0.24 auto-send 99% 4.50 1.18 0.12 auto-send escape cost 12 USD: auto-send beats draft+review above 84.0% accuracy; draft+review beats manual above 32.3% escape cost 60 USD: auto-send beats draft+review above 97.6% accuracy; draft+review beats manual above 72.3% escape cost 300 USD: auto-send beats draft+review above 99.6% accuracy; draft+review beats manual above 93.0% measured 23/24 correct: plan with the lower 95% bound, 79.8%, not 95.8% measured 230/240 correct: plan with the lower 95% bound, 92.5%, not 95.8%Four conclusions, each only as good as the assumptions. The model's cost (0.0004 USD) is irrelevant next to agent minutes; the economics are entirely about time and errors. Draft plus review beats manual work at almost any accuracy for cheap errors (above 32%), because reviewing is faster than writing. Auto-send wins for how-to questions above 84% accuracy. For expensive errors the bar climbs to 97.6% and then 99.6%, which is why refunds stayed human-only in B1. And the last two lines are the trap: 23 of 24 correct looks like 95.8%, but the interval's lower end is 79.8%, below the 84% bar. You would need about ten times more test data before auto-send is justified, even with a model that really is 96% accurate.
C4. Communicating limitations to stakeholders
Stakeholders make better decisions from a short, honest page than from a demo. A good limitations brief has five parts: what the system does, how well it works with uncertainty attached, known limits each paired with a mitigation, what would make you stop, and the decisions you need from the reader. Generate the numbers from your evals so the page cannot drift from reality.
# examples/m14_limitations_onepager.py
"""Judgement: communicating limitations to stakeholders with a one-page template.
The page is generated from measured numbers so it cannot drift from the evals.
Every number carries its sample size, and every limit says what we do about it.
"""
from __future__ import annotations
from pathlib import Path
from examples.m14_common import fmt_rate
from examples.m14_grounded_answer import KB, choose_threshold, safe_to_answer
from supportdesk.data import load_tickets
TEMPLATE = """# Help-center draft replies: what it does and where it fails
Owner: {owner} | Audience: support leadership | Data as of: {as_of} | Prompt: {prompt_version}
## What it does
Suggests a reply from the help center for agents to review. It never sends anything on its own.
## How well it works (measured, with uncertainty)
- Answers instead of abstaining on {answered} of held-out tickets.
- Of the tickets it answers, the source article is wrong for {unsafe}.
- Non-English tickets: {non_en} held-out tickets answered. Treat non-English support as unsupported for now.
## Known limits and what we do about each
| Limit | Example | Mitigation |
|---|---|---|
| Can state a policy that is not in the source | invented refund window | citation check; agent review; [Wrong source] button |
| Weak on non-English tickets | Spanish refund request missed | abstains and routes to a human |
| Help center gaps | ERP integration questions | abstains; gaps feed the article-draft queue |
## What would make us stop
Wrong-source rate above 15% on the weekly audit sample, or any customer-visible policy error.
## Decisions we need from you
1. Approve review time budget of 1.5 minutes per draft. 2. Name an owner for help-center gaps.
"""
if __name__ == "__main__":
dev, test = load_tickets("dev"), load_tickets("test")
threshold = choose_threshold(dev)
answered = [t for t in test if KB.search(t.text, 1) and KB.search(t.text, 1)[0].score >= threshold]
unsafe = sum(not safe_to_answer(t, KB.search(t.text, 1)) for t in answered)
non_en = [t for t in test if t.language != "en"]
page = TEMPLATE.format(
owner="Support AI team", as_of="2026-09-21", prompt_version="draft-v2",
answered=fmt_rate(len(answered), len(test)), unsafe=fmt_rate(unsafe, len(answered)),
non_en=f"{sum(t in answered for t in non_en)} of {len(non_en)}")
out = Path("m14_out/limitations_onepager.md")
out.parent.mkdir(exist_ok=True)
out.write_text(page)
print(page)
Code explained
- In simple words: a fill-in-the-blanks memo where the blanks are filled by the test harness, not by optimism.
- What happens:
TEMPLATEis the one-page structure. The limits table pairs every weakness with what the system does about it; a limit without a mitigation is a reason not to ship.- The main block reuses the threshold and the safety definition from A3, measures the test split, and fills the template with rates and intervals.
- The page is written to
m14_out/limitations_onepager.md, ready to paste into a document or an email.
- Comes out:text
# Help-center draft replies: what it does and where it fails Owner: Support AI team | Audience: support leadership | Data as of: 2026-09-21 | Prompt: draft-v2 ## What it does Suggests a reply from the help center for agents to review. It never sends anything on its own. ## How well it works (measured, with uncertainty) - Answers instead of abstaining on 15/24 = 62.5% (95% CI 42.7% to 78.8%) of held-out tickets. - Of the tickets it answers, the source article is wrong for 1/15 = 6.7% (95% CI 1.2% to 29.8%). - Non-English tickets: 1 of 2 held-out tickets answered. Treat non-English support as unsupported for now. ## Known limits and what we do about each | Limit | Example | Mitigation | |---|---|---| | Can state a policy that is not in the source | invented refund window | citation check; agent review; [Wrong source] button | | Weak on non-English tickets | Spanish refund request missed | abstains and routes to a human | | Help center gaps | ERP integration questions | abstains; gaps feed the article-draft queue | ## What would make us stop Wrong-source rate above 15% on the weekly audit sample, or any customer-visible policy error. ## Decisions we need from you 1. Approve review time budget of 1.5 minutes per draft. 2. Name an owner for help-center gaps.Every number carries its sample size and interval, including the uncomfortable one: 1 of 2 non-English test tickets answered, which is too few to claim anything, so the page says to treat non-English support as unsupported. "What would make us stop" is the section stakeholders value most, because it shows you have thought about failure before it happens.