Part A: Proven Application Patterns
Dark
Module 14: Applications, Patterns, and Frontiers
By the end of this module, you'll have:
- Seven proven LLM application patterns built as small, runnable Brightlane examples (triage at scale, a weekly digest, grounded answers with abstention, text-to-SQL with tests as the verifier, article drafts behind a review queue, an enrichment pipeline, and permission-aware search over past tickets), each with the real non-LLM parts measured.
- An autonomy policy that maps every Brightlane action to an autonomy level from its consequences, plus undo, fallbacks, a reviewer-fatigue model, an agent-facing draft card with measured confidence bands, and a feedback event schema.
- A real comparison of a regex, keyword rules, and a scikit-learn classifier on the ticket data, with confidence intervals, latency, and cost next to an LLM, and an expected-cost calculation that tells you when a use case clears the reliability bar.
- A one-page limitations brief for stakeholders, generated from measured numbers.
- A working method for reading the frontier: price and context trends, test-time compute, agent time horizons, small on-device models (with TinyLM's real CPU latency), and a checklist applied to a real September 2026 model announcement.
- A model-upgrade gate that keeps the assistant adaptable as models change under it.
Prerequisites: Modules 1 to 13, especially structured outputs (Module 6), retrieval with KBSearch (Module 7), agents and approval gates (Module 8), evaluation (Module 10), safety (Module 11), and operations (Module 13). Working Python. No new math beyond percentages and expected value.
Where we are: Module 13 put the Brightlane assistant into production with fallbacks, budgets, and dashboards. This module steps back and asks the questions that decide whether an LLM feature is worth building at all: which patterns work, where humans belong, when a regex beats a model, and how to keep the system sound as the models underneath it keep changing.
How this module is organized
| Part | What it covers |
|---|---|
| Part A: Proven Application Patterns | Seven patterns, each as a compact Brightlane example: classification and extraction at scale, summarization, grounded conversation, code generation with a verifier, content generation with review, enrichment pipelines, search over private data |
| Part B: Design Principles | Autonomy matched to consequence, designing for the model being wrong, review placement and reviewer fatigue, progressive disclosure, trust calibration, feedback capture |
| Part C: Judgement | Problems LLMs should not solve, a real regex vs rules vs classifier vs LLM comparison, the reliability bar as expected cost, a limitations brief for stakeholders |
| Part D: Frontier | Longer contexts and cheaper inference, test-time compute, agent time horizons, small and on-device models, reading releases critically, staying adaptable |
| Module Lab | One intake pipeline that runs the cheapest reliable step first and meters every model call |
| Project Milestone, Interview Questions, Other Tools, Coming Up | What the repository now contains, practice questions, alternatives, and the capstone preview |
Part A: Proven Application Patterns
Almost every LLM feature that survives contact with production is one of a handful of shapes. The shapes matter more than the model: each comes with a known way to check the output, a known way to fail, and a known place for a human. In this part you build all seven for Brightlane, small enough to read in one sitting.
A word about what is real here. There is no API key in the build environment for this course, so every example runs by default against ScriptedLLM, the stand-in from Module 1. Its replies are written by hand to exercise the plumbing: the parsers, gates, queues, caches, and cost math. ScriptedLLM output is not model output, and no accuracy number in this part comes from it. Everything else (token counts, retrieval scores, SQL results, queue states, timings, prices) is measured for real. Set M14_LIVE=1 with a provider key and the same scripts send the same requests to a real model through supportdesk.llm.chat.
Setup for this module
The examples run in order from the root of your supportdesk copy. Part C needs scikit-learn, which is new in this module.
cp -r supportdesk work/m14 && cd work/m14
source /home/claude/venv/bin/activate # or your own Python 3.11 venv
pip install scikit-learn==1.9.1
PYTHONPATH=. python -c "import sklearn; print(sklearn.__version__)"
Code explained
- In simple words: make a private copy of the project and add the one new library this module uses.
- What happens: the copy keeps your experiments away from the canonical code;
pip installpins scikit-learn 1.9.1 (it also pulls in scipy 1.17.1, joblib 1.6.0, and threadpoolctl 3.7.0); the last line proves the import works with the repository on the Python path. - Comes out:text
1.9.1
Every example imports a small shared file. It picks the stand-in or the live model, and it prints rates with a Wilson confidence interval, a range that is honest about small samples (24 test tickets is small, and you will see intervals 40 points wide).
# examples/m14_common.py
"""Shared setup for the Module 14 examples (run every example from the repo root with PYTHONPATH=.).
Every example works without an API key: it uses a ScriptedLLM stand-in by default.
Set M14_LIVE=1 (plus LLM_PROVIDER and the provider's key) to send the same
requests to a real model through supportdesk.llm.chat instead.
"""
from __future__ import annotations
import math
import os
from collections.abc import Callable
from supportdesk.llm import chat, resolve
from supportdesk.stand_in import ScriptedLLM
ChatFn = Callable[..., object]
def pick_llm(stand_in: ScriptedLLM) -> ChatFn:
"""Return the real `chat` when M14_LIVE=1, otherwise the scripted stand-in."""
if os.environ.get("M14_LIVE") == "1":
provider, model = resolve()
print(f"[live] using {provider}:{model}")
return chat
print("[stand-in] ScriptedLLM: tests the plumbing only, not model quality")
return stand_in
def wilson(k: int, n: int, z: float = 1.96) -> tuple[float, float]:
"""95% Wilson confidence interval for a proportion k/n (honest with small n)."""
if n == 0:
return (0.0, 1.0)
p = k / n
centre = (p + z * z / (2 * n)) / (1 + z * z / n)
half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
return (max(0.0, centre - half), min(1.0, centre + half))
def fmt_rate(k: int, n: int) -> str:
"""'17/24 = 70.8% (95% CI 50.8% to 85.1%)'."""
lo, hi = wilson(k, n)
return f"{k}/{n} = {k / n:.1%} (95% CI {lo:.1%} to {hi:.1%})"
Code explained
- In simple words: a switch between "practice mode" and "live mode", plus a ruler that always prints the error bars.
- What happens:
pick_llmreturnssupportdesk.llm.chatonly whenM14_LIVE=1, and prints which one you got so a log can never be mistaken for a real run.wilsoncomputes the 95% Wilson interval for k successes out of n; unlike the textbook "p plus or minus 1.96 standard errors", it never leaves the 0 to 1 range and behaves well at 0/n and n/n.fmt_rateformats both. - Comes out: nothing on its own. Called as
fmt_rate(17, 24)it returns17/24 = 70.8% (95% CI 50.8% to 85.1%), which already tells you that a 24-ticket test cannot tell 70% from 80%.
A1. Classification and extraction at scale
Classification puts each input into one of a fixed set of labels; extraction pulls typed fields out of free text. Together they are the most common production use of LLMs, because the output is small, checkable, and easy to aggregate. For Brightlane the job is triage: 100,000 tickets a month, each labelled with a category, a priority, a language, and a one-line summary using the Triage schema from Module 6.
At that volume three engineering questions dominate, and none of them is about prompt wording:
- Concurrency: calls spend most of their time waiting on the network, so running several at once with a thread pool shortens wall time almost linearly until you hit the provider's rate limit (Module 13).
- Failure handling: some replies will not parse. A dead-letter list collects them for retry or human handling instead of crashing the batch or silently dropping tickets.
- Cost: you multiply a per-call cost by 100,000, so you measure the prompt in tokens and compare models, the provider's batch API (asynchronous, usually 50% off, results within hours), prompt caching of the stable system prefix, and packing several tickets into one request.
# examples/m14_triage_at_scale.py
"""Pattern 1: classification and extraction at scale (Brightlane ticket triage).
Runs 72 tickets through a triage prompt with a thread pool, validates every reply
against the Triage schema, sends failures to a dead-letter list, and projects the
monthly bill for 100,000 tickets with supportdesk.pricing.
"""
from __future__ import annotations
import json
import time
from concurrent.futures import ThreadPoolExecutor
from dataclasses import dataclass
from pydantic import ValidationError
from examples.m14_common import pick_llm
from supportdesk.data import CATEGORIES, Ticket, load_tickets
from supportdesk.llm import ChatResult, Usage
from supportdesk.pricing import cost_usd
from supportdesk.schemas import Triage
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages, count_tokens
SYSTEM = (
"You triage support tickets for Brightlane, a project-management SaaS.\n"
f"Categories: {', '.join(CATEGORIES)}.\n"
"Priorities: urgent (many users blocked or security risk now), high (one user blocked or money wrong), "
"normal (needs an answer), low (question or idea).\n"
"Reply with only a JSON object with keys category, priority, language (ISO 639-1), "
"summary (one English sentence), needs_human (true or false)."
)
def build_messages(ticket: Ticket) -> list[dict]:
return [{"role": "system", "content": SYSTEM},
{"role": "user", "content": f"<ticket>\n{ticket.text}\n</ticket>"}]
@dataclass
class Outcome:
ticket_id: str
triage: Triage | None
error: str
usage: Usage
latency_ms: float
def triage_one(ticket: Ticket, llm) -> Outcome:
started = time.perf_counter()
result: ChatResult = llm(build_messages(ticket), temperature=0.0, max_tokens=300)
ms = (time.perf_counter() - started) * 1000
try:
return Outcome(ticket.id, Triage.model_validate_json(result.text), "", result.usage, ms)
except ValidationError as exc:
return Outcome(ticket.id, None, exc.errors()[0]["type"], result.usage, ms)
def triage_many(tickets: list[Ticket], llm, workers: int) -> tuple[list[Outcome], float]:
started = time.perf_counter()
with ThreadPoolExecutor(max_workers=workers) as pool:
outcomes = list(pool.map(lambda t: triage_one(t, llm), tickets))
return outcomes, time.perf_counter() - started
def scripted_responder(messages, kwargs):
"""STAND-IN: echoes gold labels after a fake 50 ms 'network' delay; every 12th reply is broken."""
time.sleep(0.05)
text = messages[-1]["content"]
ticket = next(t for t in TICKETS if t.text in text)
if int(ticket.id[-2:]) % 12 == 0:
return '{"category": "billing", "priority": "normal"' # truncated JSON, like a max_tokens cut
return json.dumps({"category": ticket.gold["category"], "priority": ticket.gold["priority"],
"language": ticket.language, "summary": f"Customer asks about: {ticket.subject}.",
"needs_human": not ticket.gold["answerable"]})
TICKETS = load_tickets()
if __name__ == "__main__":
llm = pick_llm(ScriptedLLM(responder=scripted_responder))
count_tokens("warm up the tokenizer so the first timing is fair")
for workers in (1, 8):
outcomes, seconds = triage_many(TICKETS, llm, workers)
dead = [o for o in outcomes if o.triage is None]
print(f"workers={workers}: {len(TICKETS)} tickets in {seconds:.2f} s, "
f"valid {len(outcomes) - len(dead)}, dead-letter {len(dead)} {[(o.ticket_id, o.error) for o in dead]}")
# Cost projection from measured prompt sizes (o200k_base token counts, an estimate for other tokenizers).
input_tokens = [count_messages(build_messages(t)) for t in TICKETS]
system_tokens = count_tokens(SYSTEM) + 4
example_reply = '{"category": "billing", "priority": "high", "language": "en", "summary": "Customer was charged twice for the Team plan and wants the duplicate refunded.", "needs_human": true}'
out_tokens = count_tokens(example_reply)
avg_in = sum(input_tokens) / len(input_tokens)
print(f"\nmeasured: avg input {avg_in:.0f} tokens (system prefix {system_tokens}), reply {out_tokens} tokens")
monthly = 100_000
reasoning = {"openai/gpt-oss-120b": 150, "openai/gpt-oss-20b": 150, "gemini-3.5-flash": 150,
"gemini-3.5-flash-lite": 0, "llama-3.1-8b-instant": 0} # ASSUMED hidden reasoning tokens per call
print(f"{'model':24} {'per call':>10} {'100k/month':>11} {'batch API':>10} {'cached prefix':>13}")
for model, extra in reasoning.items():
usage = Usage(input_tokens=round(avg_in), output_tokens=out_tokens + extra)
cached = Usage(input_tokens=round(avg_in), cached_tokens=system_tokens, output_tokens=out_tokens + extra)
one = cost_usd(usage, model)
print(f"{model:24} {one:>10.6f} {one * monthly:>11.2f} {cost_usd(usage, model, batch=True) * monthly:>10.2f} "
f"{cost_usd(cached, model) * monthly:>13.2f}")
# Packing 10 tickets into one request amortizes the system prompt.
packed_in = system_tokens + 3 + sum(input_tokens[:10]) - 10 * (system_tokens + 3)
print(f"\npacking 10 tickets per request: {sum(input_tokens[:10])} input tokens as 10 calls "
f"vs about {packed_in} as one call ({1 - packed_in / sum(input_tokens[:10]):.0%} fewer)")
Code explained
- In simple words: a sorting line that runs eight tickets at a time, puts damaged parcels in a separate bin, and then prices a month of work.
- What happens:
build_messagesputs the stable instructions in the system message and the ticket inside<ticket>tags (trusted instruction separated from untrusted content, Module 4).triage_onecalls whatever chat function it was given and validates the reply withTriage.model_validate_json. A validation error becomes anOutcomewithtriage=Noneand the pydantic error type, which is the dead-letter record.triage_manymaps tickets over aThreadPoolExecutor.pool.mapkeeps results in input order, so outcomes line up with tickets.- The stand-in responder sleeps 50 ms to imitate network wait and echoes the gold labels; every ticket whose number is divisible by 12 gets a truncated JSON reply, so the dead-letter path runs. Echoing gold labels means any "accuracy" here would be meaningless, which is why the script does not print one.
- The cost section measures the real prompt size of all 72 tickets with
count_messages(the o200k_base tokenizer, an estimate for other model families), measures a realistic reply, and adds an assumed 150 hidden reasoning tokens for the reasoning models (Module 3). It prices each model withcost_usd, withbatch=True, and with the system prefix billed at the cached rate. - The last lines compute what packing 10 tickets into one request saves: the 104-token system prompt is paid once instead of ten times.
- Comes out: timings vary run to run and machine to machine; the rest is deterministic.text
[stand-in] ScriptedLLM: tests the plumbing only, not model quality workers=1: 72 tickets in 3.83 s, valid 66, dead-letter 6 [('T-1012', 'json_invalid'), ('T-1024', 'json_invalid'), ('T-1036', 'json_invalid'), ('T-1048', 'json_invalid'), ('T-1060', 'json_invalid'), ('T-1072', 'json_invalid')] workers=8: 72 tickets in 0.48 s, valid 66, dead-letter 6 [('T-1012', 'json_invalid'), ('T-1024', 'json_invalid'), ('T-1036', 'json_invalid'), ('T-1048', 'json_invalid'), ('T-1060', 'json_invalid'), ('T-1072', 'json_invalid')] measured: avg input 141 tokens (system prefix 104), reply 43 tokens model per call 100k/month batch API cached prefix openai/gpt-oss-120b 0.000137 13.70 6.85 13.70 openai/gpt-oss-20b 0.000068 6.85 3.42 6.85 gemini-3.5-flash 0.001948 194.85 97.42 180.81 gemini-3.5-flash-lite 0.000150 14.98 7.49 12.17 llama-3.1-8b-instant 0.000010 1.05 0.52 1.05 packing 10 tickets per request: 1456 input tokens as 10 calls vs about 493 as one call (66% fewer)Read it in three blocks. First, eight workers cut wall time from about 3.8 s to about 0.5 s, roughly the 8x you would expect when the work is waiting, and both runs dead-letter the same six tickets, so concurrency did not change results. Second, the average prompt is 141 tokens, 104 of which are the fixed system prefix. Third, the monthly bill for 100,000 tickets ranges from about 1 USD (llama-3.1-8b-instant) to about 195 USD (gemini-3.5-flash, where the assumed reasoning tokens at 9 USD per million output tokens dominate). The batch API halves every line. Caching the prefix only helps where the cached rate is lower than the normal rate; in
pricing.pythe Groq models bill cached input at the full rate, so that column does not move for them. Prices inpricing.pywere checked on 21 September 2026 and change often; rerun this with current prices before you quote a number.
Packing looks like a free 66% saving on input tokens, but it has costs that do not show up in the token count:
| Situation | Use this | Why |
|---|---|---|
| Interactive triage as tickets arrive | One ticket per call, thread pool, prefix caching | Latency matters; one bad reply affects one ticket |
| Nightly re-labelling of a backlog | Provider batch API | Half price, and nobody waits for the result |
| Very short inputs with a long fixed prompt | Packing 5 to 20 items per call, then validate each item | The prefix dominates cost; check that items do not bleed into each other's labels |
| Any volume, but a rule or classifier already works | No LLM on that slice (Part C) | Microseconds and zero marginal cost |
A frequent failure at this scale is silent loss: an exception inside a worker thread that nobody logs, so 72 tickets go in and 66 labels come out and the gap is only noticed at month end. The dead-letter list turns that into a line you can count and alert on. Retry dead letters once with a repair prompt (Module 6), then hand them to a person.
A2. Summarization and document processing
Summarization turns many documents into fewer words. The failure mode is specific and dangerous: the summary reads fluently and contains a number or claim that is not in the sources. The pattern that works is to let code compute every fact that must be exact, and let the model write only the prose around those facts. For long inputs, use map-reduce (Module 5): summarize each group separately (map), then combine the partial summaries (reduce).
Maya wants a weekly digest of the support queue. Here all 72 tickets play the role of one week.
# examples/m14_weekly_digest.py
"""Pattern 2: summarization and document processing (Maya's weekly ticket digest).
Code computes every number; the model only writes prose around them (map per
category, then reduce). A checker then refuses any digest whose numbers are not
in the computed facts.
"""
from __future__ import annotations
import json
import re
from collections import Counter
from examples.m14_common import pick_llm
from supportdesk.data import Ticket, load_tickets
from supportdesk.stand_in import ScriptedLLM
def weekly_facts(tickets: list[Ticket]) -> dict:
"""Deterministic statistics: the part of a digest that must never be wrong."""
return {
"total": len(tickets),
"by_category": dict(Counter(t.gold["category"] for t in tickets).most_common()),
"urgent_or_high": sum(t.gold["priority"] in ("urgent", "high") for t in tickets),
"non_english": sum(t.language != "en" for t in tickets),
"not_answerable_from_kb": sum(not t.gold["answerable"] for t in tickets),
}
def map_prompt(category: str, tickets: list[Ticket]) -> list[dict]:
lines = "\n".join(f"- [{t.id}] {t.subject}: {t.body[:160]}" for t in tickets)
return [{"role": "system", "content": "Summarize the common themes in these support tickets in at most 2 "
"sentences. Mention ticket ids for examples. Do not state counts."},
{"role": "user", "content": f"Category: {category}\n{lines}"}]
def reduce_prompt(facts: dict, themes: dict[str, str]) -> list[dict]:
return [{"role": "system", "content": "Write a weekly support digest for the support lead in under 150 words. "
"Use ONLY numbers that appear in FACTS. Do not compute new numbers."},
{"role": "user", "content": f"FACTS: {json.dumps(facts)}\nTHEMES: {json.dumps(themes)}"}]
def allowed_numbers(facts: dict) -> set[str]:
return set(re.findall(r"\d+", json.dumps(facts)))
def check_numbers(digest: str, facts: dict) -> list[str]:
"""Every number in the digest must appear in the facts (ticket ids like T-1040 are allowed)."""
text = re.sub(r"T-\d{4}", "", digest)
return [n for n in re.findall(r"\d+", text) if n not in allowed_numbers(facts)]
# STAND-IN replies (written by hand to exercise the pipeline; not model output).
MAP_REPLY = "Themes vary; see for example [{first}] and [{second}]."
REDUCE_REPLY = (
"This week we received 72 tickets. How-to questions led (18), followed by account access (17) and "
"billing (16). 17 tickets were urgent or high priority, and 8 arrived in a language other than English. "
"Refund and duplicate-charge requests (T-1001, T-1031) remain the most time-sensitive billing theme. "
"11 tickets could not be answered from the help center, which points to article gaps."
)
def responder(messages, kwargs):
if "Category:" in messages[-1]["content"]:
ids = re.findall(r"\[(T-\d{4})\]", messages[-1]["content"])
return MAP_REPLY.format(first=ids[0], second=ids[-1])
return REDUCE_REPLY
if __name__ == "__main__":
tickets = load_tickets()
facts = weekly_facts(tickets)
print("FACTS", json.dumps(facts))
llm = pick_llm(ScriptedLLM(responder=responder))
themes = {}
for category in facts["by_category"]:
group = [t for t in tickets if t.gold["category"] == category]
themes[category] = llm(map_prompt(category, group), temperature=0.2).text
digest = llm(reduce_prompt(facts, themes), temperature=0.2).text
print(f"\nmap calls: {len(themes)}, reduce calls: 1")
print("\nDIGEST\n" + digest)
bad = check_numbers(digest, facts)
print("\nnumber check:", "PASS" if not bad else f"FAIL, numbers not in facts: {bad}")
Code explained
- In simple words: the accountant counts, the writer writes, and an auditor checks that the writer copied the numbers correctly.
- What happens:
weekly_factscomputes counts withCounterfrom the gold labels: the part of the digest that must never be wrong.map_promptasks for themes per category and explicitly forbids counts;reduce_promptpasses the facts as JSON and says to use only those numbers.check_numbersremoves ticket ids, pulls every number out of the digest, and reports any number that does not appear anywhere in the facts.- The stand-in's reduce reply is hand-written with one deliberate mistake: it says 11 tickets could not be answered from the help center.
- Comes out:text
FACTS {"total": 72, "by_category": {"how_to": 18, "account_access": 17, "billing": 16, "bug": 9, "cancellation": 6, "feature_request": 6}, "urgent_or_high": 17, "non_english": 8, "not_answerable_from_kb": 10} [stand-in] ScriptedLLM: tests the plumbing only, not model quality map calls: 6, reduce calls: 1 DIGEST This week we received 72 tickets. How-to questions led (18), followed by account access (17) and billing (16). 17 tickets were urgent or high priority, and 8 arrived in a language other than English. Refund and duplicate-charge requests (T-1001, T-1031) remain the most time-sensitive billing theme. 11 tickets could not be answered from the help center, which points to article gaps. number check: FAIL, numbers not in facts: ['11']The facts say 10 tickets were not answerable; the digest says 11, and the checker catches it. This is exactly the kind of error real models make when asked to restate numbers from a JSON blob. The checker is deliberately simple and has a known hole: a wrong number that happens to equal some other fact (writing "17 billing tickets" when 17 is the urgent count) passes. Closing that hole means asking the model for structured output (
{"claim": ..., "fact_key": ...}) and rendering the prose from it, which is worth doing when the digest drives decisions.
| Situation | Use this | Why |
|---|---|---|
| Counts, totals, trends | Code, never the model | Exact, free, and testable |
| Themes and examples across many tickets | Map-reduce summarization | Each call sees a manageable amount; ids keep claims traceable |
| One long document that fits the context window | A single call with the whole document (Module 2) | Simpler than chunking, if you have checked the model's effective context |
| A digest that feeds a decision | Structured claims plus a checker | Every sentence maps to a verifiable fact |
A3. Conversational assistants with grounded knowledge
A grounded assistant answers only from sources you supply, cites them, and says so when the sources do not cover the question. That last behavior is called abstaining, and it is the difference between an assistant Maya can deploy and one she has to apologize for. Module 7 built retrieval with KBSearch; here we add two gates around the model:
- A retrieval gate: if the best BM25 score is weak, do not even call the model. Route to a human.
- A citation gate: the model's
cited_articlesmust be a subset of what we actually retrieved and showed it. A citation to anything else is a fabrication, whatever the reply says.
The retrieval threshold is a number, so we choose it on dev data and measure it on test data, never by eye.
# examples/m14_grounded_answer.py
"""Pattern 3: a conversational assistant grounded in the help center, with citations and abstention.
Two gates protect the customer: a retrieval gate (abstain when the best BM25
score is weak, threshold chosen on dev and measured on test) and a citation
gate (every cited article must be one we actually retrieved and showed).
"""
from __future__ import annotations
from pydantic import ValidationError
from examples.m14_common import fmt_rate, pick_llm
from supportdesk.data import Ticket, get_article, load_tickets
from supportdesk.kb_search import Hit, KBSearch
from supportdesk.schemas import DraftReply
from supportdesk.stand_in import ScriptedLLM
KB = KBSearch()
ABSTAIN = DraftReply(reply="I could not find this in our help center, so I have passed your question to a "
"teammate who will reply personally.", cited_articles=[], confidence="low")
def safe_to_answer(ticket: Ticket, hits: list[Hit]) -> bool:
"""Gold definition of 'answering is safe': answerable AND our top hit is the right article."""
return bool(hits) and ticket.gold["answerable"] and hits[0].article_id == ticket.gold["kb_article"]
def choose_threshold(tickets: list[Ticket]) -> float:
"""Lowest threshold on dev at which at most 1 in 10 answered tickets would be unsafe."""
scored = [(KB.search(t.text, 1), t) for t in tickets]
for threshold in [x / 2 for x in range(0, 41)]:
answered = [(h, t) for h, t in scored if h and h[0].score >= threshold]
unsafe = sum(not safe_to_answer(t, h) for h, t in answered)
if answered and unsafe / len(answered) <= 0.10:
return threshold
return float("inf")
def build_messages(ticket: Ticket, hits: list[Hit]) -> list[dict]:
sources = "\n\n".join(f'<source id="{h.article_id}">\n{get_article(h.article_id).body}\n</source>' for h in hits)
return [{"role": "system", "content": (
"Answer the customer using ONLY the sources. Cite the ids you used in cited_articles. "
"If the sources do not answer the question, say so and set confidence to low. Reply as JSON "
"with keys reply, cited_articles, confidence (low, medium, high).")},
{"role": "user", "content": f"{sources}\n\n<ticket>\n{ticket.text}\n</ticket>"}]
def check_citations(draft: DraftReply, hits: list[Hit]) -> list[str]:
problems = []
shown = {h.article_id for h in hits}
if not draft.cited_articles and draft.confidence != "low":
problems.append("confident answer without citations")
problems += [f"cites {c} which was not retrieved" for c in draft.cited_articles if c not in shown]
return problems
def answer(ticket: Ticket, llm, threshold: float) -> tuple[DraftReply, str]:
hits = KB.search(ticket.text, 3)
if not hits or hits[0].score < threshold:
return ABSTAIN, f"abstain: top score {hits[0].score if hits else 0} < {threshold}"
result = llm(build_messages(ticket, hits), temperature=0.0, response_format={"type": "json_object"})
try:
draft = DraftReply.model_validate_json(result.text)
except ValidationError:
return ABSTAIN, "abstain: reply failed schema validation"
problems = check_citations(draft, hits)
if problems:
return ABSTAIN, f"abstain: {problems}"
return draft, "answered"
# STAND-IN replies (hand-written, not model output): one good, one citing an article it was never shown.
SCRIPTED = [
'{"reply": "Invoices can be reissued with a new address: update Settings > Billing > Tax details, then ask us '
'to reissue past invoices.", "cited_articles": ["billing-invoices"], "confidence": "high"}',
'{"reply": "Yes, Team includes SSO.", "cited_articles": ["billing-plans-2025"], "confidence": "high"}',
]
if __name__ == "__main__":
dev, test = load_tickets("dev"), load_tickets("test")
threshold = choose_threshold(dev)
print(f"threshold chosen on dev: {threshold}")
for name, split in (("dev", dev), ("test", test)):
answered = [t for t in split if KB.search(t.text, 1) and KB.search(t.text, 1)[0].score >= threshold]
unsafe = sum(not safe_to_answer(t, KB.search(t.text, 1)) for t in answered)
missed = sum(safe_to_answer(t, KB.search(t.text, 1)) for t in split if t not in answered)
print(f"{name}: answered {fmt_rate(len(answered), len(split))}; unsafe among answered "
f"{fmt_rate(unsafe, len(answered))}; safe tickets we abstained on: {missed}")
llm = pick_llm(ScriptedLLM(replies=list(SCRIPTED)))
for ticket_id in ("T-1041", "T-1012", "T-1036"):
ticket = next(t for t in dev + test if t.id == ticket_id)
draft, why = answer(ticket, llm, threshold)
print(f"\n{ticket_id} ({why})\n {draft.reply}\n cited={draft.cited_articles} confidence={draft.confidence}")
Code explained
- In simple words: a librarian who only answers when the right book is clearly on the desk, and a checker who confirms every footnote points to a book that was actually on the desk.
- What happens:
safe_to_answeris the gold definition of a safe answer: the ticket is answerable from the help center and our top hit is the gold article.choose_thresholdscans thresholds from 0 to 20 in steps of 0.5 on the 48 dev tickets and returns the lowest one at which at most 1 in 10 answered tickets would be unsafe. Lowest, because every step up abstains on more tickets that we could have answered.build_messageswraps each retrieved article in<source id="...">tags so the ids the model may cite are exactly the ids it was shown.answerapplies the retrieval gate, calls the model with JSON mode, validates againstDraftReply, then applies the citation gate. Every failure path returns the same politeABSTAINreply and a reason string for the logs.- The two stand-in replies are hand-written: one correct, one citing
billing-plans-2025, an article that does not exist.
- Comes out:
The table below is the decision you make for each ticket type. The measured numbers above are what let you make it with evidence.
| Situation | Use this | Why |
|---|---|---|
| Question matches a strong help-center article | Grounded answer with citations, agent reviews | Measured low wrong-source rate at high scores |
| Weak retrieval score | Abstain before calling the model | The model would answer anyway, fluently; the score is your early warning |
| Model cites something it was not shown | Treat as abstain, log it | Citation fabrication is a symptom of answering from memory |
| Non-English ticket | Translate the query or use multilingual retrieval, and measure again | Module 7 showed BM25 misses most non-English tickets |
A4. Code generation and developer tooling
Code generation is the pattern where LLMs are most useful and easiest to trust, for one reason: code can be executed, so you can check it against tests, a verifier that does not depend on the model's opinion of itself (Module 5 explained why self-critique without an external signal often fails). The shape is generate, verify, feed the error back, repeat with an attempt ceiling.
Maya asks data questions ("how many billing tickets per language?") and a text-to-SQL helper writes the query. Three guards make it safe to run model-written SQL: the statement must be a single SELECT, the connection itself refuses writes (PRAGMA query_only), and the result must match a test.
# examples/m14_sql_codegen.py
"""Pattern 4: code generation with tests as the verifier (text-to-SQL for Maya's questions).
The model writes SQL; nothing trusts it until it (1) is a single read-only
SELECT, (2) runs on a read-only database, and (3) returns the rows the test
expects. Failures go back to the model as feedback, with an attempt ceiling.
"""
from __future__ import annotations
import re
import sqlite3
from collections import Counter
from examples.m14_common import pick_llm
from supportdesk.data import load_tickets
from supportdesk.stand_in import ScriptedLLM
SCHEMA = ("CREATE TABLE tickets (id TEXT PRIMARY KEY, subject TEXT, body TEXT, customer_tier TEXT, "
"language TEXT, category TEXT, priority TEXT, answerable INTEGER)")
def make_db() -> sqlite3.Connection:
conn = sqlite3.connect(":memory:")
conn.execute(SCHEMA)
conn.executemany("INSERT INTO tickets VALUES (?,?,?,?,?,?,?,?)", [
(t.id, t.subject, t.body, t.customer_tier, t.language, t.gold["category"], t.gold["priority"],
int(t.gold["answerable"])) for t in load_tickets()])
conn.execute("PRAGMA query_only = ON") # the database itself refuses writes from here on
return conn
def extract_sql(text: str) -> str:
fenced = re.search(r"```(?:sql)?\s*(.+?)```", text, re.S)
return (fenced.group(1) if fenced else text).strip().rstrip(";")
def verify(conn: sqlite3.Connection, sql: str, expected: set[tuple]) -> str:
"""Return '' if the SQL passes, otherwise feedback the model can act on."""
if not re.fullmatch(r"(?is)\s*(select|with)\b.*", sql) or ";" in sql:
return "Only a single SELECT statement is allowed."
try:
rows = set(conn.execute(sql).fetchall())
except sqlite3.Error as exc:
return f"SQLite error: {exc}"
if rows != expected:
return f"Wrong result. Got {sorted(rows)[:5]} ({len(rows)} rows); expected {len(expected)} rows."
return ""
def generate_sql(question: str, llm, conn, expected: set[tuple], max_attempts: int = 3) -> tuple[str, int]:
messages = [{"role": "system", "content": f"Write one SQLite SELECT for this schema. Reply with SQL only.\n{SCHEMA}\n"
"category values: billing, cancellation, account_access, bug, how_to, "
"feature_request. priority values: low, normal, high, urgent."},
{"role": "user", "content": question}]
for attempt in range(1, max_attempts + 1):
reply = llm(messages, temperature=0.0).text
sql = extract_sql(reply)
feedback = verify(conn, sql, expected)
print(f" attempt {attempt}: {sql!r}\n verifier: {feedback or 'PASS'}")
if not feedback:
return sql, attempt
messages += [{"role": "assistant", "content": reply}, {"role": "user", "content": feedback + " Try again."}]
raise RuntimeError(f"no passing SQL after {max_attempts} attempts")
# STAND-IN replies (hand-written to show each verifier branch; not model output).
SCRIPTED = [
"DELETE FROM tickets WHERE category = 'billing'",
"```sql\nSELECT language, COUNT(*) FROM tickets WHERE category = 'Billing' GROUP BY language\n```",
"```sql\nSELECT language, COUNT(*) FROM tickets WHERE category = 'billing' GROUP BY language;\n```",
]
if __name__ == "__main__":
conn = make_db()
# The test: expected rows computed in plain Python from the same data, independently of any SQL.
expected = set(Counter(t.language for t in load_tickets() if t.gold["category"] == "billing").items())
print("expected:", sorted(expected))
llm = pick_llm(ScriptedLLM(replies=list(SCRIPTED)))
sql, attempts = generate_sql("How many billing tickets per language?", llm, conn, expected)
print(f"accepted after {attempts} attempts: {sql}")
try:
conn.execute("DELETE FROM tickets")
except sqlite3.OperationalError as exc:
print("direct write attempt:", exc)
Code explained
- In simple words: the model is a junior analyst whose queries run in a read-only sandbox and are graded against an answer key before anyone sees the result.
- What happens:
make_dbloads the 72 tickets into an in-memory SQLite table, then turns onPRAGMA query_only, so even a query that slips past the text check cannot modify data.extract_sqlaccepts SQL inside a code fence or bare, and strips a trailing semicolon.verifyrejects anything that is not oneSELECT(orWITH) statement, runs it, and compares the rows with the expected set. Each failure returns a sentence the model can act on.generate_sqlloops up to three attempts, appending the model's reply and the verifier's feedback to the conversation each time.- The expected rows come from plain Python (
Counterover the gold labels), independent of any SQL: a test must not share the code path it tests. - The three stand-in replies are hand-written to hit each branch: a destructive statement, a subtle case bug (
'Billing'), and a correct query.
- Comes out:text
expected: [('en', 14), ('es', 1), ('hi', 1)] [stand-in] ScriptedLLM: tests the plumbing only, not model quality attempt 1: "DELETE FROM tickets WHERE category = 'billing'" verifier: Only a single SELECT statement is allowed. attempt 2: "SELECT language, COUNT(*) FROM tickets WHERE category = 'Billing' GROUP BY language" verifier: Wrong result. Got [] (0 rows); expected 3 rows. attempt 3: "SELECT language, COUNT(*) FROM tickets WHERE category = 'billing' GROUP BY language" verifier: PASS accepted after 3 attempts: SELECT language, COUNT(*) FROM tickets WHERE category = 'billing' GROUP BY language direct write attempt: attempt to write a readonly databaseAttempt 1 is refused before it runs. Attempt 2 runs and returns zero rows, because SQLite string comparison is case sensitive; this is a realistic model mistake that looks perfectly plausible in review, and only the test catches it. Attempt 3 passes. The final line shows the second layer of defense: even a direct
DELETEon the connection fails.
In a developer tool, "the test" is usually a fixture the developer writes once: a tiny database with known answers, a set of example strings for a regex, a unit test for a function. The model then iterates against it. For questions with no known answer (a genuinely new analysis), the verifier shrinks to the safety checks plus a human reading the SQL, and you should present the query next to the result so the reader can check it.
A5. Content generation with human review
When the model writes something customers will read, such as a help-center article, the output is too open-ended to verify automatically. The pattern is to make human review a hard gate, enforced by code, not by a checklist in someone's head. A review queue is a table of drafts with a state machine: a fixed set of states and the only transitions allowed between them.
Brightlane's 10 unanswerable tickets are help-center gaps. The assistant drafts an article per gap, marks uncertain facts with [VERIFY], and the queue refuses to approve a draft that still contains a marker or to accept approval from anyone but a named human.
# examples/m14_review_queue.py
"""Pattern 5: content generation with human review (help-center article drafts).
Tickets the help center could not answer are grouped into gaps; the model drafts
an article per gap; drafts enter a review queue whose state machine makes it
impossible to publish anything a named human has not approved.
"""
from __future__ import annotations
import json
import sqlite3
from collections import defaultdict
from datetime import datetime, timezone
from examples.m14_common import pick_llm
from supportdesk.data import load_tickets
from supportdesk.stand_in import ScriptedLLM
TRANSITIONS = {"draft": {"in_review"}, "in_review": {"approved", "changes_requested", "rejected"},
"changes_requested": {"in_review"}, "approved": {"published"}, "rejected": set(), "published": set()}
class ReviewQueue:
def __init__(self) -> None:
self.db = sqlite3.connect(":memory:")
self.db.execute("CREATE TABLE items (id INTEGER PRIMARY KEY, title TEXT, body TEXT, sources TEXT, "
"state TEXT, reviewer TEXT, history TEXT)")
def add(self, title: str, body: str, sources: list[str]) -> int:
cur = self.db.execute("INSERT INTO items (title, body, sources, state, reviewer, history) VALUES (?,?,?,?,?,?)",
(title, body, json.dumps(sources), "draft", "", "[]"))
return cur.lastrowid
def move(self, item_id: int, new_state: str, actor: str, note: str = "") -> None:
state, history = self.db.execute("SELECT state, history FROM items WHERE id=?", (item_id,)).fetchone()
if new_state not in TRANSITIONS[state]:
raise ValueError(f"item {item_id}: {state} -> {new_state} is not allowed")
if new_state == "approved" and not actor.startswith("human:"):
raise PermissionError("only a named human reviewer can approve")
body = self.db.execute("SELECT body FROM items WHERE id=?", (item_id,)).fetchone()[0]
if new_state == "approved" and "[VERIFY]" in body:
raise ValueError(f"item {item_id}: resolve every [VERIFY] marker before approving")
events = json.loads(history) + [{"from": state, "to": new_state, "actor": actor, "note": note,
"at": datetime.now(timezone.utc).isoformat(timespec="seconds")}]
self.db.execute("UPDATE items SET state=?, reviewer=?, history=? WHERE id=?",
(new_state, actor if new_state == "approved" else "", json.dumps(events), item_id))
def edit(self, item_id: int, body: str, actor: str) -> None:
self.db.execute("UPDATE items SET body=? WHERE id=?", (body, item_id))
history = json.loads(self.db.execute("SELECT history FROM items WHERE id=?", (item_id,)).fetchone()[0])
history.append({"edit_by": actor})
self.db.execute("UPDATE items SET history=? WHERE id=?", (json.dumps(history), item_id))
def counts(self) -> dict[str, int]:
return dict(self.db.execute("SELECT state, COUNT(*) FROM items GROUP BY state ORDER BY state").fetchall())
def draft_messages(category: str, tickets) -> list[dict]:
asks = "\n".join(f"- [{t.id}] {t.subject}: {t.body}" for t in tickets)
return [{"role": "system", "content": "Draft a short help-center article (title line, then 3 to 6 sentences) that "
"would answer these customer questions. Mark any fact you are unsure of "
"with [VERIFY]. Never invent prices, dates, or policies."},
{"role": "user", "content": f"Category: {category}\n{asks}"}]
def responder(messages, kwargs):
"""STAND-IN: a templated draft per gap, not model output."""
category = messages[-1]["content"].splitlines()[0].split(": ")[1]
return (f"Help with {category.replace('_', ' ')} requests we cannot resolve automatically\n"
"Some requests need a person on our team. [VERIFY] Typical reply time is one business day.")
if __name__ == "__main__":
gaps = defaultdict(list)
for t in load_tickets():
if not t.gold["answerable"]:
gaps[t.gold["category"]].append(t)
print({c: [t.id for t in ts] for c, ts in gaps.items()})
llm = pick_llm(ScriptedLLM(responder=responder))
queue = ReviewQueue()
ids = {}
for category, tickets in gaps.items():
text = llm(draft_messages(category, tickets), temperature=0.3).text
title, body = text.split("\n", 1)
ids[category] = queue.add(title, body, [t.id for t in tickets])
queue.move(ids[category], "in_review", actor="system")
print("after drafting:", queue.counts())
try:
queue.move(ids["billing"], "approved", actor="human:maya")
except ValueError as exc:
print("refused:", exc)
queue.edit(ids["billing"], "Some billing requests need a person on our team. We reply within the times "
"listed in your plan.", actor="human:maya")
queue.move(ids["billing"], "approved", actor="human:maya", note="checked against billing policy")
queue.move(ids["billing"], "published", actor="system")
queue.move(ids["bug"], "changes_requested", actor="human:maya", note="remove the [VERIFY] reply-time claim")
queue.move(ids["feature_request"], "rejected", actor="human:maya", note="roadmap questions go to product")
for attempt in ((ids["account_access"], "approved", "model:auto-reviewer"), (ids["how_to"], "published", "system")):
try:
queue.move(*attempt)
except (ValueError, PermissionError) as exc:
print("refused:", exc)
print("final:", queue.counts())
unverified = queue.db.execute("SELECT COUNT(*) FROM items WHERE state='published' AND body LIKE '%[VERIFY]%'")
print("published items still containing [VERIFY]:", unverified.fetchone()[0])
Code explained
- In simple words: a publishing desk where drafts can only move along approved paths, and the stamp that says "approved" only works in a human hand.
- What happens:
TRANSITIONSlists every legal move. There is no path fromin_reviewtopublishedthat skipsapproved, andrejectedandpublishedare final.movechecks the transition, then two business rules: approvals must come from an actor whose id starts withhuman:, and a body containing[VERIFY]cannot be approved. Every move appends a timestamped event to the item's history, which is your audit trail.editrecords who changed the text, so you can later measure how much reviewers rewrite (Part B).- The main block drafts one article per gap category with the stand-in, then plays out a day: Maya's first approval is refused because of the marker, she edits and approves, one draft goes back for changes, one is rejected, and two illegal moves are attempted.
- Comes out:
A6. Data transformation and enrichment pipelines
An enrichment pipeline adds fields to records in bulk: a CRM gets industries, a product catalogue gets attributes, a ticket table gets invoice ids and an English gist for analytics. Three properties separate a pipeline you can rerun from one that corrupts your warehouse:
- Deterministic first: anything a regex or lookup can extract exactly (invoice ids, VAT numbers, plan names, seat counts) never goes to the model.
- Content-addressed caching: the cache key is a hash of the exact request (prompt version, messages, and parameters), so rerunning costs nothing and changing the prompt version invalidates the cache on purpose.
- Idempotence: running twice gives the same output, because the output file is rewritten from scratch rather than appended to.
# examples/m14_enrich_pipeline.py
"""Pattern 6: data transformation and enrichment (turning raw tickets into analytics rows).
Deterministic fields come from code (regexes, lookups). Only the fuzzy field
(a one-line English gist) comes from the model, cached by a hash of the exact
request so reruns cost nothing, validated, and routed to a dead-letter file
when invalid. The pipeline is idempotent: running it twice gives the same rows.
"""
from __future__ import annotations
import hashlib
import json
import re
from pathlib import Path
from pydantic import BaseModel, Field, ValidationError
from examples.m14_common import pick_llm
from supportdesk.data import Ticket, load_tickets
from supportdesk.stand_in import ScriptedLLM
INVOICE = re.compile(r"\bINV-\d{4}-\d{6}\b")
VAT = re.compile(r"\b(?:DE\d{9}|FR[A-Z0-9]{2}\d{9}|GB\d{9})\b")
PLAN = re.compile(r"\b(Free|Team|Business|Enterprise)\b")
SEATS = re.compile(r"\b(\d{1,5})\s+(?:people|users|seats)\b", re.I)
PROMPT_VERSION = "gist-v1"
class Gist(BaseModel):
gist_en: str = Field(min_length=8, max_length=140)
class Row(BaseModel):
ticket_id: str
invoice_ids: list[str]
vat_ids: list[str]
plans_mentioned: list[str]
seats: int | None
gist_en: str
prompt_version: str
def gist_messages(ticket: Ticket) -> list[dict]:
return [{"role": "system", "content": 'Return JSON {"gist_en": "..."}: the request in at most 15 English words.'},
{"role": "user", "content": ticket.text}]
class CachedLLM:
"""Content-addressed cache: same model + same messages = same key = no second call."""
def __init__(self, llm, path: Path) -> None:
self.llm, self.path = llm, path
self.store = json.loads(path.read_text()) if path.exists() else {}
self.hits = self.misses = 0
def __call__(self, messages: list[dict], **kwargs) -> str:
key = hashlib.sha256(json.dumps([PROMPT_VERSION, messages, kwargs], sort_keys=True).encode()).hexdigest()
if key in self.store:
self.hits += 1
else:
self.misses += 1
self.store[key] = self.llm(messages, **kwargs).text
self.path.write_text(json.dumps(self.store))
return self.store[key]
def enrich(ticket: Ticket, cached: CachedLLM) -> Row:
raw = cached(gist_messages(ticket), temperature=0.0)
gist = Gist.model_validate_json(raw) # raises ValidationError; the caller dead-letters it
seats = SEATS.search(ticket.body)
return Row(ticket_id=ticket.id, invoice_ids=INVOICE.findall(ticket.text), vat_ids=VAT.findall(ticket.text),
plans_mentioned=sorted(set(PLAN.findall(ticket.text))), seats=int(seats.group(1)) if seats else None,
gist_en=gist.gist_en, prompt_version=PROMPT_VERSION)
def run(tickets: list[Ticket], cached: CachedLLM, out: Path, dead: Path) -> tuple[int, int]:
rows, failures = [], []
for t in tickets:
try:
rows.append(enrich(t, cached).model_dump())
except ValidationError as exc:
failures.append({"ticket_id": t.id, "error": exc.errors()[0]["type"]})
out.write_text("".join(json.dumps(r, ensure_ascii=False) + "\n" for r in rows)) # overwrite, never append
dead.write_text("".join(json.dumps(f) + "\n" for f in failures))
return len(rows), len(failures)
def responder(messages, kwargs):
"""STAND-IN: uses the subject as the gist; Japanese subjects come back too short, to exercise dead-lettering."""
subject = messages[-1]["content"].splitlines()[0].removeprefix("Subject: ")
return json.dumps({"gist_en": subject if subject.isascii() else "?"})
if __name__ == "__main__":
work = Path("m14_out")
work.mkdir(exist_ok=True)
cache_file = work / "gist_cache.json"
cache_file.unlink(missing_ok=True)
cached = CachedLLM(pick_llm(ScriptedLLM(responder=responder)), cache_file)
tickets = load_tickets()
for run_no in (1, 2):
before = (cached.hits, cached.misses)
ok, bad = run(tickets, cached, work / "enriched.jsonl", work / "dead_letter.jsonl")
print(f"run {run_no}: rows {ok}, dead-letter {bad}, cache hits {cached.hits - before[0]}, "
f"model calls {cached.misses - before[1]}")
rows = [json.loads(line) for line in (work / "enriched.jsonl").read_text().splitlines()]
for r in rows:
if r["invoice_ids"] or r["vat_ids"] or r["seats"]:
print(r)
print("dead-letter:", (work / "dead_letter.jsonl").read_text().strip().replace("\n", " | "))
Code explained
- In simple words: a factory line where the measuring machines do the exact work, the model adds one fuzzy label, and every model answer is kept in a labelled drawer so you never pay for it twice.
- What happens:
- Four compiled regexes pull invoice ids, VAT ids, plan names, and seat counts. Part C tests the invoice regex properly.
CachedLLMhashes[PROMPT_VERSION, messages, kwargs]with SHA-256 and stores the raw reply under that key in a JSON file.enrichvalidates the model's JSON with the smallGistschema and builds a typedRow; any validation error propagates.runcatches validation errors into a dead-letter file and overwrites both output files, which is what makes reruns idempotent.- The stand-in uses the ticket subject as the gist and returns
"?"for subjects that are not ASCII, which fails the 8-character minimum and exercises dead-lettering.
- Comes out:text
[stand-in] ScriptedLLM: tests the plumbing only, not model quality run 1: rows 69, dead-letter 3, cache hits 0, model calls 72 run 2: rows 69, dead-letter 3, cache hits 72, model calls 0 {'ticket_id': 'T-1001', 'invoice_ids': ['INV-2026-004512'], 'vat_ids': [], 'plans_mentioned': ['Team'], 'seats': None, 'gist_en': 'Charged twice this month', 'prompt_version': 'gist-v1'} {'ticket_id': 'T-1002', 'invoice_ids': [], 'vat_ids': [], 'plans_mentioned': ['Business'], 'seats': 14, 'gist_en': 'How much is Business?', 'prompt_version': 'gist-v1'} {'ticket_id': 'T-1006', 'invoice_ids': [], 'vat_ids': ['DE811234567'], 'plans_mentioned': [], 'seats': None, 'gist_en': 'VAT number on invoice', 'prompt_version': 'gist-v1'} {'ticket_id': 'T-1007', 'invoice_ids': [], 'vat_ids': [], 'plans_mentioned': ['Business'], 'seats': 40, 'gist_en': 'Pay by bank transfer?', 'prompt_version': 'gist-v1'} {'ticket_id': 'T-1031', 'invoice_ids': ['INV-2026-004871'], 'vat_ids': [], 'plans_mentioned': ['Team'], 'seats': None, 'gist_en': 'Cobro duplicado', 'prompt_version': 'gist-v1'} {'ticket_id': 'T-1043', 'invoice_ids': [], 'vat_ids': [], 'plans_mentioned': ['Enterprise'], 'seats': 500, 'gist_en': 'Enterprise pricing', 'prompt_version': 'gist-v1'} dead-letter: {"ticket_id": "T-1034", "error": "string_too_short"} | {"ticket_id": "T-1035", "error": "string_too_short"} | {"ticket_id": "T-1062", "error": "string_too_short"}The first run makes 72 model calls; the second makes none and produces identical files. The printed rows show the deterministic fields doing real work:
INV-2026-004871pulled from a Spanish ticket,DE811234567recognized as a VAT id, 14, 40, and 500 seats. The three dead letters are the Japanese and Hindi tickets. Two honest caveats. First, the cache stores the bad replies too, so a rerun reproduces the failures without retrying them; in production, cache only replies that validated. Second, schema validation cannot tell that "Cobro duplicado" is not English. If a field's language matters, check it with a language detector, or accept that a schema only checks shape.
A7. Search and question answering over private corpora
The last pattern is search over data only your organization has: past tickets, contracts, internal wikis. The new problem is permissions. A support agent may read most past tickets, but Enterprise tickets can contain contract terms that only the Enterprise pod may see. The rule is to filter before ranking: restricted documents must never enter the candidate set, not be removed from the results afterward.
Filtering afterward leaks in two ways. The obvious one is a bug that forgets the filter. The subtle one is that restricted documents still shape the scores of everything else, through the word statistics a ranker computes over the whole collection.
# examples/m14_ticket_search.py
"""Pattern 7: search and question answering over a private corpus (past tickets).
Agents ask "have we seen this before?". We index past (dev) tickets with the same
BM25 class the help center uses, filter by permission BEFORE ranking, and
measure how often the most similar past ticket shares the new ticket's category.
"""
from __future__ import annotations
from examples.m14_common import fmt_rate
from supportdesk.data import Article, Ticket, load_tickets
from supportdesk.kb_search import KBSearch
# Who may read what: Enterprise tickets can contain contract terms, so only the enterprise pod sees them.
ACL = {"free": {"all-agents"}, "team": {"all-agents"}, "business": {"all-agents"}, "enterprise": {"enterprise-pod"}}
def as_article(t: Ticket) -> Article:
return Article(id=t.id, title=t.subject, tags=(t.customer_tier,), body=t.body)
class TicketSearch:
def __init__(self, past: list[Ticket]) -> None:
self.past = {t.id: t for t in past}
self._indexes: dict[frozenset[str], KBSearch] = {}
def _index_for(self, groups: frozenset[str]) -> KBSearch:
"""One BM25 index per permission set, built only from tickets those groups may read.
Filtering before ranking matters twice: restricted text never reaches the results, and it
never influences the word statistics (IDF) that score everyone else's results either.
"""
if groups not in self._indexes:
visible = [as_article(t) for t in self.past.values() if ACL[t.customer_tier] & groups]
self._indexes[groups] = KBSearch(visible)
return self._indexes[groups]
def search(self, query: str, groups: set[str], k: int = 3):
return self._index_for(frozenset(groups)).search(query, k)
if __name__ == "__main__":
dev, test = load_tickets("dev"), load_tickets("test")
search = TicketSearch(dev)
agree = found = 0
for t in test:
hits = search.search(t.text, {"all-agents", "enterprise-pod"}, k=1)
if hits:
found += 1
agree += search.past[hits[0].article_id].gold["category"] == t.gold["category"]
print(f"test tickets with any similar past ticket: {found}/{len(test)}")
print(f"top-1 past ticket has the same category: {fmt_rate(agree, found)}")
query = "custom SAML login domain for our Okta users"
for groups in ({"all-agents"}, {"all-agents", "enterprise-pod"}):
print(sorted(groups), [(h.article_id, h.title, h.score) for h in search.search(query, groups)])
Code explained
- In simple words: each team gets its own card catalogue built only from the books it may read, instead of one catalogue with some cards blacked out.
- What happens:
ACLmaps each customer tier to the groups allowed to read those tickets.as_articlewraps a past ticket in theArticletype so the canonicalKBSearchclass can index tickets without any new search code._index_forbuilds, and caches, one BM25 index per permission set from only the visible tickets.- The evaluation indexes the 48 dev tickets as "the past" and, for each test ticket, checks whether the single most similar past ticket has the same category.
- The last query runs twice, as an ordinary agent and as a member of the Enterprise pod.
- Comes out:text
test tickets with any similar past ticket: 24/24 top-1 past ticket has the same category: 11/24 = 45.8% (95% CI 27.9% to 64.9%) ['all-agents'] [('T-1047', 'SSO users see reset link', 3.649), ('T-1013', 'SCIM provisioning', 2.597), ('T-1007', 'Pay by bank transfer?', 2.236)] ['all-agents', 'enterprise-pod'] [('T-1065', 'Custom domain', 14.473), ('T-1011', 'SAML signature invalid', 9.291), ('T-1047', 'SSO users see reset link', 3.495)]Every test ticket finds some similar past ticket, but the top one shares its category only 11 times in 24, with an interval from 28% to 65%. Similar-ticket search is useful for an agent to read ("we solved this before in T-1065"), not as a classifier. The permission check behaves as designed: an ordinary agent never sees T-1065 or T-1011, the two Enterprise tickets that best match the query. Now look at T-1047, which both users can see: its score is 3.649 for the agent and 3.495 for the Enterprise pod. The only difference is which other documents were in the index. Rank against a shared index and filter afterward, and restricted documents quietly change the order of everyone's results.
| Situation | Use this | Why |
|---|---|---|
| Few permission groups (tens) | One index per group or per permission set | Simple, and statistics never mix |
| Many users with individual rights | One index with the ACL as a pre-filter inside the query | Vector databases and search engines support filtered search |
| Answers generated from search results | Filter, then retrieve, then generate | The model can only leak what it was shown |