Part B: Design Principles
The patterns tell you what to build. The principles in this part decide whether people can live with it: who acts, what happens when the model is wrong, where humans look, what users are told, and what you learn from every interaction.
B1. Matching autonomy to consequence
Autonomy is how much the system does without a person approving it. The mistake teams make is deciding autonomy per feature by feel ("the model is good at refunds now"). The robust approach derives autonomy from properties of the action itself, because consequence is a property of the action, not of the model:
- Can it be undone quickly and completely?
- Does the customer see it?
- Does it move money, delete data, or change access?
- How many customers can one mistake reach (the blast radius)?
The same rules in code, with the undo log and the fallback that implement the next principle:
# examples/m14_autonomy.py
"""Design principles: match autonomy to consequence, and design for the model being wrong.
Autonomy is derived from properties of the action (can it be undone, does the
customer see it, does it move money or data), not decided case by case. Every
automatic action is logged with its inverse so a human can undo it, and a model
failure falls back to a safe default instead of blocking the queue.
"""
from __future__ import annotations
from collections.abc import Callable
from dataclasses import dataclass, field
from pydantic import ValidationError
from supportdesk.schemas import Triage
from supportdesk.stand_in import ScriptedLLM
LEVELS = ("auto", "auto_then_notify", "suggest_and_approve", "human_only")
@dataclass(frozen=True)
class Action:
name: str
reversible: bool # can we fully undo it within minutes?
customer_visible: bool # does the customer see or receive it?
money_or_data: bool # does it move money, delete data, or change access?
blast_radius: int # how many customers one mistake can reach
def autonomy_for(a: Action) -> str:
if a.money_or_data and not a.reversible:
return "human_only"
if a.money_or_data or a.customer_visible or a.blast_radius > 1:
return "suggest_and_approve"
if not a.reversible:
return "auto_then_notify"
return "auto"
ACTIONS = [
Action("add_category_tag", True, False, False, 1),
Action("set_priority", True, False, False, 1),
Action("route_to_queue", True, False, False, 1),
Action("post_internal_note", False, False, False, 1),
Action("send_reply_with_kb_link", False, True, False, 1),
Action("send_bulk_incident_update", False, True, False, 500),
Action("unlock_account", True, True, True, 1),
Action("issue_refund", False, True, True, 1),
Action("delete_personal_data", False, True, True, 1),
]
@dataclass
class Ticketish:
id: str
tags: set[str] = field(default_factory=set)
priority: str = "normal"
class ActionLog:
"""Every automatic change is stored with the operation that reverses it."""
def __init__(self) -> None:
self.entries: list[tuple[str, str, Callable[[], None]]] = []
def apply(self, ticket: Ticketish, name: str, do, undo) -> None:
do()
self.entries.append((ticket.id, name, undo))
def undo_last(self, ticket_id: str) -> str:
for i in range(len(self.entries) - 1, -1, -1):
if self.entries[i][0] == ticket_id:
_, name, undo = self.entries.pop(i)
undo()
return name
return "nothing to undo"
def triage_or_fallback(llm, text: str) -> tuple[Triage | None, str]:
"""If the model call fails or returns junk, degrade to 'normal priority, human queue', never block."""
try:
return Triage.model_validate_json(llm([{"role": "user", "content": text}]).text), "model"
except (ValidationError, RuntimeError, TimeoutError) as exc:
return None, f"fallback ({type(exc).__name__}): route to human queue at normal priority"
if __name__ == "__main__":
print(f"{'action':28} {'rev':>4} {'seen':>5} {'$/data':>6} {'reach':>6} autonomy")
for a in ACTIONS:
print(f"{a.name:28} {str(a.reversible)[0]:>4} {str(a.customer_visible)[0]:>5} "
f"{str(a.money_or_data)[0]:>6} {a.blast_radius:>6} {autonomy_for(a)}")
log, ticket = ActionLog(), Ticketish("T-1001")
old = ticket.priority
log.apply(ticket, "add_category_tag", lambda: ticket.tags.add("billing"), lambda: ticket.tags.discard("billing"))
log.apply(ticket, "set_priority", lambda: setattr(ticket, "priority", "low"),
lambda: setattr(ticket, "priority", old))
print("\nafter model actions:", ticket)
print("agent clicks undo:", log.undo_last("T-1001"), "->", ticket)
# STAND-IN replies: one valid, one truncated, then an outage (not model output).
replies = ['{"category": "billing", "priority": "high", "language": "en", "summary": "Duplicate charge.", '
'"needs_human": true}', '{"category": "billing", "prio']
llm = ScriptedLLM(replies=replies)
for _ in range(3):
triage, source = triage_or_fallback(llm, "Charged twice this month")
print(source if triage is None else f"{source}: {triage.category}/{triage.priority}")
Code explained
- In simple words: a rulebook that looks at what an action can break, not at how clever the model seems, and a notebook that remembers how to take back everything the model did.
- What happens:
Actionrecords the four consequence properties;autonomy_forturns them into one of four levels, checking the most dangerous combination first.ACTIONSlists nine real Brightlane actions from the earlier modules.ActionLog.applyperforms a change and stores its inverse;undo_lastpops the most recent change for a ticket and runs the inverse. This is what the agent's Undo button calls.triage_or_fallbackcatches validation errors, a stand-in outage (RuntimeErrorwhen the script runs out of replies), and timeouts, and returns a safe default instead of raising.- The stand-in replies are hand-written: one valid triage, one truncated reply, then nothing (an outage).
- Comes out:
Autonomy should move only with evidence. Start an action one level more conservative than the table says, measure the approval rate and the edit rate from the feedback events in B6, and relax the level when the numbers justify it.
B2. Designing for the model being wrong
Every model is wrong some of the time, so the design question is never "how do we stop errors?" but "what happens next?". Four mechanisms cover most cases, and you have now built all of them:
| Situation | Use this | Why |
|---|---|---|
| Model output is malformed or the call fails | Fallback to a safe default (human queue, rules, previous answer) | The queue keeps moving; nobody is blocked by an outage |
| Model takes a reversible action that is wrong | Undo, logged with the action | Mistakes cost one click instead of a support escalation |
| Model is unsure or the evidence is weak | Abstain (A3) | "I don't know" is cheaper than a confident wrong answer |
| Model's output feeds another system | Validation at the boundary (A1, A4, A6) | Errors stop at the boundary instead of spreading |
The design test is simple: pick any model call in your system, assume it returns the worst plausible output, and trace what the user sees. If the answer is "a wrong refund" or "a crash", that call needs one of the rows above.
B3. Human review placement and reviewer fatigue
Putting "a human in the loop" is not a safety measure until you know how often that human catches errors. Reviewers who approve hundreds of mostly correct drafts get faster and less careful; this is sometimes called automation bias, the tendency to accept a system's suggestion because it is usually right. So the design questions are where to place review and how much to put in front of one person at once.
We have no measured Brightlane reviewer data, so the next script makes its assumption explicit and then does exact math on it: the chance of catching a wrong draft starts at 95% and decays toward 60% as a session goes on, with a time constant of 40 items. The numbers are invented for teaching; the method is what you reuse.
# examples/m14_reviewer_fatigue.py
"""Design principles: where to place human review, and what reviewer fatigue does to it.
ASSUMPTION (not measured, replace with your own audit data): the chance a reviewer
catches a wrong draft decays with the number of items already reviewed in the
session, p(i) = floor + (start - floor) * exp(-i / tau). The math that follows
from the assumption is exact, and a seeded Monte Carlo run checks it.
"""
from __future__ import annotations
import math
import numpy as np
from examples.m14_common import fmt_rate
START, FLOOR, TAU = 0.95, 0.60, 40.0 # ASSUMED reviewer behaviour
ERROR_RATE = 0.08 # ASSUMED share of drafts that are wrong
DAILY = 300 # drafts per day needing a decision
def catch_prob(i: int) -> float:
return FLOOR + (START - FLOOR) * math.exp(-i / TAU)
def expected_escapes(n_items: int, error_rate: float, session: int) -> float:
"""Expected wrong drafts that get through when n_items are reviewed in sessions of `session` items."""
return sum(error_rate * (1 - catch_prob(i % session)) for i in range(n_items))
def simulate(n_items: int, error_rate: float, session: int, days: int, seed: int = 14) -> float:
rng = np.random.default_rng(seed)
p = np.array([catch_prob(i % session) for i in range(n_items)])
wrong = rng.random((days, n_items)) < error_rate
caught = rng.random((days, n_items)) < p
return float((wrong & ~caught).sum(axis=1).mean())
if __name__ == "__main__":
print("session length vs average catch rate (under the assumption)")
for n in (10, 25, 50, 100, 200, 300):
avg = sum(catch_prob(i) for i in range(n)) / n
print(f" {n:>3} items in one sitting: average catch {avg:.1%}, last item {catch_prob(n - 1):.1%}")
print(f"\n{DAILY} drafts/day, {ERROR_RATE:.0%} wrong = {DAILY * ERROR_RATE:.0f} wrong drafts/day")
plans = {
"A review all, one sitting": (DAILY, ERROR_RATE, DAILY, 0.0),
"B review all, breaks every 40": (DAILY, ERROR_RATE, 40, 0.0),
}
# C and D: review only drafts flagged low-confidence. ASSUMED: 30% are flagged, holding 70% or 90% of errors.
flagged = int(DAILY * 0.30)
for label, share in (("C", 0.70), ("D", 0.90)):
rate = DAILY * ERROR_RATE * share / flagged
plans[f"{label} review flagged 30% ({share:.0%} of errors)"] = (flagged, rate, 40, DAILY * ERROR_RATE * (1 - share))
for name, (n, rate, session, unreviewed) in plans.items():
exact = expected_escapes(n, rate, session) + unreviewed
sim = simulate(n, rate, session, days=20_000) + unreviewed
print(f" {name:40} reviewed {n:>3} escapes/day exact {exact:5.2f} simulated {sim:5.2f}")
print("\nmeasuring the real curve: plant known-wrong drafts and count catches")
for caught, planted in ((17, 20), (85, 100)):
print(f" caught {fmt_rate(caught, planted)}")
Code explained
- In simple words: if attention fades during a session, count how many mistakes slip through for each way of organizing the work, and double-check the count by simulating many days.
- What happens:
catch_prob(i)is the assumed catch rate for the i-th item in a session.expected_escapessums, over every reviewed item, the chance it is wrong times the chance the reviewer misses it. Sessions restart the fatigue clock (i % session).simulatedraws 20,000 random days with a seeded generator and counts wrong drafts that were not caught, to check the algebra.- Four plans are compared at 300 drafts a day and an assumed 8% error rate: A reviews everything in one sitting, B takes a break every 40 items, and C and D review only the 30% of drafts flagged as low confidence, where the flag is assumed to capture 70% (C) or 90% (D) of the wrong drafts. Unflagged drafts go out unreviewed, so their errors all escape.
- The last lines show how you would measure the real curve: plant known-wrong drafts in the queue and count catches, with a confidence interval.
- Comes out:text
session length vs average catch rate (under the assumption) 10 items in one sitting: average catch 91.4%, last item 87.9% 25 items in one sitting: average catch 86.4%, last item 79.2% 50 items in one sitting: average catch 80.2%, last item 70.3% 100 items in one sitting: average catch 73.0%, last item 62.9% 200 items in one sitting: average catch 67.0%, last item 60.2% 300 items in one sitting: average catch 64.7%, last item 60.0% 300 drafts/day, 8% wrong = 24 wrong drafts/day A review all, one sitting reviewed 300 escapes/day exact 8.47 simulated 8.49 B review all, breaks every 40 reviewed 300 escapes/day exact 4.14 simulated 4.13 C review flagged 30% (70% of errors) reviewed 90 escapes/day exact 9.99 simulated 9.99 D review flagged 30% (90% of errors) reviewed 90 escapes/day exact 5.99 simulated 6.00 measuring the real curve: plant known-wrong drafts and count catches caught 17/20 = 85.0% (95% CI 64.0% to 94.8%) caught 85/100 = 85.0% (95% CI 76.7% to 90.7%)Under this assumption, one sitting of 300 items catches only 64.7% of errors on average, and about 8.5 of the 24 daily wrong drafts escape. Breaks every 40 items halve that to about 4.1 with no extra headcount. Reviewing only flagged drafts cuts review work by 70%, but whether it is safe depends entirely on the flag: a flag that captures 70% of errors lets about 10 a day through (worse than reviewing everything), while a flag that captures 90% lets about 6 through. The exact and simulated columns agree to within 0.02, which confirms the arithmetic, not the assumption. The planted-error lines show why the audit needs volume: 17 of 20 caught gives an interval from 64% to 95%, too wide to choose between these plans; 100 planted errors narrows it to 77% to 91%.
| Situation | Use this | Why |
|---|---|---|
| Few items, high consequence (refunds, deletions) | Review every item, short sessions | Volume is low enough to keep attention high |
| Many items, a confidence signal you have measured | Review flagged items, audit a random sample of the rest | Spends attention where the errors are; the audit catches a bad flag |
| Many items, no measured signal | Review all, with breaks and planted errors | You cannot target review until you know where errors are |
| Reviewers approve over 98% of items | Suspect fatigue before you celebrate | Plant errors and check the catch rate |
B4. Progressive disclosure of AI involvement
Progressive disclosure means telling each audience as much about the AI's role as it needs, at the moment it needs it. Agents need to know a draft came from the model, which sources it used, and how much to trust it. Customers need to know when an answer came from a machine and how to reach a person. Module 11 covered disclosure obligations at a working level; the design point here is that the disclosure should match what actually happened, so it has to be generated from the same state machine that decided who acted.
B5. Trust calibration in the interface
Trust calibration is getting users to trust the output as much as it deserves, no more and no less. Two interface habits do most of the work. Show the evidence (the cited articles, by title) rather than asking users to trust a sentence. And show a confidence label only if you have measured what it means. A model's self-reported "confidence: high" is not calibrated (Module 1); a retrieval score band whose hit rate you measured is.
The draft card below does both, and also produces the customer disclosure lines from B4.
# examples/m14_draft_card.py
"""Design principles: trust calibration and progressive disclosure in the interface.
The agent sidebar shows WHY to trust a draft (sources, and a confidence label
whose meaning was measured), not just the draft. The customer sees a disclosure
line that matches how much the AI did.
"""
from __future__ import annotations
from examples.m14_common import fmt_rate
from supportdesk.data import get_article, load_tickets
from supportdesk.kb_search import KBSearch
from supportdesk.schemas import DraftReply
KB = KBSearch()
BANDS = [(12.0, "strong match"), (6.5, "partial match"), (0.0, "weak match")]
def band(score: float) -> str:
return next(label for floor, label in BANDS if score >= floor)
def calibration_table(tickets) -> dict[str, tuple[int, int]]:
"""How often the top article is the right one, per displayed band (measured, answerable tickets only)."""
table = {label: [0, 0] for _, label in BANDS}
for t in tickets:
hits = KB.search(t.text, 1)
if t.gold["answerable"] and hits:
cell = table[band(hits[0].score)]
cell[0] += hits[0].article_id == t.gold["kb_article"]
cell[1] += 1
return {k: (v[0], v[1]) for k, v in table.items()}
def agent_card(ticket_id: str, draft: DraftReply, top_score: float, table) -> str:
right, n = table[band(top_score)]
lines = [f"AI DRAFT for {ticket_id} (review before sending)",
f"Source match: {band(top_score)}; historically the right article {right} of {n} times",
"Sources:"]
lines += [f" [{a}] {get_article(a).title}" for a in draft.cited_articles] or [" none: treat as unsupported"]
lines += ["Draft:", f" {draft.reply}", "[Send] [Edit] [Discard] [Wrong source]"]
return "\n".join(lines)
DISCLOSURE = {
"human_wrote": "",
"ai_drafted_human_sent": "This reply was drafted with AI assistance and reviewed by {agent}.",
"ai_sent": "This is an automated answer from our help center. Reply 'agent' to reach a person.",
}
def customer_footer(mode: str, agent: str = "Maya") -> str:
return DISCLOSURE[mode].format(agent=agent)
if __name__ == "__main__":
table = calibration_table(load_tickets("dev"))
for label, (right, n) in table.items():
print(f"{label:14} top article correct: {fmt_rate(right, n) if n else 'no tickets'}")
ticket = next(t for t in load_tickets("test") if t.id == "T-1048")
top = KB.search(ticket.text, 1)[0]
draft = DraftReply(reply="After 5 failed sign-in attempts an account locks for 15 minutes. For security "
"reasons our team cannot unlock it sooner.",
cited_articles=[top.article_id], confidence="medium") # hand-written example draft
print("\n" + agent_card(ticket.id, draft, top.score, table))
for mode in DISCLOSURE:
print(f"\nfooter[{mode}]: {customer_footer(mode) or '(none)'}")
Code explained
- In simple words: instead of a label that says "trust me", the card says "drafts with this kind of source match were right 18 times out of 19", and it shows the sources.
- What happens:
BANDSsplits the top BM25 score into three bands. The 6.5 boundary is the abstain threshold from A3; 12 is a round number above it.calibration_tablemeasures, on the answerable dev tickets, how often the top article is the gold one within each band.agent_cardrenders the sidebar card: an explicit "AI DRAFT" label, the measured band statement, the cited articles by title, the draft, and four buttons. "Wrong source" is a feedback button, not decoration (B6).DISCLOSUREmaps who actually acted to the line the customer sees.- The draft for T-1048 is written by hand from the
account-loginarticle.
- Comes out:text
strong match top article correct: 9/9 = 100.0% (95% CI 70.1% to 100.0%) partial match top article correct: 18/19 = 94.7% (95% CI 75.4% to 99.1%) weak match top article correct: 8/14 = 57.1% (95% CI 32.6% to 78.6%) AI DRAFT for T-1048 (review before sending) Source match: partial match; historically the right article 18 of 19 times Sources: [account-login] Login problems and password resets Draft: After 5 failed sign-in attempts an account locks for 15 minutes. For security reasons our team cannot unlock it sooner. [Send] [Edit] [Discard] [Wrong source] footer[human_wrote]: (none) footer[ai_drafted_human_sent]: This reply was drafted with AI assistance and reviewed by Maya. footer[ai_sent]: This is an automated answer from our help center. Reply 'agent' to reach a person.The bands are genuinely informative: strong matches were right 9 of 9 times, partial 18 of 19, weak only 8 of 14. That is a calibration statement you can defend, with the sample sizes attached. Note what it measures: whether retrieval found the right article, not whether the drafted reply is correct. The card for T-1048 shows its evidence (the source title) next to the draft. The three footers show progressive disclosure: nothing when a human wrote the reply, an assistance line when a human reviewed a draft, and an explicit automated-answer line with a way out when nobody did.
B6. Feedback capture from day one
Every interaction with a draft is a free label: sent unchanged means good, heavily edited means partly wrong, discarded means wrong, "wrong source" pinpoints retrieval. Teams that skip this on day one spend month three unable to answer "did the new prompt help?". The fix is a fixed event schema that every surface writes to, carrying the prompt version and model so any change can be compared later (Module 10 used these online signals; this is where they come from).
# examples/m14_feedback_events.py
"""Design principles: capture feedback from day one with a fixed event schema.
Every draft shown to an agent produces events. Because each event carries the
prompt version and model, next month's question "did v2 help?" is a query, not
an archaeology project. Edit size is measured, not self-reported.
"""
from __future__ import annotations
import difflib
import hashlib
import json
import random
import uuid
from collections import defaultdict
from datetime import datetime, timezone
from pathlib import Path
from typing import Literal
from pydantic import BaseModel, Field
from examples.m14_common import fmt_rate
EventType = Literal["draft_shown", "draft_sent", "draft_discarded", "thumbs_up", "thumbs_down", "escalated",
"wrong_source_flagged", "undo"]
class FeedbackEvent(BaseModel):
event_id: str = Field(default_factory=lambda: uuid.uuid4().hex)
at: str = Field(default_factory=lambda: datetime.now(timezone.utc).isoformat(timespec="seconds"))
event: EventType
ticket_id: str
draft_id: str
feature: Literal["triage", "draft_reply", "kb_article_draft"]
prompt_version: str
model: str
agent: str # pseudonymous id, never a name or email
edit_ratio: float | None = None # 0 = sent unchanged, 1 = fully rewritten
cited_articles: list[str] = []
cost_usd: float | None = None
latency_ms: float | None = None
def pseudonym(agent_email: str, salt: str = "rotate-me-quarterly") -> str:
return "agent_" + hashlib.sha256((salt + agent_email).encode()).hexdigest()[:10]
def edit_ratio(draft: str, sent: str) -> float:
"""1 minus difflib's similarity: how much of the draft the agent changed."""
return round(1 - difflib.SequenceMatcher(None, draft, sent).ratio(), 3)
def log(event: FeedbackEvent, path: Path) -> None:
with path.open("a", encoding="utf-8") as f:
f.write(event.model_dump_json() + "\n")
def summarize(path: Path) -> dict[str, dict[str, float]]:
by_version: dict[str, list[FeedbackEvent]] = defaultdict(list)
for line in path.read_text().splitlines():
e = FeedbackEvent.model_validate_json(line)
by_version[e.prompt_version].append(e)
report = {}
for version, events in sorted(by_version.items()):
shown = sum(e.event == "draft_shown" for e in events)
sent = [e for e in events if e.event == "draft_sent"]
unchanged = sum(e.edit_ratio is not None and e.edit_ratio < 0.05 for e in sent)
report[version] = {"shown": shown, "sent": len(sent), "sent_unchanged": unchanged,
"discarded": sum(e.event == "draft_discarded" for e in events),
"mean_edit_ratio": round(sum(e.edit_ratio for e in sent) / max(len(sent), 1), 3)}
return report
if __name__ == "__main__":
draft = "After 5 failed sign-in attempts an account locks for 15 minutes."
for sent in (draft, draft + " Sorry for the trouble!",
"Your account locks for 15 minutes after 5 failed attempts; please wait and try again."):
print(f"edit_ratio={edit_ratio(draft, sent):.3f} sent: {sent}")
# SIMULATED traffic (seeded random, NOT real agent behaviour) just to exercise the pipeline end to end.
path = Path("m14_out/feedback.jsonl")
path.parent.mkdir(exist_ok=True)
path.unlink(missing_ok=True)
rng = random.Random(14)
agent = pseudonym("maya@brightlane.example")
for version, p_discard, typical_edit in (("draft-v1", 0.30, 0.35), ("draft-v2", 0.15, 0.20)):
for i in range(40):
common = dict(ticket_id=f"T-{2000 + i}", draft_id=f"{version}-{i}", feature="draft_reply",
prompt_version=version, model="openai/gpt-oss-120b", agent=agent)
log(FeedbackEvent(event="draft_shown", **common), path)
if rng.random() < p_discard:
log(FeedbackEvent(event="draft_discarded", **common), path)
else:
ratio = max(0.0, min(1.0, rng.gauss(typical_edit, 0.15)))
log(FeedbackEvent(event="draft_sent", edit_ratio=round(ratio, 3), **common), path)
print(f"\n{sum(1 for _ in path.open())} events written for agent {agent}")
report = summarize(path)
print(json.dumps(report, indent=1))
for version, r in report.items():
print(f"{version} discard rate {fmt_rate(r['discarded'], r['shown'])}")
Code explained
- In simple words: a receipt printed for every draft, with enough detail to reconstruct later which prompt and model produced it and what the agent did with it.
- What happens:
FeedbackEventis a pydantic model with an enumeratedeventtype, so a typo such as"draft_send"is rejected at write time instead of silently becoming a new category.pseudonymreplaces the agent's email with a salted hash, so you can count per-agent behaviour without storing who they are (Module 11 on logging policy).edit_ratiois one minusdifflib.SequenceMatcher.ratio(): 0 when the draft was sent unchanged, near 1 when it was rewritten. It is measured from text, not self-reported.summarizegroups events by prompt version.- The traffic is simulated with a seeded random generator, not real agent behaviour: 40 drafts per version, with v2 given a lower discard chance and smaller edits.
- Comes out:text
edit_ratio=0.000 sent: After 5 failed sign-in attempts an account locks for 15 minutes. edit_ratio=0.152 sent: After 5 failed sign-in attempts an account locks for 15 minutes. Sorry for the trouble! edit_ratio=0.584 sent: Your account locks for 15 minutes after 5 failed attempts; please wait and try again. 160 events written for agent agent_0e1a1c66f4 { "draft-v1": { "shown": 40, "sent": 29, "sent_unchanged": 0, "discarded": 11, "mean_edit_ratio": 0.393 }, "draft-v2": { "shown": 40, "sent": 35, "sent_unchanged": 2, "discarded": 5, "mean_edit_ratio": 0.203 } } draft-v1 discard rate 11/40 = 27.5% (95% CI 16.1% to 42.8%) draft-v2 discard rate 5/40 = 12.5% (95% CI 5.5% to 26.1%)The first three lines are real
edit_ratiovalues on real text pairs: adding an apology scores 0.152, a paraphrase 0.584. The summary shows the query you will want every week. And the last two lines show why the confidence interval belongs in that report: with 40 drafts each, v2's discard rate of 12.5% versus v1's 27.5% has overlapping intervals, so even on this simulated data you should call it promising, not proven, and collect more events.
Log the events from the first day of the pilot, even before anyone looks at them. A good minimum set is: draft shown, sent (with edit ratio), discarded, escalated, wrong source flagged, undo, plus the prompt version, model, cost, and latency on every event.