Part D: Human evaluation
When humans are unavoidable
Deterministic checks cannot tell whether a reply is helpful, whether its tone fits an angry customer, or whether a correct-looking sentence is subtly wrong for this customer. Someone has to judge. Humans are unavoidable in four places:
- Defining the standard. A rubric is a human decision. No model can tell you what Maya considers "good enough to send".
- Calibrating automated judges. An LLM judge is only as trustworthy as its measured agreement with human labels (Part E).
- High-stakes or novel failures. Legal, security, and money questions, and any failure type you have not seen before.
- Periodic audits. Judges and checks drift as the product changes; a monthly human sample catches it.
Human time is expensive, so spend it where it buys the most: labelling a few dozen items well beats skimming hundreds.
Rating design
Good human labels come from a clear question, a small scale with written anchors, and examples. For Brightlane drafts the question is the one agents actually face: can I send this?
| Design choice | Brightlane choice | Why |
|---|---|---|
| Scale | 3 points: 1 Reject, 2 Edit, 3 Send | Each point is a different action; a 1 to 10 scale invites disagreement between 6 and 7 that means nothing |
| Anchors | Written rules for each point, including tie-breakers | Raters apply rules, not moods |
| Unit | One draft for one ticket, with the help-center article visible | Raters must check claims against the source |
| Notes | A one-line reason per label | Reasons make adjudication fast and turn into judge rubric text |
| Blind | Raters do not see which system wrote the draft | Prevents "the new model is probably better" bias |
The rubric lives in examples/m10_human.py so the judge prompt in Part E uses the exact same words:
3 = Send: every claim is supported by the help-center article, it answers what was asked
directly, it follows policy, and it needs at most a trivial edit.
2 = Edit: right direction but needs a real edit before sending: part of the question is
unanswered, the answer has to be inferred, it is padded with irrelevant sentences,
the tone is off, or it states something the article does not support.
1 = Reject: wrong or unrelated facts, a claim the customer could act on wrongly (invented
settings, features, or capabilities), a policy violation (promising a roadmap date,
claiming a refund, unlock, or deletion was done), or it does not engage with the problem.
Code explained
- In simple words: three scores, each tied to what the agent does next, with the tricky cases spelled out.
- What happens: the rubric separates "unsupported but harmless" (at most 2) from "unsupported and the customer would act on it" (1), because an invented settings path sends a customer somewhere that does not exist. Adjudication below added no new rules; it applied these.
- Comes out: the text that raters read and that
pointwise_messagesembeds in the judge prompt.
The labelled set, and what it really is
examples/m10_human.py build produces 30 drafts for the first 30 English, answerable dev tickets: 10 from the v2 baseline (real output), 5 sampled from TinyLM with seed 0 (real output), and 15 written by hand for this module to cover failures our systems do not produce: policy violations, confident wrong facts, padding, curt tone, invented settings paths, and genuinely good replies.
Be clear about the labels. evals/labels.jsonl holds two label sets, "rater A" and "rater B", plus an adjudicated label and notes. They were written by the author of this module, not by Brightlane support agents, in two separate labelling passes with the rubric above. They are a worked example that makes every computation below real and checkable, not evidence about how real raters behave: two passes by one person agree more than two different people usually would. Replace them with labels from two of your own agents before you trust any number in this part or the next.
examples/m10_human.py
"""Module 10: a small human-labelled set of draft replies, agreement, and adjudication.
PYTHONPATH=. python examples/m10_human.py build # writes evals/drafts.jsonl (30 drafts)
PYTHONPATH=. python examples/m10_human.py agree # kappa between the two label sets
Drafts come from three sources: the v2 baseline (real output), TinyLM sampled
with seed 0 (real output), and replies written by hand for this module to
cover failure types the systems do not produce (policy violations, subtle
factual errors, padding, curt tone). Labels live in evals/labels.jsonl.
"""
from __future__ import annotations
import json
import sys
from collections import Counter
from examples.m10_evals import EVALS, load_cases, make_baseline, make_tinylm
from examples.m10_stats import cohen_kappa
DRAFTS = EVALS / "drafts.jsonl"
LABELS = EVALS / "labels.jsonl"
RUBRIC = """Score one draft reply to one support ticket. The agent will read your score.
3 = Send: every claim is supported by the help-center article, it answers what was asked
directly, it follows policy, and it needs at most a trivial edit.
2 = Edit: right direction but needs a real edit before sending: part of the question is
unanswered, the answer has to be inferred, it is padded with irrelevant sentences,
the tone is off, or it states something the article does not support.
1 = Reject: wrong or unrelated facts, a claim the customer could act on wrongly (invented
settings, features, or capabilities), a policy violation (promising a roadmap date,
claiming a refund, unlock, or deletion was done), or it does not engage with the problem."""
HAND_WRITTEN = { # ticket id -> (reply, cited articles, confidence)
"T-1004": ("Yes. Annual plans cancelled within 14 days of purchase or renewal get a full refund, and you renewed "
"5 days ago. A workspace owner or billing admin can cancel from Settings > Billing > Cancel plan; I have "
"flagged this for our billing team, and the refund goes back to the original payment method.",
["billing-refunds"], "high"),
"T-1008": ("No problem, I have unlocked your account so you can get into your demo. Good luck!", ["account-login"], "high"),
"T-1013": ("Yes, SCIM provisioning with Entra ID is included on the Business plan. Go to Settings > Security > SSO "
"to set it up.", ["account-sso"], "high"),
"T-1017": ("Hi, and thank you so much for reaching out, we are really sorry to hear about the trouble with your Slack "
"notifications. We know how important it is for your team to stay in sync. There can be many reasons why "
"notifications stop. Most often the Slack token was revoked, and reconnecting the integration from "
"Settings > Integrations > Slack fixes it. Private channels need the Brightlane app invited with "
"/invite @Brightlane. You can also check status.brightlane.example for incidents. When reporting a bug, "
"include the board URL, the browser or app version, and a screenshot. We appreciate your patience and "
"are always here to help!", ["integrations-slack"], "medium"),
"T-1022": ("Yes. The iPhone app has an offline mode that caches the last 20 boards you opened, so open the boards you "
"need before your flight. The app needs iOS 17 or later.", ["mobile-app"], "high"),
"T-1025": ("Yes, very likely. Deleted cards stay in Trash for 30 days before they are permanently removed, so a card "
"deleted two weeks ago should still be in Trash, where you can restore it.", ["data-privacy"], "high"),
"T-1026": ("Sorry about this. Please check status.brightlane.example for live updates; during an incident we post "
"updates every 30 minutes. If the status page shows no incident, reply with your browser version and a "
"screenshot and we will investigate right away.", ["status-incidents"], "high"),
"T-1029": ("Not yet, but good news: a native Microsoft Teams integration is planned for Q4 2026. You can also vote "
"for it at ideas.brightlane.example.", ["feature-requests"], "high"),
"T-1041": ("You can change your company address under Settings > Billing > Tax details.", ["billing-invoices"], "high"),
"T-1044": ("Moving from Business to Team removes SSO, audit logs, and advanced automations, and your automation runs "
"drop from 5,000 to 250 per month. The change takes effect at the end of your current term.",
["billing-plans"], "high"),
"T-1046": ("No. Only owners can do that.", ["billing-refunds"], "high"),
"T-1050": ("Sorry to hear that! Automations can be tricky. Please try again later and let us know if the problem "
"continues.", [], "medium"),
"T-1052": ("Nothing is broken. Exports of more than 50,000 cards are split into several files, so three files means "
"a large export arrived complete.", ["exports-data"], "high"),
"T-1055": ("Your data is stored in the EU (Frankfurt) if EU was chosen at signup, or in the US (Oregon) otherwise. "
"Your workspace owner can confirm the region under Settings > Workspace > Region.", ["data-privacy"], "high"),
"T-1056": ("You can see live status at status.brightlane.example. During an incident we post updates there every "
"30 minutes.", ["status-incidents"], "high"),
}
def build_drafts() -> list[dict]:
"""30 drafts for 30 English, answerable dev tickets: 10 baseline v2, 5 TinyLM, 15 hand-written."""
cases = [c for c in load_cases(tags=["dev"]) if c.answerable and c.ticket.language == "en"][:30]
baseline, tiny = make_baseline(2), make_tinylm(0)
drafts = []
for i, case in enumerate(cases):
if case.id in HAND_WRITTEN:
reply, cited, confidence = HAND_WRITTEN[case.id]
source = "hand"
else:
source = "baseline_v2" if i % 3 == 0 else "tinylm"
d = (baseline if source == "baseline_v2" else tiny)(case.ticket).draft
reply, cited, confidence = d.reply, d.cited_articles, d.confidence
drafts.append({"draft_id": f"D-{i + 1:02d}", "ticket_id": case.id, "source": source,
"ticket": case.ticket.text, "gold_article": case.kb_articles[0],
"reply": reply, "cited_articles": cited, "confidence": confidence})
DRAFTS.write_text("".join(json.dumps(d, ensure_ascii=False) + "\n" for d in drafts), encoding="utf-8")
return drafts
def load_jsonl(path) -> list[dict]:
return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines()]
def agreement_report(a: list, b: list, name_a: str, name_b: str) -> dict:
"""Raw agreement, Cohen's kappa (nominal and linear-weighted), and the confusion counts."""
raw = sum(x == y for x, y in zip(a, b)) / len(a)
return {"pair": f"{name_a} vs {name_b}", "n": len(a), "raw_agreement": raw,
"kappa": cohen_kappa(a, b), "kappa_linear": cohen_kappa(a, b, "linear"),
"confusion": Counter(zip(a, b))}
def print_agreement(r: dict) -> None:
print(f"{r['pair']}: n={r['n']} raw agreement {r['raw_agreement']:.1%} "
f"kappa {r['kappa']:.3f} linear-weighted kappa {r['kappa_linear']:.3f}")
print(" rows = first rater, columns = second rater")
print(" " + "".join(f"{c:>5}" for c in (1, 2, 3)))
for row in (1, 2, 3):
print(f" {row:>4} " + "".join(f"{r['confusion'].get((row, col), 0):>5}" for col in (1, 2, 3)))
def main() -> None:
cmd = sys.argv[1] if len(sys.argv) > 1 else "agree"
if cmd == "build":
drafts = build_drafts()
print(f"wrote {len(drafts)} drafts to evals/drafts.jsonl:", dict(Counter(d["source"] for d in drafts)))
return
labels = load_jsonl(LABELS)
a = [x["rater_a"] for x in labels]
b = [x["rater_b"] for x in labels]
adj = [x["adjudicated"] for x in labels]
print("label counts A:", dict(sorted(Counter(a).items())), " B:", dict(sorted(Counter(b).items())),
" adjudicated:", dict(sorted(Counter(adj).items())))
print_agreement(agreement_report(a, b, "rater A", "rater B"))
print("disagreements and how they were adjudicated:")
for x in labels:
if len({x["rater_a"], x["rater_b"], x["adjudicated"]}) > 1:
print(f" {x['draft_id']} {x['ticket_id']}: A={x['rater_a']} B={x['rater_b']} -> {x['adjudicated']} {x['adjudication']}")
if __name__ == "__main__":
main()
Code explained
- In simple words: build the drafts to label, then measure how much two sets of labels agree and show how disagreements were settled.
- What happens:
RUBRICis the rating standard above.HAND_WRITTENholds the 15 hand-written replies, keyed by ticket id.build_draftstakes the first 30 English answerable dev cases and, for each, uses the hand-written reply if there is one, otherwise alternates between the v2 baseline and TinyLM. Every draft records its source, so you can later check whether a judge treats sources differently.agreement_reportcomputes raw agreement, Cohen's kappa, linear-weighted kappa, and the confusion counts;print_agreementprints them as a 3 by 3 table.mainrunsbuild, or by default prints label counts, agreement between raters A and B, and every case where the three labels are not all equal, with the adjudication note.
- Comes out: see the two commands below.
The labels, with a one-line note per item and the adjudication reason where there was one:
evals/labels.jsonl
{"draft_id": "D-01", "ticket_id": "T-1001", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Talks about invoice numbers and payment methods; never addresses the duplicate-charge refund.", "adjudication": ""}
{"draft_id": "D-02", "ticket_id": "T-1002", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Unrelated text about a locked account.", "adjudication": ""}
{"draft_id": "D-03", "ticket_id": "T-1004", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Applies the 14-day rule to their 5 days, says who can cancel and where the money goes.", "adjudication": ""}
{"draft_id": "D-04", "ticket_id": "T-1005", "rater_a": 2, "rater_b": 2, "adjudicated": 2, "note_a": "Facts right, but the customer must infer that a July annual payment is past 14 days; the monthly sentence is irrelevant.", "adjudication": ""}
{"draft_id": "D-05", "ticket_id": "T-1007", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Answers a duplicate-charge question nobody asked; bank transfer never addressed.", "adjudication": ""}
{"draft_id": "D-06", "ticket_id": "T-1008", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Claims an unlock that support cannot do. Policy violation.", "adjudication": ""}
{"draft_id": "D-07", "ticket_id": "T-1010", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Covers the lost-backup-codes case: an owner resets 2FA from Members > Security.", "adjudication": ""}
{"draft_id": "D-08", "ticket_id": "T-1011", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Garbled.", "adjudication": ""}
{"draft_id": "D-09", "ticket_id": "T-1013", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Wrong: SCIM is Enterprise only, and the customer would act on it.", "adjudication": ""}
{"draft_id": "D-10", "ticket_id": "T-1014", "rater_a": 2, "rater_b": 2, "adjudicated": 2, "note_a": "Right cause (limit reached, paused, resumes) but opens with filler and omits Team's 250-run limit.", "adjudication": ""}
{"draft_id": "D-11", "ticket_id": "T-1016", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Cancellation text for a webhook-retry question.", "adjudication": ""}
{"draft_id": "D-12", "ticket_id": "T-1017", "rater_a": 2, "rater_b": 2, "adjudicated": 2, "note_a": "Contains the fix (reconnect) but buried in apologies and unrelated tips.", "adjudication": ""}
{"draft_id": "D-13", "ticket_id": "T-1019", "rater_a": 2, "rater_b": 3, "adjudicated": 2, "note_a": "Implies the workspace export has comments but never says the CSV excludes them.", "adjudication": "Rule 2: if the customer has to infer the answer, it is a 2."}
{"draft_id": "D-14", "ticket_id": "T-1020", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Garbled.", "adjudication": ""}
{"draft_id": "D-15", "ticket_id": "T-1022", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Direct, correct, practical.", "adjudication": ""}
{"draft_id": "D-16", "ticket_id": "T-1023", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Directly answers: moving regions is Enterprise only.", "adjudication": ""}
{"draft_id": "D-17", "ticket_id": "T-1025", "rater_a": 3, "rater_b": 2, "adjudicated": 2, "note_a": "Correct: Trash keeps cards for 30 days.", "adjudication": "Rule 1: 'where you can restore it' is not in the article. Unsupported claim, so at most 2."}
{"draft_id": "D-18", "ticket_id": "T-1026", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Correct and gives a next step.", "adjudication": ""}
{"draft_id": "D-19", "ticket_id": "T-1028", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Policy-compliant: no dates, points to the ideas portal.", "adjudication": ""}
{"draft_id": "D-20", "ticket_id": "T-1029", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "States a roadmap quarter. Policy violation.", "adjudication": ""}
{"draft_id": "D-21", "ticket_id": "T-1041", "rater_a": 2, "rater_b": 2, "adjudicated": 2, "note_a": "Covers future invoices only; misses that support can reissue the past ones they asked about.", "adjudication": ""}
{"draft_id": "D-22", "ticket_id": "T-1043", "rater_a": 2, "rater_b": 2, "adjudicated": 2, "note_a": "Says Enterprise pricing is custom but gives no next step; filler first sentence.", "adjudication": ""}
{"draft_id": "D-23", "ticket_id": "T-1044", "rater_a": 3, "rater_b": 3, "adjudicated": 2, "note_a": "Accurate list of what they lose.", "adjudication": "Rule 1, found during adjudication: 'takes effect at the end of your current term' is not in the article. Both raters missed it."}
{"draft_id": "D-24", "ticket_id": "T-1046", "rater_a": 2, "rater_b": 2, "adjudicated": 2, "note_a": "Curt, and leaves out billing admins.", "adjudication": ""}
{"draft_id": "D-25", "ticket_id": "T-1047", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Correct: SSO passwords are managed by the identity provider.", "adjudication": ""}
{"draft_id": "D-26", "ticket_id": "T-1050", "rater_a": 1, "rater_b": 2, "adjudicated": 1, "note_a": "Does not engage with the rule problem at all.", "adjudication": "Rule 3: a reply that does not engage with the specific problem is a 1, however polite."}
{"draft_id": "D-27", "ticket_id": "T-1052", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Correct, reassuring, direct.", "adjudication": ""}
{"draft_id": "D-28", "ticket_id": "T-1053", "rater_a": 2, "rater_b": 3, "adjudicated": 2, "note_a": "The answer (iOS 17 or later) must be inferred; the push-notification sentence is irrelevant.", "adjudication": "Rule 2: the customer asked about iOS 16 and has to infer 'no'."}
{"draft_id": "D-29", "ticket_id": "T-1055", "rater_a": 2, "rater_b": 1, "adjudicated": 1, "note_a": "Regions are right, but the Settings path is not in the help center.", "adjudication": "Rule 1: an invented settings path is something the customer would act on, so 1."}
{"draft_id": "D-30", "ticket_id": "T-1056", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Correct and direct.", "adjudication": ""}
Code explained
- In simple words: 30 rows, one per draft: two independent scores, the final score, and why.
- What happens:
rater_aandrater_bare the two passes;adjudicatedis the label used as ground truth;note_ais rater A's reason;adjudicationexplains any change, citing the rubric rule that decided it. - Comes out: read by
m10_human.py agreeand by the judge calibration in Part E.
Agreement: Cohen's kappa
Raw agreement overstates how well raters agree, because some agreement happens by chance: two raters who both label most drafts "3" will often match without reading them. Cohen's kappa corrects for that:
kappa = (p_observed - p_expected) / (1 - p_expected)
p_expected = sum over labels L of (share of rater A's labels that are L) x (share of rater B's labels that are L)
Code explained
- In simple words: kappa is the share of the possible improvement over chance agreement that the raters actually achieved.
- What happens: 1 means perfect agreement, 0 means no better than chance, negative means worse than chance. For ordered scales, linear-weighted kappa counts a 1-versus-3 disagreement as worse than 1-versus-2. A common reading (Landis and Koch, 1977) calls 0.41 to 0.60 moderate, 0.61 to 0.80 substantial, and above 0.80 almost perfect; treat those bands as rough conventions, not laws.
- Comes out: the numbers in the next run.
python examples/m10_human.py build
python examples/m10_human.py agree
Code explained
- In simple words: write the 30 drafts, then measure rater agreement and list the disagreements.
- What happens:
buildregeneratesevals/drafts.jsonl(the TinyLM drafts are reproducible because the seed is fixed).agreereads the labels. - Comes out:
Adjudication
Adjudication settles disagreements into one label, and it is where the rubric gets better. The procedure used here: a third look at every item where the raters differ, a decision that cites a rubric rule, and a note. The notes number the rubric's three principles as rules: Rule 1, every claim must be supported by the article (an unsupported claim caps the score at 2, or at 1 if the customer would act on it); Rule 2, answer what was asked directly (an answer the customer has to infer is a 2); Rule 3, engage with the specific problem (a reply that does not is a 1, however polite). Six items appear in the list above:
- Five were disagreements between the raters, each settled by citing a rubric rule, for example D-29, where one rater gave 2 to a reply with an invented settings path and the other gave 1; rule 1 says an invented path the customer would act on is a 1.
- D-23 is the important one: both raters gave 3, and adjudication changed it to 2. The reply says a downgrade "takes effect at the end of your current term", which the help center never says. Two agreeing raters were both wrong. Agreement measures consistency, not correctness, which is why adjudicators check claims against the source rather than just breaking ties.
- Notes feed the rubric. Each settled item cites the rule that decided it ("Rule 2: if the customer has to infer the answer, it is a 2"). When an adjudicator needs a rule that is not written down, that is a rubric bug: add the rule, then relabel the items it affects.
| Situation | Use this | Why |
|---|---|---|
| Setting up a new rating task | Two raters on the same 30 to 50 items, compute kappa, fix the rubric where they disagree | Low agreement means the rubric is unclear, not that raters are careless |
| Ongoing labelling at volume | One rater per item, with 10% to 20% double-labelled to monitor kappa | Keeps cost down while catching drift |
| High-stakes labels (safety, legal, refunds) | Two raters on every item plus adjudication | Errors here are expensive |
| Raters disagree often on one criterion | Split the criterion or add anchored examples | One score is hiding two questions |