CourseLarge Language Models · Module 10: Evaluation · part 54 of 80
Part 54 · Module 10: Evaluation

Part D: Human evaluation

17 min read·22 Sept 2026

When humans are unavoidable

Deterministic checks cannot tell whether a reply is helpful, whether its tone fits an angry customer, or whether a correct-looking sentence is subtly wrong for this customer. Someone has to judge. Humans are unavoidable in four places:

  • Defining the standard. A rubric is a human decision. No model can tell you what Maya considers "good enough to send".
  • Calibrating automated judges. An LLM judge is only as trustworthy as its measured agreement with human labels (Part E).
  • High-stakes or novel failures. Legal, security, and money questions, and any failure type you have not seen before.
  • Periodic audits. Judges and checks drift as the product changes; a monthly human sample catches it.

Human time is expensive, so spend it where it buys the most: labelling a few dozen items well beats skimming hundreds.

Rating design

Good human labels come from a clear question, a small scale with written anchors, and examples. For Brightlane drafts the question is the one agents actually face: can I send this?

Design choiceBrightlane choiceWhy
Scale3 points: 1 Reject, 2 Edit, 3 SendEach point is a different action; a 1 to 10 scale invites disagreement between 6 and 7 that means nothing
AnchorsWritten rules for each point, including tie-breakersRaters apply rules, not moods
UnitOne draft for one ticket, with the help-center article visibleRaters must check claims against the source
NotesA one-line reason per labelReasons make adjudication fast and turn into judge rubric text
BlindRaters do not see which system wrote the draftPrevents "the new model is probably better" bias

The rubric lives in examples/m10_human.py so the judge prompt in Part E uses the exact same words:

text
3 = Send: every claim is supported by the help-center article, it answers what was asked
    directly, it follows policy, and it needs at most a trivial edit.
2 = Edit: right direction but needs a real edit before sending: part of the question is
    unanswered, the answer has to be inferred, it is padded with irrelevant sentences,
    the tone is off, or it states something the article does not support.
1 = Reject: wrong or unrelated facts, a claim the customer could act on wrongly (invented
    settings, features, or capabilities), a policy violation (promising a roadmap date,
    claiming a refund, unlock, or deletion was done), or it does not engage with the problem.

Code explained

  • In simple words: three scores, each tied to what the agent does next, with the tricky cases spelled out.
  • What happens: the rubric separates "unsupported but harmless" (at most 2) from "unsupported and the customer would act on it" (1), because an invented settings path sends a customer somewhere that does not exist. Adjudication below added no new rules; it applied these.
  • Comes out: the text that raters read and that pointwise_messages embeds in the judge prompt.

The labelled set, and what it really is

examples/m10_human.py build produces 30 drafts for the first 30 English, answerable dev tickets: 10 from the v2 baseline (real output), 5 sampled from TinyLM with seed 0 (real output), and 15 written by hand for this module to cover failures our systems do not produce: policy violations, confident wrong facts, padding, curt tone, invented settings paths, and genuinely good replies.

Be clear about the labels. evals/labels.jsonl holds two label sets, "rater A" and "rater B", plus an adjudicated label and notes. They were written by the author of this module, not by Brightlane support agents, in two separate labelling passes with the rubric above. They are a worked example that makes every computation below real and checkable, not evidence about how real raters behave: two passes by one person agree more than two different people usually would. Replace them with labels from two of your own agents before you trust any number in this part or the next.

examples/m10_human.py

python
"""Module 10: a small human-labelled set of draft replies, agreement, and adjudication.

    PYTHONPATH=. python examples/m10_human.py build     # writes evals/drafts.jsonl (30 drafts)
    PYTHONPATH=. python examples/m10_human.py agree     # kappa between the two label sets

Drafts come from three sources: the v2 baseline (real output), TinyLM sampled
with seed 0 (real output), and replies written by hand for this module to
cover failure types the systems do not produce (policy violations, subtle
factual errors, padding, curt tone). Labels live in evals/labels.jsonl.
"""
from __future__ import annotations

import json
import sys
from collections import Counter

from examples.m10_evals import EVALS, load_cases, make_baseline, make_tinylm
from examples.m10_stats import cohen_kappa

DRAFTS = EVALS / "drafts.jsonl"
LABELS = EVALS / "labels.jsonl"

RUBRIC = """Score one draft reply to one support ticket. The agent will read your score.
3 = Send: every claim is supported by the help-center article, it answers what was asked
    directly, it follows policy, and it needs at most a trivial edit.
2 = Edit: right direction but needs a real edit before sending: part of the question is
    unanswered, the answer has to be inferred, it is padded with irrelevant sentences,
    the tone is off, or it states something the article does not support.
1 = Reject: wrong or unrelated facts, a claim the customer could act on wrongly (invented
    settings, features, or capabilities), a policy violation (promising a roadmap date,
    claiming a refund, unlock, or deletion was done), or it does not engage with the problem."""

HAND_WRITTEN = {  # ticket id -> (reply, cited articles, confidence)
    "T-1004": ("Yes. Annual plans cancelled within 14 days of purchase or renewal get a full refund, and you renewed "
               "5 days ago. A workspace owner or billing admin can cancel from Settings > Billing > Cancel plan; I have "
               "flagged this for our billing team, and the refund goes back to the original payment method.",
               ["billing-refunds"], "high"),
    "T-1008": ("No problem, I have unlocked your account so you can get into your demo. Good luck!", ["account-login"], "high"),
    "T-1013": ("Yes, SCIM provisioning with Entra ID is included on the Business plan. Go to Settings > Security > SSO "
               "to set it up.", ["account-sso"], "high"),
    "T-1017": ("Hi, and thank you so much for reaching out, we are really sorry to hear about the trouble with your Slack "
               "notifications. We know how important it is for your team to stay in sync. There can be many reasons why "
               "notifications stop. Most often the Slack token was revoked, and reconnecting the integration from "
               "Settings > Integrations > Slack fixes it. Private channels need the Brightlane app invited with "
               "/invite @Brightlane. You can also check status.brightlane.example for incidents. When reporting a bug, "
               "include the board URL, the browser or app version, and a screenshot. We appreciate your patience and "
               "are always here to help!", ["integrations-slack"], "medium"),
    "T-1022": ("Yes. The iPhone app has an offline mode that caches the last 20 boards you opened, so open the boards you "
               "need before your flight. The app needs iOS 17 or later.", ["mobile-app"], "high"),
    "T-1025": ("Yes, very likely. Deleted cards stay in Trash for 30 days before they are permanently removed, so a card "
               "deleted two weeks ago should still be in Trash, where you can restore it.", ["data-privacy"], "high"),
    "T-1026": ("Sorry about this. Please check status.brightlane.example for live updates; during an incident we post "
               "updates every 30 minutes. If the status page shows no incident, reply with your browser version and a "
               "screenshot and we will investigate right away.", ["status-incidents"], "high"),
    "T-1029": ("Not yet, but good news: a native Microsoft Teams integration is planned for Q4 2026. You can also vote "
               "for it at ideas.brightlane.example.", ["feature-requests"], "high"),
    "T-1041": ("You can change your company address under Settings > Billing > Tax details.", ["billing-invoices"], "high"),
    "T-1044": ("Moving from Business to Team removes SSO, audit logs, and advanced automations, and your automation runs "
               "drop from 5,000 to 250 per month. The change takes effect at the end of your current term.",
               ["billing-plans"], "high"),
    "T-1046": ("No. Only owners can do that.", ["billing-refunds"], "high"),
    "T-1050": ("Sorry to hear that! Automations can be tricky. Please try again later and let us know if the problem "
               "continues.", [], "medium"),
    "T-1052": ("Nothing is broken. Exports of more than 50,000 cards are split into several files, so three files means "
               "a large export arrived complete.", ["exports-data"], "high"),
    "T-1055": ("Your data is stored in the EU (Frankfurt) if EU was chosen at signup, or in the US (Oregon) otherwise. "
               "Your workspace owner can confirm the region under Settings > Workspace > Region.", ["data-privacy"], "high"),
    "T-1056": ("You can see live status at status.brightlane.example. During an incident we post updates there every "
               "30 minutes.", ["status-incidents"], "high"),
}


def build_drafts() -> list[dict]:
    """30 drafts for 30 English, answerable dev tickets: 10 baseline v2, 5 TinyLM, 15 hand-written."""
    cases = [c for c in load_cases(tags=["dev"]) if c.answerable and c.ticket.language == "en"][:30]
    baseline, tiny = make_baseline(2), make_tinylm(0)
    drafts = []
    for i, case in enumerate(cases):
        if case.id in HAND_WRITTEN:
            reply, cited, confidence = HAND_WRITTEN[case.id]
            source = "hand"
        else:
            source = "baseline_v2" if i % 3 == 0 else "tinylm"
            d = (baseline if source == "baseline_v2" else tiny)(case.ticket).draft
            reply, cited, confidence = d.reply, d.cited_articles, d.confidence
        drafts.append({"draft_id": f"D-{i + 1:02d}", "ticket_id": case.id, "source": source,
                       "ticket": case.ticket.text, "gold_article": case.kb_articles[0],
                       "reply": reply, "cited_articles": cited, "confidence": confidence})
    DRAFTS.write_text("".join(json.dumps(d, ensure_ascii=False) + "\n" for d in drafts), encoding="utf-8")
    return drafts


def load_jsonl(path) -> list[dict]:
    return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines()]


def agreement_report(a: list, b: list, name_a: str, name_b: str) -> dict:
    """Raw agreement, Cohen's kappa (nominal and linear-weighted), and the confusion counts."""
    raw = sum(x == y for x, y in zip(a, b)) / len(a)
    return {"pair": f"{name_a} vs {name_b}", "n": len(a), "raw_agreement": raw,
            "kappa": cohen_kappa(a, b), "kappa_linear": cohen_kappa(a, b, "linear"),
            "confusion": Counter(zip(a, b))}


def print_agreement(r: dict) -> None:
    print(f"{r['pair']}: n={r['n']} raw agreement {r['raw_agreement']:.1%}  "
          f"kappa {r['kappa']:.3f}  linear-weighted kappa {r['kappa_linear']:.3f}")
    print("   rows = first rater, columns = second rater")
    print("        " + "".join(f"{c:>5}" for c in (1, 2, 3)))
    for row in (1, 2, 3):
        print(f"   {row:>4} " + "".join(f"{r['confusion'].get((row, col), 0):>5}" for col in (1, 2, 3)))


def main() -> None:
    cmd = sys.argv[1] if len(sys.argv) > 1 else "agree"
    if cmd == "build":
        drafts = build_drafts()
        print(f"wrote {len(drafts)} drafts to evals/drafts.jsonl:", dict(Counter(d["source"] for d in drafts)))
        return
    labels = load_jsonl(LABELS)
    a = [x["rater_a"] for x in labels]
    b = [x["rater_b"] for x in labels]
    adj = [x["adjudicated"] for x in labels]
    print("label counts  A:", dict(sorted(Counter(a).items())), " B:", dict(sorted(Counter(b).items())),
          " adjudicated:", dict(sorted(Counter(adj).items())))
    print_agreement(agreement_report(a, b, "rater A", "rater B"))
    print("disagreements and how they were adjudicated:")
    for x in labels:
        if len({x["rater_a"], x["rater_b"], x["adjudicated"]}) > 1:
            print(f"   {x['draft_id']} {x['ticket_id']}: A={x['rater_a']} B={x['rater_b']} -> {x['adjudicated']}  {x['adjudication']}")


if __name__ == "__main__":
    main()

Code explained

  • In simple words: build the drafts to label, then measure how much two sets of labels agree and show how disagreements were settled.
  • What happens:
    • RUBRIC is the rating standard above. HAND_WRITTEN holds the 15 hand-written replies, keyed by ticket id.
    • build_drafts takes the first 30 English answerable dev cases and, for each, uses the hand-written reply if there is one, otherwise alternates between the v2 baseline and TinyLM. Every draft records its source, so you can later check whether a judge treats sources differently.
    • agreement_report computes raw agreement, Cohen's kappa, linear-weighted kappa, and the confusion counts; print_agreement prints them as a 3 by 3 table.
    • main runs build, or by default prints label counts, agreement between raters A and B, and every case where the three labels are not all equal, with the adjudication note.
  • Comes out: see the two commands below.

The labels, with a one-line note per item and the adjudication reason where there was one:

evals/labels.jsonl

json
{"draft_id": "D-01", "ticket_id": "T-1001", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Talks about invoice numbers and payment methods; never addresses the duplicate-charge refund.", "adjudication": ""}
{"draft_id": "D-02", "ticket_id": "T-1002", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Unrelated text about a locked account.", "adjudication": ""}
{"draft_id": "D-03", "ticket_id": "T-1004", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Applies the 14-day rule to their 5 days, says who can cancel and where the money goes.", "adjudication": ""}
{"draft_id": "D-04", "ticket_id": "T-1005", "rater_a": 2, "rater_b": 2, "adjudicated": 2, "note_a": "Facts right, but the customer must infer that a July annual payment is past 14 days; the monthly sentence is irrelevant.", "adjudication": ""}
{"draft_id": "D-05", "ticket_id": "T-1007", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Answers a duplicate-charge question nobody asked; bank transfer never addressed.", "adjudication": ""}
{"draft_id": "D-06", "ticket_id": "T-1008", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Claims an unlock that support cannot do. Policy violation.", "adjudication": ""}
{"draft_id": "D-07", "ticket_id": "T-1010", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Covers the lost-backup-codes case: an owner resets 2FA from Members > Security.", "adjudication": ""}
{"draft_id": "D-08", "ticket_id": "T-1011", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Garbled.", "adjudication": ""}
{"draft_id": "D-09", "ticket_id": "T-1013", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Wrong: SCIM is Enterprise only, and the customer would act on it.", "adjudication": ""}
{"draft_id": "D-10", "ticket_id": "T-1014", "rater_a": 2, "rater_b": 2, "adjudicated": 2, "note_a": "Right cause (limit reached, paused, resumes) but opens with filler and omits Team's 250-run limit.", "adjudication": ""}
{"draft_id": "D-11", "ticket_id": "T-1016", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Cancellation text for a webhook-retry question.", "adjudication": ""}
{"draft_id": "D-12", "ticket_id": "T-1017", "rater_a": 2, "rater_b": 2, "adjudicated": 2, "note_a": "Contains the fix (reconnect) but buried in apologies and unrelated tips.", "adjudication": ""}
{"draft_id": "D-13", "ticket_id": "T-1019", "rater_a": 2, "rater_b": 3, "adjudicated": 2, "note_a": "Implies the workspace export has comments but never says the CSV excludes them.", "adjudication": "Rule 2: if the customer has to infer the answer, it is a 2."}
{"draft_id": "D-14", "ticket_id": "T-1020", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "Garbled.", "adjudication": ""}
{"draft_id": "D-15", "ticket_id": "T-1022", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Direct, correct, practical.", "adjudication": ""}
{"draft_id": "D-16", "ticket_id": "T-1023", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Directly answers: moving regions is Enterprise only.", "adjudication": ""}
{"draft_id": "D-17", "ticket_id": "T-1025", "rater_a": 3, "rater_b": 2, "adjudicated": 2, "note_a": "Correct: Trash keeps cards for 30 days.", "adjudication": "Rule 1: 'where you can restore it' is not in the article. Unsupported claim, so at most 2."}
{"draft_id": "D-18", "ticket_id": "T-1026", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Correct and gives a next step.", "adjudication": ""}
{"draft_id": "D-19", "ticket_id": "T-1028", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Policy-compliant: no dates, points to the ideas portal.", "adjudication": ""}
{"draft_id": "D-20", "ticket_id": "T-1029", "rater_a": 1, "rater_b": 1, "adjudicated": 1, "note_a": "States a roadmap quarter. Policy violation.", "adjudication": ""}
{"draft_id": "D-21", "ticket_id": "T-1041", "rater_a": 2, "rater_b": 2, "adjudicated": 2, "note_a": "Covers future invoices only; misses that support can reissue the past ones they asked about.", "adjudication": ""}
{"draft_id": "D-22", "ticket_id": "T-1043", "rater_a": 2, "rater_b": 2, "adjudicated": 2, "note_a": "Says Enterprise pricing is custom but gives no next step; filler first sentence.", "adjudication": ""}
{"draft_id": "D-23", "ticket_id": "T-1044", "rater_a": 3, "rater_b": 3, "adjudicated": 2, "note_a": "Accurate list of what they lose.", "adjudication": "Rule 1, found during adjudication: 'takes effect at the end of your current term' is not in the article. Both raters missed it."}
{"draft_id": "D-24", "ticket_id": "T-1046", "rater_a": 2, "rater_b": 2, "adjudicated": 2, "note_a": "Curt, and leaves out billing admins.", "adjudication": ""}
{"draft_id": "D-25", "ticket_id": "T-1047", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Correct: SSO passwords are managed by the identity provider.", "adjudication": ""}
{"draft_id": "D-26", "ticket_id": "T-1050", "rater_a": 1, "rater_b": 2, "adjudicated": 1, "note_a": "Does not engage with the rule problem at all.", "adjudication": "Rule 3: a reply that does not engage with the specific problem is a 1, however polite."}
{"draft_id": "D-27", "ticket_id": "T-1052", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Correct, reassuring, direct.", "adjudication": ""}
{"draft_id": "D-28", "ticket_id": "T-1053", "rater_a": 2, "rater_b": 3, "adjudicated": 2, "note_a": "The answer (iOS 17 or later) must be inferred; the push-notification sentence is irrelevant.", "adjudication": "Rule 2: the customer asked about iOS 16 and has to infer 'no'."}
{"draft_id": "D-29", "ticket_id": "T-1055", "rater_a": 2, "rater_b": 1, "adjudicated": 1, "note_a": "Regions are right, but the Settings path is not in the help center.", "adjudication": "Rule 1: an invented settings path is something the customer would act on, so 1."}
{"draft_id": "D-30", "ticket_id": "T-1056", "rater_a": 3, "rater_b": 3, "adjudicated": 3, "note_a": "Correct and direct.", "adjudication": ""}

Code explained

  • In simple words: 30 rows, one per draft: two independent scores, the final score, and why.
  • What happens: rater_a and rater_b are the two passes; adjudicated is the label used as ground truth; note_a is rater A's reason; adjudication explains any change, citing the rubric rule that decided it.
  • Comes out: read by m10_human.py agree and by the judge calibration in Part E.

Agreement: Cohen's kappa

Raw agreement overstates how well raters agree, because some agreement happens by chance: two raters who both label most drafts "3" will often match without reading them. Cohen's kappa corrects for that:

text
kappa = (p_observed - p_expected) / (1 - p_expected)
p_expected = sum over labels L of (share of rater A's labels that are L) x (share of rater B's labels that are L)

Code explained

  • In simple words: kappa is the share of the possible improvement over chance agreement that the raters actually achieved.
  • What happens: 1 means perfect agreement, 0 means no better than chance, negative means worse than chance. For ordered scales, linear-weighted kappa counts a 1-versus-3 disagreement as worse than 1-versus-2. A common reading (Landis and Koch, 1977) calls 0.41 to 0.60 moderate, 0.61 to 0.80 substantial, and above 0.80 almost perfect; treat those bands as rough conventions, not laws.
  • Comes out: the numbers in the next run.

bash
python examples/m10_human.py build
python examples/m10_human.py agree

Code explained

  • In simple words: write the 30 drafts, then measure rater agreement and list the disagreements.
  • What happens: build regenerates evals/drafts.jsonl (the TinyLM drafts are reproducible because the seed is fixed). agree reads the labels.
  • Comes out:

Adjudication

Adjudication settles disagreements into one label, and it is where the rubric gets better. The procedure used here: a third look at every item where the raters differ, a decision that cites a rubric rule, and a note. The notes number the rubric's three principles as rules: Rule 1, every claim must be supported by the article (an unsupported claim caps the score at 2, or at 1 if the customer would act on it); Rule 2, answer what was asked directly (an answer the customer has to infer is a 2); Rule 3, engage with the specific problem (a reply that does not is a 1, however polite). Six items appear in the list above:

  • Five were disagreements between the raters, each settled by citing a rubric rule, for example D-29, where one rater gave 2 to a reply with an invented settings path and the other gave 1; rule 1 says an invented path the customer would act on is a 1.
  • D-23 is the important one: both raters gave 3, and adjudication changed it to 2. The reply says a downgrade "takes effect at the end of your current term", which the help center never says. Two agreeing raters were both wrong. Agreement measures consistency, not correctness, which is why adjudicators check claims against the source rather than just breaking ties.
  • Notes feed the rubric. Each settled item cites the rule that decided it ("Rule 2: if the customer has to infer the answer, it is a 2"). When an adjudicator needs a rule that is not written down, that is a rubric bug: add the rule, then relabel the items it affects.
SituationUse thisWhy
Setting up a new rating taskTwo raters on the same 30 to 50 items, compute kappa, fix the rubric where they disagreeLow agreement means the rubric is unclear, not that raters are careless
Ongoing labelling at volumeOne rater per item, with 10% to 20% double-labelled to monitor kappaKeeps cost down while catching drift
High-stakes labels (safety, legal, refunds)Two raters on every item plus adjudicationErrors here are expensive
Raters disagree often on one criterionSplit the criterion or add anchored examplesOne score is hiding two questions