Part E: LLM-as-judge
What a judge is, and how to write its rubric
An LLM-as-judge is a model prompted to grade another model's output against a rubric. It scales human judgment: once calibrated, it can score thousands of drafts for the price of a few cents each. It is also a model, with the same non-determinism and blind spots as the system it grades, so it is itself a system under test.
Rules that hold across providers:
- Use the human rubric verbatim. The judge prompt embeds
RUBRICfrom Part D, so humans and judge answer the same question. - Use a small, anchored scale. 1 to 3 with written anchors, or a binary pass or fail per criterion. Judges, like people, are inconsistent on 1 to 10 scales.
- Reasoning before the score. Ask for a short reasoning field first, then the score, in JSON. The score is then conditioned on the reasoning rather than the other way around, and the reasoning makes disagreements debuggable.
- One criterion per call when criteria conflict. "Is it grounded?" and "is the tone right?" are different questions; a single score blurs them.
- Give the judge the evidence. A grounding judge needs the help-center article; without it, it grades plausibility, not truth.
Pointwise, pairwise, and reference-based judging
| Situation | Use this | Why |
|---|---|---|
| Track quality over time, gate releases | Pointwise: score one output on the rubric | Scores are comparable across runs and systems |
| Choose between two prompts or models | Pairwise: which of two outputs is better | Relative judgments are easier and more consistent than absolute ones, but must be run in both orders |
| A known correct answer exists | Reference-based: compare the output to a reference answer | Turns an open judgment into "does it convey these facts?", which judges do more reliably |
| The check can be written as code | No judge | Deterministic checks are free and exact |
All three prompt builders, the stand-in judges, the bias harness, calibration, and cost math are in one file.
examples/m10_judge.py
"""Module 10: LLM-as-judge prompts, a deterministic stand-in judge, bias tests, calibration, and cost.
PYTHONPATH=. python examples/m10_judge.py demo # the three judge prompts and their size
PYTHONPATH=. python examples/m10_judge.py calibrate # stand-in judge vs human labels
PYTHONPATH=. python examples/m10_judge.py bias # position and verbosity harness
PYTHONPATH=. python examples/m10_judge.py cost # judge cost from real token counts
LLM_PROVIDER=groq PYTHONPATH=. python examples/m10_judge.py calibrate --live
LLM_PROVIDER=groq PYTHONPATH=. python examples/m10_judge.py bias --live
Without --live nothing here is a language model. The stand-in judge is a rule
(key facts, policy regexes, padding). The "biased" pairwise judge is a
ScriptedLLM that SIMULATES position and length bias so we can check that the
harness detects bias; it says nothing about how real models behave.
"""
from __future__ import annotations
import json
import re
import sys
from collections import Counter
from collections.abc import Callable
from examples.m10_evals import KEY_FACTS, ROADMAP_DATE, UNAUTHORIZED_ACTION, make_baseline
from examples.m10_human import DRAFTS, LABELS, RUBRIC, agreement_report, load_jsonl, print_agreement
from supportdesk.data import Ticket, get_article
from supportdesk.kb_search import tokenize
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages
# Judge prompts (pointwise, pairwise, reference-based) ------------------------------------
def pointwise_messages(ticket: str, article: str, reply: str) -> list[dict]:
return [
{"role": "system", "content": "You are a strict reviewer of support replies.\n" + RUBRIC +
'\nThink about each claim first. Return only JSON: {"reasoning": "<two sentences>", "score": 1|2|3}'},
{"role": "user", "content": f"Help-center article:\n{article}\n\nTicket:\n{ticket}\n\nDraft reply:\n{reply}"},
]
def pairwise_messages(ticket: str, article: str, reply_a: str, reply_b: str) -> list[dict]:
return [
{"role": "system", "content": "You compare two draft replies to the same support ticket.\n" + RUBRIC +
"\nPrefer the reply an agent could send with fewer edits. Length is not a virtue. The order of "
'the replies is random. Return only JSON: {"reasoning": "<two sentences>", "winner": "A"|"B"|"tie"}'},
{"role": "user", "content": f"Help-center article:\n{article}\n\nTicket:\n{ticket}\n\n"
f"Reply A:\n{reply_a}\n\nReply B:\n{reply_b}"},
]
def reference_messages(ticket: str, reference: str, reply: str) -> list[dict]:
return [
{"role": "system", "content": "You check a draft support reply against a reference answer written by a "
"senior agent. The draft may use different words. Score 3 if it conveys every fact in the reference and "
"nothing contradicting it, 2 if it misses a fact but contradicts nothing, 1 if it contradicts the reference "
'or adds a claim the reference does not support. Return only JSON: {"reasoning": "<two sentences>", "score": 1|2|3}'},
{"role": "user", "content": f"Ticket:\n{ticket}\n\nReference answer:\n{reference}\n\nDraft reply:\n{reply}"},
]
def parse_judge(text: str) -> dict:
"""Pull the first JSON object out of a judge reply; judges sometimes wrap it in prose."""
match = re.search(r"\{.*\}", text or "", re.S)
if not match:
raise ValueError(f"no JSON in judge output: {text[:80]!r}")
return json.loads(match.group(0))
def llm_pointwise_judge(chat_fn: Callable | None = None, **kw) -> Callable[[dict], int]:
"""A pointwise judge on a real model through supportdesk.llm.chat."""
if chat_fn is None:
from supportdesk.llm import chat as chat_fn
def judge(d: dict) -> int:
messages = pointwise_messages(d["ticket"], get_article(d["gold_article"]).body, d["reply"])
return int(parse_judge(chat_fn(messages, response_format={"type": "json_object"}, **kw).text)["score"])
return judge
# A deterministic stand-in judge ------------------------------------------------------------
def heuristic_score(ticket: str, reply: str, gold_article: str) -> int:
"""A rule-based, reference-based stand-in judge (not a model).
1 if it breaks policy or states no key fact of the right article; 2 if it is long
or padded (under half its sentences share a word with the ticket); else 3.
"""
if ROADMAP_DATE.search(reply) or UNAUTHORIZED_ACTION.search(reply):
return 1
if not re.search(KEY_FACTS[gold_article], reply, re.I):
return 1
ticket_words = {w for w in tokenize(ticket) if len(w) > 3}
sentences = [s for s in re.split(r"(?<=[.!?])\s+", reply) if s.strip()]
on_topic = sum(bool(ticket_words & set(tokenize(s))) for s in sentences) / len(sentences)
if len(reply.split()) > 80 or on_topic < 0.5:
return 2
return 3
def heuristic_judge(d: dict) -> int:
return heuristic_score(d["ticket"], d["reply"], d["gold_article"])
def calibrate(judge: Callable[[dict], int], name: str) -> dict:
"""Agreement between a judge and the adjudicated human labels, plus the 'send it' decision."""
drafts = {d["draft_id"]: d for d in load_jsonl(DRAFTS)}
labels = load_jsonl(LABELS)
human = [x["adjudicated"] for x in labels]
judged = [judge(drafts[x["draft_id"]]) for x in labels]
report = agreement_report(human, judged, "human (adjudicated)", name)
print_agreement(report)
tp = sum(h == 3 and j == 3 for h, j in zip(human, judged))
said_send = sum(j == 3 for j in judged)
should_send = sum(h == 3 for h in human)
print(f"'send as is' (score 3): judge said {said_send}, humans said {should_send}, both {tp} "
f"precision {tp / max(said_send, 1):.2f} recall {tp / max(should_send, 1):.2f}")
misses = [(x["draft_id"], h, j) for x, h, j in zip(labels, human, judged) if h != j]
print("disagreements (draft, human, judge):", misses)
return report
# Position and verbosity bias harness --------------------------------------------------------
PairJudge = Callable[[str, str, str, str], str] # (ticket, article, reply_a, reply_b) -> "A" | "B" | "tie"
def heuristic_pair_judge(ticket: str, article_id: str, a: str, b: str) -> str:
sa, sb = heuristic_score(ticket, a, article_id), heuristic_score(ticket, b, article_id)
return "A" if sa > sb else "B" if sb > sa else "tie"
def simulated_biased_judge() -> PairJudge:
"""SIMULATED bias for testing the harness: prefers a reply 60% longer, otherwise says 'A'."""
def responder(messages, kwargs):
text = messages[-1]["content"]
a = text.split("Reply A:\n", 1)[1].split("\n\nReply B:\n", 1)[0]
b = text.split("\n\nReply B:\n", 1)[1]
winner = "B" if len(b) > 1.6 * len(a) else "A"
return json.dumps({"reasoning": "simulated", "winner": winner})
fake = ScriptedLLM(responder=responder)
return llm_pair_judge(fake)
def llm_pair_judge(chat_fn: Callable | None = None, **kw) -> PairJudge:
if chat_fn is None:
from supportdesk.llm import chat as chat_fn
def judge(ticket: str, article_id: str, a: str, b: str) -> str:
messages = pairwise_messages(ticket, get_article(article_id).body, a, b)
return parse_judge(chat_fn(messages, response_format={"type": "json_object"}, **kw).text)["winner"]
return judge
def swap_test(judge: PairJudge, pairs: list[tuple[str, str, str, str]]) -> dict:
"""Judge every pair in both orders. A verdict counts only if it survives the swap."""
tally = Counter()
for ticket, article_id, x, y in pairs:
first = judge(ticket, article_id, x, y) # x shown as A
second = judge(ticket, article_id, y, x) # x shown as B
winner1 = {"A": "x", "B": "y"}.get(first, "tie")
winner2 = {"A": "y", "B": "x"}.get(second, "tie")
if winner1 == winner2:
tally["consistent"] += 1
tally[f"consistent_{winner1}"] += 1
else:
tally["inconsistent"] += 1
tally["first_position_won_both"] += first == "A" and second == "A"
tally["second_position_won_both"] += first == "B" and second == "B"
tally["pairs"] = len(pairs)
return dict(tally)
def pad(reply: str) -> str:
"""Make a reply longer without adding information (a verbosity-bias probe)."""
return (reply + " We truly appreciate your patience and understanding. Our team is always committed to "
"providing you with the best possible experience, and we are constantly working to improve. "
"Please do not hesitate to reach out if there is anything else at all we can help with.")
def verbosity_test(judge: PairJudge, pairs: list[tuple[str, str, str, str]]) -> dict:
"""Where the judge prefers x over y, pad y. A fair judge keeps preferring x."""
kept = flipped = 0
for ticket, article_id, x, y in pairs:
if judge(ticket, article_id, x, y) == "A" and judge(ticket, article_id, y, x) == "B":
padded = pad(y)
if judge(ticket, article_id, x, padded) == "A" and judge(ticket, article_id, padded, x) == "B":
kept += 1
else:
flipped += 1
return {"clear_preferences": kept + flipped, "kept_after_padding": kept, "flipped_or_tied": flipped}
def build_pairs() -> list[tuple[str, str, str, str]]:
"""Same-ticket pairs: each non-baseline draft against the v2 baseline's reply to that ticket."""
baseline = make_baseline(2)
pairs = []
for d in load_jsonl(DRAFTS):
if d["source"] == "baseline_v2":
continue
ticket_id, rest = d["ticket_id"], d["ticket"].split("\n\n", 1)
t = Ticket(ticket_id, rest[0].removeprefix("Subject: "), rest[1], "team", "en", "dev", {})
pairs.append((d["ticket"], d["gold_article"], d["reply"], baseline(t).draft.reply))
return pairs
# Judge cost ------------------------------------------------------------------------------
def judge_cost(output_tokens: int = 150, reasoning_tokens: int = 400) -> None:
"""Real input-token counts for this module's judge prompts; output and reasoning tokens are ASSUMPTIONS."""
drafts = load_jsonl(DRAFTS)
point = [count_messages(pointwise_messages(d["ticket"], get_article(d["gold_article"]).body, d["reply"]))
for d in drafts]
pair = [count_messages(pairwise_messages(t, get_article(a).body, x, y)) for t, a, x, y in build_pairs()]
mean_point, mean_pair = sum(point) / len(point), sum(pair) / len(pair)
print(f"input tokens per call: pointwise mean {mean_point:.0f} (n={len(point)}), pairwise mean {mean_pair:.0f} (n={len(pair)})")
print(f"assumed per call: {output_tokens} output tokens, plus {reasoning_tokens} reasoning tokens on reasoning models")
print(f"{'model':<26}{'pointwise':>11}{'pairwise x2':>13}{'CI/month':>10} (USD per 91-case suite; month = pairwise x2)")
for model in ["openai/gpt-oss-120b", "openai/gpt-oss-20b", "llama-3.3-70b-versatile", "gemini-3.5-flash",
"gemini-3.5-flash-lite"]:
extra = reasoning_tokens if model.startswith(("openai/gpt-oss", "gemini")) else 0
one_point = cost_usd(Usage(input_tokens=round(mean_point), output_tokens=output_tokens + extra), model)
one_pair = cost_usd(Usage(input_tokens=round(mean_pair), output_tokens=output_tokens + extra), model)
suite_point, suite_pair = one_point * 91, one_pair * 91 * 2
monthly = suite_pair * 20 * 22 # 20 CI runs a day, 22 working days
print(f"{model:<26}{suite_point:>11.4f}{suite_pair:>13.4f}{monthly:>10.2f}")
print("Prices from supportdesk/pricing.py (checked 21 Sep 2026; verify before relying on them).")
def demo(live: bool) -> None:
"""Build all three judge prompts for one draft; with --live, send them through supportdesk.llm.chat."""
drafts = {d["draft_id"]: d for d in load_jsonl(DRAFTS)}
bad, good = drafts["D-06"], drafts["D-03"] # a policy violation, and a hand-written 'send' reply
article = get_article(bad["gold_article"]).body
prompts = {"pointwise": pointwise_messages(bad["ticket"], article, bad["reply"]),
"pairwise": pairwise_messages(bad["ticket"], article, bad["reply"],
"After 5 failed attempts an account is locked for 15 minutes. Support "
"cannot unlock it sooner, so please try again then or use Forgot password."),
"reference": reference_messages(good["ticket"], good["reply"],
"You renewed 5 days ago, so you qualify for a full refund.")}
for name, messages in prompts.items():
print(f"{name:<10} {count_messages(messages):>4} input tokens user message starts: "
f"{messages[1]['content'][:60]!r}")
if live:
from supportdesk.llm import chat
result = chat(messages, response_format={"type": "json_object"})
print(f" judge said: {parse_judge(result.text)} ({result.usage.input_tokens} in, "
f"{result.usage.output_tokens} out, {result.latency_ms:.0f} ms)")
def main() -> None:
cmd = sys.argv[1] if len(sys.argv) > 1 else "calibrate"
live = "--live" in sys.argv
if cmd == "demo":
demo(live)
elif cmd == "calibrate":
calibrate(llm_pointwise_judge() if live else heuristic_judge, "LLM judge" if live else "stand-in judge")
elif cmd == "bias":
pairs = build_pairs()
judges = {"LLM judge": llm_pair_judge()} if live else {
"stand-in rule judge": heuristic_pair_judge, "SIMULATED biased judge": simulated_biased_judge()}
for name, judge in judges.items():
print(f"{name}: swap test {swap_test(judge, pairs)}")
print(f"{name}: verbosity test {verbosity_test(judge, pairs)}")
elif cmd == "cost":
judge_cost()
if __name__ == "__main__":
main()
Code explained
- In simple words: three ways to ask a model for a grade, a rule-based stand-in for when no model is available, and tools to check any judge for bias and against human labels.
- What happens:
pointwise_messages,pairwise_messages, andreference_messagesbuild the three judge prompts. The pairwise prompt tells the judge that order is random and length is not a virtue; that reduces bias but never removes it, which is why the swap test exists.parse_judgeextracts the first JSON object, because judges sometimes wrap JSON in prose even in JSON mode; the tests cover both cases.llm_pointwise_judgeandllm_pair_judgerun the prompts throughsupportdesk.llm.chat(or any function with the same signature) in JSON mode.heuristic_scoreis the stand-in judge: a rule, not a model. It scores 1 for a policy violation or when no key fact of the gold article appears, 2 for a long or padded reply, otherwise 3. It exists so the calibration and bias machinery runs here for real.calibratescores the 30 labelled drafts with any judge, prints agreement with the adjudicated labels, and the precision and recall of the "send as is" decision.swap_testjudges each pair in both orders and counts verdicts that survive the swap.verbosity_testpads the losing reply with polite filler and checks whether the judge changes its mind.simulated_biased_judgeis aScriptedLLMresponder that deliberately prefers position A and much longer replies; it exists only to prove the harness detects bias.build_pairspairs each non-baseline draft with the v2 baseline's reply to the same ticket (20 pairs).judge_costcounts real input tokens for the prompts and prices them.demobuilds one prompt of each kind and, with--live, sends them to your provider.
- Comes out: the commands below.
python examples/m10_judge.py demo
Code explained
- In simple words: build a pointwise, a pairwise, and a reference-based prompt for real drafts and count their tokens.
- What happens: the pointwise and pairwise prompts grade D-06 ("I have unlocked your account", a policy violation) against the login article; the pairwise prompt compares it with a correct reply; the reference-based prompt checks a short reply against the hand-written 3-rated answer to T-1004.
- Comes out:text
pointwise 392 input tokens user message starts: 'Help-center article:\nUse Forgot password on the sign-in page' pairwise 448 input tokens user message starts: 'Help-center article:\nUse Forgot password on the sign-in page' reference 215 input tokens user message starts: 'Ticket:\nSubject: Refund for annual plan\n\nWe renewed our annu'The reference-based prompt is about half the size because it carries a reference answer instead of the whole article. With a key,
python examples/m10_judge.py demo --livesends all three and prints each verdict with its token usage; a sound judge scores D-06 a 1, prefers the correct reply, and scores the short refund reply a 2 (it omits who can cancel and where the money goes). We could not run the live version here.
Position, verbosity, and self-preference bias
Judges have measured, repeatable biases. The ones that matter most:
- Position bias: preferring whichever answer is shown first (or second). Zheng et al. (2023), "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (NeurIPS 2023 Datasets and Benchmarks), documented it along with verbosity and self-enhancement bias, while also finding that GPT-4 as a judge agreed with human preferences over 80% of the time, about the level at which humans agree with each other. Wang et al. (2023), "Large Language Models are not Fair Evaluators", showed how strong it can be: with ChatGPT as the judge, simply reordering the candidates let Vicuna-13B beat ChatGPT on 66 of 80 queries. Their fix, which we use, is to judge both orders and combine.
- Verbosity bias: preferring longer answers even when the extra text adds nothing (Zheng et al., 2023). Support drafts should be short, so this bias works directly against Maya's standard.
- Self-preference: favoring outputs from the judge's own model. Panickssery, Bowman, and Feng (2024), "LLM Evaluators Recognize and Favor Their Own Generations", found that models can recognize their own text and that the strength of self-preference tracks how well they recognize it. The practical rule: do not let the model that wrote the drafts be the only judge of them, especially in model comparisons.
The swap harness tests the first two on any pairwise judge. Here it runs on the deterministic rule judge and on the deliberately biased ScriptedLLM stand-in (simulated bias, not a real model's behavior):
python examples/m10_judge.py bias
Code explained
- In simple words: show every pair to the judge twice with the order swapped, then pad the losing reply with filler, and count how often the verdict changes.
- What happens: for each of 20 same-ticket pairs,
swap_testcalls the judge with x first, then y first. If both calls pick the same underlying reply (or both say tie), the verdict is consistent. If the first-shown reply wins both times, that is position bias.verbosity_testtakes pairs where the judge clearly preferred x in both orders, adds three sentences of polite filler to y, and checks whether x still wins. - Comes out:
Calibrating a judge against human labels
A judge is calibrated when its scores agree with adjudicated human labels at a level you have measured and accepted. Measure it with kappa, the confusion table, and the decision you will actually use the score for. Here that decision is "send as is" (score 3), so precision matters: of drafts the judge calls sendable, how many really are?
python examples/m10_judge.py calibrate
Code explained
- In simple words: grade the 30 labelled drafts with the stand-in judge and compare to the adjudicated labels.
- What happens:
calibraterunsheuristic_judgeon every draft, thenagreement_reportand the send-decision counts. Rows of the table are the human label, columns the judge. - Comes out: (stand-in rule judge against the hand labels from Part D)text
human (adjudicated) vs stand-in judge: n=30 raw agreement 70.0% kappa 0.551 linear-weighted kappa 0.603 rows = first rater, columns = second rater 1 2 3 1 9 0 2 2 2 5 3 3 0 2 7 'send as is' (score 3): judge said 12, humans said 9, both 7 precision 0.58 recall 0.78 disagreements (draft, human, judge): [('D-09', 1, 3), ('D-10', 2, 1), ('D-13', 2, 3), ('D-15', 3, 2), ('D-17', 2, 3), ('D-18', 3, 2), ('D-21', 2, 3), ('D-23', 2, 1), ('D-29', 1, 3)]Kappa 0.551 is "moderate", well below the 0.748 between the two label passes, and the disagreements explain why. The judge gave 3 to D-09 ("SCIM ... is included on the Business plan"), which is wrong: the reply contains the key fact pattern "SCIM", so the rule sees grounding where there is a false claim. The same happened with D-29's invented settings path. In the other direction it gave 1 to D-23 and D-10 because their correct facts are not in its short key-fact list. The "send" precision of 0.58 means 5 of the 12 drafts it would auto-approve should not be sent. That is the typical failure of a cheap judge: it is too lenient on confident, specific, wrong statements, which are the most dangerous ones.
To calibrate a real model judge, run the same command live:
LLM_PROVIDER=groq python examples/m10_judge.py calibrate --live
Code explained
- In simple words: the same 30 drafts and labels, graded by a real model through
supportdesk.llm.chat. - What happens:
llm_pointwise_judgesends each draft with its gold article and the rubric, in JSON mode, and parses the score. Thirty calls of about 390 input tokens each. - Comes out: we could not run this here. Read your output the same way: kappa against the adjudicated labels, the send precision, and every disagreement. Do not accept a judge whose kappa is well below your human-human kappa, or whose send precision is below what Maya would accept from a new agent. Improve the prompt on disagreements (add the rubric rule the judge missed), then re-measure on fresh labels, not the same 30, or you will overfit the judge to its calibration set.
Judge cost, and when it is not worth it
A judge call costs about as much as the call it grades, sometimes more, because the judge reads the ticket, the article, and the draft. The cost script counts real input tokens for the prompts in this module and prices them with pricing.py. Output and reasoning token counts are assumptions, stated in the output.
python examples/m10_judge.py cost
Code explained
- In simple words: what judging the 91-case suite would cost per run and per month on several models.
- What happens:
judge_costcounts tokens withcount_messagesfor all 30 pointwise and 20 pairwise prompts, assumes 150 output tokens plus 400 reasoning tokens on reasoning models, and prices one pointwise suite (91 calls), one pairwise suite judged in both orders (182 calls), and a month of pairwise CI runs (20 runs a day, 22 working days). - Comes out:
| Situation | Use this | Why |
|---|---|---|
| The criterion can be checked by code (schema, citation, forbidden phrases) | Deterministic check, no judge | Free, exact, no drift |
| Every CI run on every commit | Deterministic checks always; judge only a sample or only changed slices | Keeps CI fast and cheap |
| Release decision between two systems | Pairwise judge, both orders, on the full set | Worth the cost once per release |
| The judge's kappa against humans is low | Fix the rubric and prompt, or use humans | An uncalibrated judge adds noise, not signal |
| High-stakes decisions (refunds, legal) | Humans, possibly with judge pre-sorting | The cost of a wrong "send" exceeds any saving |