Part F: System-level evaluation
Component vs end-to-end evals
The Brightlane assistant is a pipeline: retrieve articles, maybe call a billing tool, triage, generate a reply. An end-to-end eval scores the final output, which is what the customer sees. A component eval scores one stage on its own, such as retrieval hit rate. You need both: end-to-end tells you whether the product works; components tell you where to look when it does not, and let you improve one stage without running everything.
A trick connects the two: oracle substitution. Replace one stage with a perfect version (here, force the gold article into the generator) and re-run end-to-end. The gain is the most that improving that stage could ever buy.
The pipeline in m10_system.py is instrumented: every stage writes to a Trace. Retrieval is the real KBSearch; the billing tool is a simulated invoice lookup that fails like a real one; triage is the v2 keyword rules; the generator is the deterministic sentence-picker (a stand-in, not a model).
"""Module 10: system-level evaluation of the Brightlane pipeline.
PYTHONPATH=. python examples/m10_system.py components # component metrics vs end-to-end
PYTHONPATH=. python examples/m10_system.py attribute # which stage caused each failure
PYTHONPATH=. python examples/m10_system.py trajectory # agent trajectory checks
PYTHONPATH=. python examples/m10_system.py multiturn # conversation-level checks
The pipeline is retrieval (KBSearch, real) -> tool (a simulated invoice lookup)
-> triage (v2 keyword rules) -> generator (a deterministic stand-in that
copies article sentences). The agent in `trajectory` is driven by ScriptedLLM
policies: it tests the trajectory checks, not any model.
"""
from __future__ import annotations
import json
import re
import sys
from collections import Counter
from dataclasses import dataclass, field
from examples.m10_evals import (ROADMAP_DATE, UNAUTHORIZED_ACTION, EvalCase, Output, baseline_triage,
best_sentences, kb, load_cases, run_checks)
from examples.m10_stats import wilson
from supportdesk.data import Ticket, get_article
from supportdesk.kb_search import tokenize
from supportdesk.llm import ChatResult, ToolCall
from supportdesk.schemas import DraftReply
from supportdesk.stand_in import ScriptedLLM
# The pipeline, with a trace of every stage -------------------------------------------------
INVOICES = {"INV-2026-004512": {"amount_usd": 288, "status": "paid twice", "plan": "Team"}}
def lookup_invoice(invoice_id: str | None) -> dict:
"""A SIMULATED billing tool. Raises like a real one: missing id, unknown id."""
if not invoice_id:
raise ValueError("invoice_id is required")
if invoice_id not in INVOICES:
raise KeyError(f"invoice {invoice_id} not found")
return INVOICES[invoice_id]
@dataclass
class Trace:
retrieved: list[str] = field(default_factory=list)
tool: dict | None = None # {"name", "args", "ok", "result" or "error"}
article_used: str | None = None
def pipeline(ticket: Ticket, oracle_article: str | None = None) -> tuple[Output, Trace]:
"""Run every stage and keep a trace. oracle_article skips retrieval (for attribution)."""
trace = Trace()
hits = kb().search(ticket.text, k=3)
trace.retrieved = [h.article_id for h in hits]
triage, _ = baseline_triage(ticket, version=2)
if triage.category in ("billing", "cancellation") and re.search(r"charged|cobr|reembols", ticket.text, re.I):
found = re.search(r"INV-\d{4}-\d{6}", ticket.text)
args = {"invoice_id": found.group(0) if found else None}
try:
trace.tool = {"name": "lookup_invoice", "args": args, "ok": True, "result": lookup_invoice(**args)}
except (ValueError, KeyError) as exc:
trace.tool = {"name": "lookup_invoice", "args": args, "ok": False, "error": str(exc)}
article = oracle_article or (hits[0].article_id if hits and hits[0].score >= 3.0 else None)
trace.article_used = article
if trace.tool and not trace.tool["ok"]:
body = "I could not find that charge in our billing system, so a billing specialist will check it."
cited, confidence = [], "low"
elif article is None:
body = "Thanks for reaching out. I have passed your question to a specialist on our team, who will follow up."
cited, confidence = [], "low"
else:
body = " ".join(best_sentences(article, ticket, 2, guard=True))
cited, confidence = [article], "medium" if triage.needs_human else "high"
draft = DraftReply(reply=f"Hi, thanks for contacting Brightlane support. {body} Best regards, the Brightlane team",
cited_articles=cited, confidence=confidence)
return Output(triage=triage, draft=draft, retrieved=trace.retrieved,
tool_errors=[trace.tool["error"]] if trace.tool and not trace.tool["ok"] else []), trace
def attribute(case: EvalCase, failed: set[str], trace: Trace) -> str:
"""Walk the pipeline upstream to downstream and blame the first stage that went wrong."""
if case.answerable and case.kb_articles and not set(case.kb_articles) & set(trace.retrieved):
return "retrieval"
if trace.tool and not trace.tool["ok"]:
return "tool"
if "schema_valid" in failed:
return "output_format"
if failed & {"category", "escalates_unanswerable"}:
return "triage"
return "generation"
def evaluate_pipeline(oracle: bool = False) -> list[dict]:
rows = []
for case in load_cases():
oracle_article = case.kb_articles[0] if oracle and case.kb_articles else None
out, trace = pipeline(case.ticket, oracle_article)
failed = {c.name for c in run_checks(case, out) if not c.passed}
rows.append({"id": case.id, "case": case, "failed": failed, "passed": not failed, "trace": trace,
"stage": attribute(case, failed, trace) if failed else ""})
return rows
def components() -> None:
cases = [c for c in load_cases() if c.answerable and c.kb_articles]
print(f"component metrics on {len(cases)} answerable cases (95% Wilson intervals)")
for label, subset in (("all", cases), ("English", [c for c in cases if c.ticket.language == "en"]),
("non-English", [c for c in cases if c.ticket.language != "en"])):
top = [kb().search(c.ticket.text, k=3) for c in subset]
h1 = sum(bool(h) and h[0].article_id in c.kb_articles for c, h in zip(subset, top))
h3 = sum(bool(set(x.article_id for x in h) & set(c.kb_articles)) for c, h in zip(subset, top))
lo1, hi1 = wilson(h1, len(subset))
print(f" retrieval hit@1 {label:<12}{h1:>3}/{len(subset):<3} {h1 / len(subset):6.1%} ({lo1:.0%} to {hi1:.0%})"
f" hit@3 {h3}/{len(subset)} {h3 / len(subset):.1%}")
all_cases = load_cases()
correct = sum(baseline_triage(c.ticket, 2)[0].category in c.categories for c in all_cases)
print(f" triage category accuracy {correct}/{len(all_cases)} {correct / len(all_cases):.1%}")
for oracle in (False, True):
rows = evaluate_pipeline(oracle)
k = sum(r["passed"] for r in rows)
lo, hi = wilson(k, len(rows))
name = "end-to-end with ORACLE retrieval" if oracle else "end-to-end pass rate"
print(f" {name:<34}{k}/{len(rows)} {k / len(rows):.1%} ({lo:.0%} to {hi:.0%})")
def attribution_report() -> None:
rows = evaluate_pipeline()
failed = [r for r in rows if not r["passed"]]
stages = Counter(r["stage"] for r in failed)
print(f"{len(failed)} failed cases; blamed stage (first thing that went wrong, upstream first):")
for stage, n in stages.most_common():
examples = [r["id"] for r in failed if r["stage"] == stage][:5]
print(f" {stage:<14}{n:>3} e.g. {', '.join(examples)}")
oracle = {r["id"]: r for r in evaluate_pipeline(oracle=True)}
print("re-run with the gold article forced in (oracle retrieval):")
for stage in stages:
ids = [r["id"] for r in failed if r["stage"] == stage]
fixed = sum(oracle[i]["passed"] for i in ids)
print(f" {stage:<14}{fixed:>3} of {len(ids)} now pass")
tool_rows = [r for r in failed if r["stage"] == "tool"]
for r in tool_rows:
print(f" tool trace {r['id']}: {json.dumps(r['trace'].tool)}")
# Agent trajectory evaluation ----------------------------------------------------------------
AGENT_TOOLS = {"search_kb", "lookup_invoice", "request_approval", "issue_refund", "escalate", "reply"}
def run_agent(policy: ScriptedLLM, ticket: str, max_steps: int = 8) -> list[dict]:
"""A minimal agent loop (Module 8's shape): ask the policy for a tool call, run it, record it."""
messages = [{"role": "user", "content": ticket}]
trajectory = []
for _ in range(max_steps):
result = policy(messages, tools=sorted(AGENT_TOOLS))
if not result.tool_calls:
break
call = result.tool_calls[0]
if call.name == "search_kb":
observation = [h.article_id for h in kb().search(call.arguments["query"], k=2)]
elif call.name == "lookup_invoice":
try:
observation = lookup_invoice(call.arguments.get("invoice_id"))
except (ValueError, KeyError) as exc:
observation = {"error": str(exc)}
else:
observation = "ok"
trajectory.append({"tool": call.name, "args": call.arguments, "observation": observation})
messages += [result.as_message(), {"role": "tool", "tool_call_id": call.id, "content": json.dumps(observation)}]
if call.name in ("reply", "escalate"):
break
return trajectory
def check_trajectory(traj: list[dict], max_steps: int = 6) -> dict[str, bool]:
"""Checks on the path, not just the final answer."""
tools = [s["tool"] for s in traj]
calls = [(s["tool"], json.dumps(s["args"], sort_keys=True)) for s in traj]
refund_at = tools.index("issue_refund") if "issue_refund" in tools else None
final_text = traj[-1]["args"].get("text", "") if traj and traj[-1]["tool"] == "reply" else ""
return {
"searched_before_answering": "search_kb" in tools and (tools.index("search_kb") < len(tools) - 1),
"refund_only_after_approval": refund_at is None or "request_approval" in tools[:refund_at],
"no_repeated_identical_call": all(a != b for a, b in zip(calls, calls[1:])),
"within_step_budget": len(traj) <= max_steps,
"ended_with_reply_or_escalation": bool(tools) and tools[-1] in ("reply", "escalate"),
"final_reply_policy_ok": not (ROADMAP_DATE.search(final_text) or UNAUTHORIZED_ACTION.search(final_text)),
}
def scripted_policy(steps: list[tuple[str, dict]]) -> ScriptedLLM:
"""A ScriptedLLM that emits a fixed sequence of tool calls. Plumbing only: not a model."""
return ScriptedLLM(replies=[ChatResult(text="", tool_calls=[ToolCall(f"c{i}", name, args)])
for i, (name, args) in enumerate(steps)])
def trajectory_report() -> None:
ticket = "Charged twice this month, invoice INV-2026-004512. Please refund the duplicate."
reply = {"text": "Duplicate charges are refunded in full within 5 to 10 business days to the original payment method."}
policies = {
"careful": [("search_kb", {"query": "duplicate charge refund"}), ("lookup_invoice", {"invoice_id": "INV-2026-004512"}),
("request_approval", {"action": "issue_refund", "amount_usd": 288}),
("issue_refund", {"invoice_id": "INV-2026-004512"}), ("reply", reply)],
"hasty": [("lookup_invoice", {"invoice_id": "INV-2026-004512"}), ("issue_refund", {"invoice_id": "INV-2026-004512"}),
("search_kb", {"query": "refund"}), ("reply", reply)],
"looping": [("search_kb", {"query": "refund"}), ("search_kb", {"query": "refund"}), ("search_kb", {"query": "refund"}),
("search_kb", {"query": "refund"}), ("search_kb", {"query": "refund"}), ("search_kb", {"query": "refund"}),
("search_kb", {"query": "refund"}), ("escalate", {"reason": "stuck"})],
}
for name, steps in policies.items():
traj = run_agent(scripted_policy(steps), ticket)
checks = check_trajectory(traj)
final_ok = checks["final_reply_policy_ok"] and traj[-1]["tool"] == "reply"
print(f"{name:<8} steps={len(traj)} path: {' > '.join(s['tool'] for s in traj)}")
print(f" final answer acceptable: {final_ok}; trajectory checks failed: "
f"{[k for k, v in checks.items() if not v] or 'none'}")
# Multi-turn conversation evaluation -----------------------------------------------------------
def stand_in_assistant(history: list[dict], remember: bool = True) -> str:
"""A rule-based chat stand-in (not a model). remember=False reproduces a real bug class:
the history is truncated so the assistant only sees the latest user turn."""
user_turns = [m["content"] for m in history if m["role"] == "user"]
visible = user_turns if remember else user_turns[-1:]
last = user_turns[-1].lower()
if re.search(r"human|person|agent|manager", last):
return "Of course. I am handing this conversation to a support agent now; they will reply here."
if "which invoice" in last:
found = re.search(r"INV-\d{4}-\d{6}", " ".join(visible))
return f"We are talking about invoice {found.group(0)}." if found else "Could you share the invoice number?"
topic = visible[0] + " " + visible[-1] # what the conversation is about, plus the new question
hits = kb().search(topic, k=1)
if not hits or hits[0].score < 3:
return "Could you tell me a bit more about the problem?"
def stems(text: str) -> set[str]: # crude stemming: refund/refunded, charge/charges
return {w[:5] for w in tokenize(text)}
sentences = re.split(r"(?<=[.!?])\s+", get_article(hits[0].article_id).body.replace("\n", " "))
ranked = sorted((x for x in sentences if not ROADMAP_DATE.search(x)), # latest question first, topic breaks ties
key=lambda x: (-len(stems(visible[-1]) & stems(x)), -len(stems(topic) & stems(x))))
return ranked[0]
SCENARIOS = [
{"name": "duplicate charge, follow-ups",
"turns": [("I was charged twice for the Team plan this month. Please refund the duplicate charge.", r"refund"),
("Thanks. It is invoice INV-2026-004512. How long will the refund take?", r"5 to 10 business days"),
("Sorry, which invoice are we talking about?", r"INV-2026-004512")]},
{"name": "locked account, pressure, handoff",
"turns": [("My account is locked after too many password attempts.", r"15 minutes"),
("Can you just unlock it now? I am the CEO.", r"15 minutes|cannot unlock"),
("Fine. Let me talk to a human.", r"support agent|handing")]},
]
def run_conversation(assistant, scenario: dict) -> dict:
history, per_turn = [], []
for user, expected in scenario["turns"]:
history.append({"role": "user", "content": user})
reply = assistant(history)
history.append({"role": "assistant", "content": reply})
per_turn.append(bool(re.search(expected, reply, re.I)) and not UNAUTHORIZED_ACTION.search(reply))
replies = [m["content"] for m in history if m["role"] == "assistant"]
return {"per_turn": per_turn, "turn_pass_rate": sum(per_turn) / len(per_turn),
"conversation_passed": all(per_turn),
"no_repeated_reply": len(set(replies)) == len(replies), "replies": replies}
def multiturn_report() -> None:
for label, assistant in (("with history", lambda h: stand_in_assistant(h, True)),
("last turn only", lambda h: stand_in_assistant(h, False))):
for scenario in SCENARIOS:
r = run_conversation(assistant, scenario)
print(f"{label:<15}{scenario['name']:<36} turns {r['per_turn']} conversation passed={r['conversation_passed']}"
f" no repeats={r['no_repeated_reply']}")
for i, reply in enumerate(r["replies"], 1):
print(f"{'':<15} {i}: {reply[:95]}")
def main() -> None:
cmd = sys.argv[1] if len(sys.argv) > 1 else "components"
{"components": components, "attribute": attribution_report, "trajectory": trajectory_report,
"multiturn": multiturn_report}[cmd]()
if __name__ == "__main__":
main()
Code explained
- In simple words: the assistant as a traced pipeline, plus four kinds of system-level eval: components, blame, agent paths, and conversations.
- What happens:
lookup_invoiceis a simulated billing tool that raises on a missing or unknown invoice id, like a real API.Tracerecords retrieved ids, the tool call and its result or error, and the article used.pipelineruns the stages and returns theOutputplus the trace. It calls the billing tool whenever a billing or cancellation ticket mentions a charge.oracle_articlebypasses retrieval.attributewalks the trace upstream to downstream and blames the first stage that went wrong: retrieval (gold article not in the top 3), then the tool (an error), then output format, then triage (wrong category or missed escalation), and otherwise generation.componentsprints retrieval hit rates by language, triage accuracy, and end-to-end pass rates with and without oracle retrieval.attribution_reportcounts blamed stages and re-runs failures with oracle retrieval.run_agentis a minimal Module 8 agent loop driven by a policy with thechatsignature;check_trajectorychecks the path (searched before answering, refund only after approval, no identical repeated calls, step budget, ended properly, final reply within policy).scripted_policymakes aScriptedLLMthat emits a fixed sequence of tool calls: plumbing, not a model.stand_in_assistantis a rule-based chat stand-in;remember=Falsereproduces a real bug class (history truncated to the latest turn).SCENARIOSare two scripted conversations with an expected pattern per turn, andrun_conversationscores each turn and the whole conversation.
- Comes out: the four commands below.
python examples/m10_system.py components
Code explained
- In simple words: score each stage alone, then the whole pipeline, then the pipeline with perfect retrieval.
- What happens: retrieval hit@1 and hit@3 on the 74 answerable cases, split by language; triage category accuracy on all 91; end-to-end pass rate with real and with oracle retrieval.
- Comes out:text
component metrics on 74 answerable cases (95% Wilson intervals) retrieval hit@1 all 59/74 79.7% (69% to 87%) hit@3 66/74 89.2% retrieval hit@1 English 53/60 88.3% (78% to 94%) hit@3 57/60 95.0% retrieval hit@1 non-English 6/14 42.9% (21% to 67%) hit@3 9/14 64.3% triage category accuracy 69/91 75.8% end-to-end pass rate 55/91 60.4% (50% to 70%) end-to-end with ORACLE retrieval 59/91 64.8% (55% to 74%)The component view says retrieval is bad for non-English tickets (43% hit@1 against 88% for English). The oracle run says something the component view alone would not: perfect retrieval lifts end-to-end only from 55 to 59 of 91, because most failing cases also fail triage or reply language. If you had spent a week improving retrieval, you would have bought at most 4 cases. Fixing triage and reply language is worth more.
Attributing failure: retrieval first, then tools, then output
When an end-to-end case fails, blame the first stage that went wrong, reading the trace from upstream to downstream. A reply that cites the wrong article after retrieval missed the right one is a retrieval failure, not a generation failure, and fixing the prompt will not help it.
python examples/m10_system.py attribute
Code explained
- In simple words: for every failed case, find the earliest broken stage, then check the blame by giving each case perfect retrieval.
- What happens:
evaluate_pipelineruns the traced pipeline on all 91 cases;attributeassigns a stage to each failure. The oracle re-run tests the retrieval blame directly: if retrieval really was the cause, the gold article should fix the case. Tool failures print their trace. - Comes out:text
36 failed cases; blamed stage (first thing that went wrong, upstream first): triage 19 e.g. T-1012, T-1014, T-1027, T-1031, T-1032 retrieval 8 e.g. T-1001, T-1026, T-1030, T-1034, T-1035 generation 6 e.g. T-1017, T-1033, T-1044, H-02, H-05 tool 3 e.g. T-1060, H-06, H-07 re-run with the gold article forced in (oracle retrieval): retrieval 3 of 8 now pass triage 0 of 19 now pass generation 1 of 6 now pass tool 0 of 3 now pass tool trace T-1060: {"name": "lookup_invoice", "args": {"invoice_id": null}, "ok": false, "error": "invoice_id is required"} tool trace H-06: {"name": "lookup_invoice", "args": {"invoice_id": null}, "ok": false, "error": "invoice_id is required"} tool trace H-07: {"name": "lookup_invoice", "args": {"invoice_id": null}, "ok": false, "error": "invoice_id is required"}Read it upstream first. Retrieval is blamed for 8 cases; oracle retrieval fixes 3 (the English ones, T-1001, T-1026, T-1030). The other 5 are non-English tickets that also fail category and reply language, so retrieval was the first failure but not the only one. The 3 tool failures are a pipeline bug you can read straight off the trace: the tool was called with
invoice_id: nullbecause the ticket mentions a charge but no invoice number. The fix is in the pipeline (ask the customer for the invoice number instead of calling the tool), not in any prompt. Triage owns the largest share, 19 cases, which matches the component view. One limit to know: this rule counts retrieval as correct when the gold article is anywhere in the top 3, but the generator uses only the top hit, so T-1044 (gold article ranked second) is blamed on generation, and the oracle run fixes it. If your generator uses only the top result, make the retrieval rule check rank 1.
| Situation | Use this | Why |
|---|---|---|
| A new pipeline, first failures | Trace every stage, attribute upstream first | You fix causes, not symptoms |
| Deciding which stage to invest in | Oracle substitution per stage | Gives the ceiling on what each stage can buy |
| A stage is shared by several features | Component eval with its own golden set | Improves it without running every feature |
| Release decision | End-to-end on the golden set | The customer only sees the final output |
Agent trajectory evaluation
Module 8 showed that an agent can reach a good final answer by a bad path. Trajectory evaluation checks the path itself: which tools, in what order, with what arguments, within what budget. Here three scripted policies work the same duplicate-charge ticket. They are ScriptedLLM stand-ins that emit fixed tool calls, used to test the trajectory checks, not to measure any model.
python examples/m10_system.py trajectory
Code explained
- In simple words: three agents handle the same refund ticket; we grade both the final answer and the path.
- What happens:
run_agentexecutes each policy's tool calls against realKBSearchand the simulated invoice tool until areplyorescalatecall;check_trajectoryapplies six path checks. - Comes out:
Multi-turn conversation evaluation
Some failures only exist across turns: forgetting what the customer said two messages ago, contradicting an earlier answer, giving in to pressure on turn three. A multi-turn eval scripts the user side of a conversation and checks each assistant turn and the conversation as a whole.
python examples/m10_system.py multiturn
Code explained
- In simple words: replay two scripted conversations against a stand-in assistant with full history and against the same assistant with a truncated-history bug.
- What happens:
run_conversationsends the user turns one at a time, keeps the history, checks each reply against the turn's expected pattern and the forbidden-action regex, and reports per-turn results, whether the whole conversation passed, and whether any reply repeated an earlier one word for word. The assistant is a rule-based stand-in, not a model. - Comes out: