CourseLarge Language Models · Module 8: Agents · part 42 of 80
Part 42 · Module 8: Agents

Part 7: Multi-agent systems

14 min read·22 Sept 2026

Orchestrator and workers

In the orchestrator-worker pattern, one agent (the orchestrator) breaks the task into pieces and delegates each piece to a worker agent with a specialized role: its own system prompt, its own small tool set, its own context. Workers report back; the orchestrator combines the reports. For Brightlane we split T-1001 into an account worker (only get_account) and a billing worker (policy search, invoice, refund).

Flowchart

The file below builds it from the same run_agent loop. A worker is a tool whose function runs a whole agent. It also contains two more experiments: failure compounding and parallel exploration.

examples/m08_multi_agent.py

python
"""Module 8: orchestrator-worker agents, their token bill, parallel exploration, and failure compounding.

Orchestrator and workers are the same run_agent loop with different tools and
prompts. Replies come from ScriptedLLM policies (not a model); token counts are
real counts of the prompts each design sends.
"""
from __future__ import annotations

import json
import math
import random
import time
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path

from m08_agent import (EMAIL_RE, INVOICE_RE, T1001, BillingStore, ScriptedApprover, SupportPolicy, Tool,
                       history, make_tools, prompt_tokens, run_agent)
from supportdesk.data import load_tickets
from supportdesk.kb_search import KBSearch
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_tokens

WORK = Path("runs/m08/multi")
WORK.mkdir(parents=True, exist_ok=True)
for old in WORK.glob("*"):
    old.unlink()

ORCH_SYSTEM = ("You coordinate Brightlane support. Delegate account questions to delegate_account and "
               "billing checks or refunds to delegate_billing. Write each brief so the worker needs nothing else. "
               "Then reply to the customer.")
WORKER_SYSTEM = {
    "account": "You are the account worker. Use get_account. Report account_id, plan, seats and role in one line.",
    "billing": ("You are the billing worker. Confirm the refund policy with search_kb and read_article, check the "
                "invoice with get_invoice, and refund only a duplicate charge (issue_refund needs human approval). "
                "Report the refund id and amount in one line."),
}


class AccountWorker(SupportPolicy):
    def __call__(self, messages, kwargs):
        brief, done, tools = messages[1]["content"], history(messages), kwargs.get("tools")
        if not done:
            return self.reply(messages, "Thought: look up the requester.", "get_account",
                              {"email": EMAIL_RE.search(brief).group(0)}, tools)
        if not done[-1][2].get("ok"):
            return self.reply(messages, f"ACCOUNT LOOKUP FAILED: {done[-1][2].get('error')}")
        a = done[-1][2]["result"]
        return self.reply(messages, f"ACCOUNT {a['account_id']} plan={a['plan']} seats={a['seats']} role={a['role']}")


class BillingWorker(SupportPolicy):
    def __call__(self, messages, kwargs):
        brief, done, tools = messages[1]["content"], history(messages), kwargs.get("tools")
        names = [d[0] for d in done]
        r = lambda *a: self.reply(messages, *a, tools=tools)  # noqa: E731
        if "search_kb" not in names:
            return r("Thought: confirm policy.", "search_kb", {"query": "duplicate charge refund"})
        if "read_article" not in names:
            return r("Thought: read it.", "read_article", {"article_id": done[0][2]["result"][0]["article_id"]})
        inv = [d for d in done if d[0] == "get_invoice" and d[2].get("ok")]
        if not inv:
            return r("Thought: check charges.", "get_invoice", {"invoice_id": INVOICE_RE.search(brief).group(0)})
        invoice = inv[-1][2]["result"]
        if "issue_refund" not in names:
            dup = invoice["charges"][-1]
            return r("Thought: duplicate confirmed; refund it.", "issue_refund",
                     {"invoice_id": invoice["invoice_id"], "charge_id": dup["charge_id"],
                      "amount_usd": invoice["overpaid"], "reason": "duplicate charge"})
        rf = done[-1][2]
        if not rf.get("ok"):
            return r(f"REFUND NOT DONE: {rf.get('error')}")
        rf = rf["result"]
        return r(f"REFUND {rf['refund_id']} {rf['amount']} USD on {rf['invoice_id']} charge {rf['charge_id']}")


class Orchestrator(SupportPolicy):
    def __call__(self, messages, kwargs):
        ticket, done, tools = messages[1]["content"], history(messages), kwargs.get("tools")
        names = [d[0] for d in done]
        r = lambda *a: self.reply(messages, *a, tools=tools)  # noqa: E731
        if "delegate_account" not in names:
            return r("Thought: account first, then billing.", "delegate_account",
                     {"brief": f"Requester {EMAIL_RE.search(ticket).group(0)}. Confirm the account, plan, seats and role."})
        if "delegate_billing" not in names:
            acct = done[0][2]["result"]
            body = ticket.split("\n\n", 1)[1]
            return r("Thought: account is fine; hand billing the facts.", "delegate_billing",
                     {"brief": f"{acct}. Customer wrote: {body} Verify the duplicate on the invoice and refund it."})
        report = done[-1][2]["result"]
        return r(f"Hi, sorry about the double charge. Our billing team confirmed it: {report}. "
                 "The money returns to your original card within 5 to 10 business days.")


def run_multi_agent(ticket: str, store: BillingStore, approver) -> dict:
    workers: list = []  # (name, state, llm) for every worker run

    def delegate(kind: str, policy_cls, tool_names: list[str], flaky: int):
        all_tools = make_tools(store, flaky_failures=flaky)
        tools = {n: all_tools[n] for n in tool_names}

        def run(brief: str) -> str:
            llm = ScriptedLLM(responder=policy_cls())
            state = run_agent(brief, f"T-1001/{kind}", tools, llm, approver=approver, system_prompt=WORKER_SYSTEM[kind],
                              trace_path=WORK / f"{kind}.jsonl", sleep=lambda s: None)
            workers.append((kind, state, llm))
            return state.final
        return run

    brief_param = {"type": "object", "properties": {"brief": {"type": "string"}}, "required": ["brief"],
                   "additionalProperties": False}
    orch_tools = {
        "delegate_account": Tool("delegate_account", "Send a self-contained brief to the account worker; returns its one-line report.",
                                 brief_param, delegate("account", AccountWorker, ["get_account"], 0)),
        "delegate_billing": Tool("delegate_billing", "Send a self-contained brief to the billing worker; returns its one-line report.",
                                 brief_param, delegate("billing", BillingWorker,
                                                       ["search_kb", "read_article", "get_invoice", "issue_refund"], 2)),
    }
    orch_llm = ScriptedLLM(responder=Orchestrator())
    orch = run_agent(ticket, "T-1001", orch_tools, orch_llm, approver=approver, system_prompt=ORCH_SYSTEM,
                     trace_path=WORK / "orchestrator.jsonl")
    return {"orchestrator": (orch, orch_llm), "workers": workers}


def duplicated_tokens(llm: ScriptedLLM, snippet: str) -> int:
    """Tokens of `snippet` re-sent across every call of one context (it rides along in each prompt)."""
    n = count_tokens(snippet)
    return sum(n for call in llm.calls if any(snippet in (m.get("content") or "") for m in call["messages"]))


def static_prefix_tokens(llm: ScriptedLLM) -> int:
    """System prompt, first user message, and tool definitions: paid again on every call."""
    return sum(prompt_tokens(call["messages"][:2], call.get("tools")) for call in llm.calls)


def compounding_table() -> None:
    print("\nP(all k steps succeed) = p ** k        (analytic | simulated, 20,000 trials)")
    rng = random.Random(8)
    print("   p    k=1     k=3     k=5     k=10")
    for p in (0.99, 0.95, 0.90):
        cells = []
        for k in (1, 3, 5, 10):
            sim = sum(all(rng.random() < p for _ in range(k)) for _ in range(20_000)) / 20_000
            cells.append(f"{p ** k:.3f}|{sim:.3f}")
        print(f"  {p:.2f} " + " ".join(cells))
    p, k, recall = 0.90, 5, 0.8
    p_checked = p + (1 - p) * recall * p   # a checker catches 80% of failures and the step is retried once
    print(f"with a checker (recall {recall}) and one retry per step: p={p} becomes {p_checked:.3f}; "
          f"k={k} chain: {p ** k:.3f} -> {p_checked ** k:.3f}")


def parallel_exploration() -> None:
    """Three retrieval 'explorers' per ticket, run in parallel, aggregated three ways."""
    kb = KBSearch()
    tickets = [t for t in load_tickets() if t.gold["answerable"]]
    explorers = {"full_text": lambda t: t.text, "subject": lambda t: t.subject, "body": lambda t: t.body}
    hits = {name: 0 for name in [*explorers, "fused_rrf", "fused_max_score", "any_explorer"]}

    def explore(args):
        name, query = args
        time.sleep(0.02)  # stands in for an LLM worker's network wait (simulated, not measured)
        return name, kb.search(query, k=5)

    for parallel in (False, True):
        started = time.perf_counter()
        for t in tickets:
            jobs = [(n, f(t)) for n, f in explorers.items()]
            if parallel:
                with ThreadPoolExecutor(max_workers=3) as pool:
                    results = dict(pool.map(explore, jobs))
            else:
                results = dict(map(explore, jobs))
            if not parallel:
                continue
            gold = t.gold["kb_article"]
            rrf: dict[str, float] = {}
            for name, ranked in results.items():
                hits[name] += bool(ranked) and ranked[0].article_id == gold
                for i, h in enumerate(ranked):
                    rrf[h.article_id] = rrf.get(h.article_id, 0) + 1 / (60 + i + 1)
            hits["fused_rrf"] += bool(rrf) and max(rrf, key=rrf.get) == gold
            tops = [r[0] for r in results.values() if r]
            hits["fused_max_score"] += bool(tops) and max(tops, key=lambda h: h.score).article_id == gold
            hits["any_explorer"] += any(r and r[0].article_id == gold for r in results.values())
        print(f"{'parallel' if parallel else 'sequential'} wall time for {len(tickets)} tickets: "
              f"{time.perf_counter() - started:.2f} s")
    n = len(tickets)
    for name, h in hits.items():
        lo, hi = wilson(h, n)
        print(f"  {name:16} top-1 {h}/{n} = {h / n:.2f}  (95% CI {lo:.2f} to {hi:.2f})")


def wilson(k: int, n: int, z: float = 1.96) -> tuple[float, float]:
    p = k / n
    centre = (p + z * z / (2 * n)) / (1 + z * z / n)
    half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
    return centre - half, centre + half


if __name__ == "__main__":
    # Single agent baseline
    single_llm = ScriptedLLM(responder=SupportPolicy())
    single = run_agent(T1001, "T-1001", make_tools(BillingStore(WORK / "single.json"), flaky_failures=2), single_llm,
                       approver=ScriptedApprover([True]), sleep=lambda s: None)
    # Orchestrator plus two workers
    out = run_multi_agent(T1001, BillingStore(WORK / "multi.json"), ScriptedApprover([True]))
    orch, orch_llm = out["orchestrator"]
    body = T1001.split("\n\n", 1)[1]
    print("multi-agent final:", orch.final[:110], "...")
    print(f"\n{'context':13} {'calls':>5} {'input':>6} {'output':>6} {'static prefix':>13} {'ticket body re-sent':>19}")
    contexts = [("single agent", single, single_llm), ("orchestrator", orch, orch_llm)] + \
               [(f"{k} worker", s, llm) for k, s, llm in out["workers"]]
    totals = {"single": 0, "multi": 0}
    for name, st, llm in contexts:
        print(f"{name:13} {st.llm_calls:>5} {st.input_tokens:>6} {st.output_tokens:>6} "
              f"{static_prefix_tokens(llm):>13} {duplicated_tokens(llm, body):>19}")
        totals["single" if name == "single agent" else "multi"] += st.input_tokens + st.output_tokens
    briefs = [c["args"]["brief"] for e in map(json.loads, (WORK / "orchestrator.jsonl").read_text().splitlines())
              if e["type"] == "llm_call" for c in e["tool_calls"]]
    reports = [s.final for _, s, _ in out["workers"]]
    comm = sum(count_tokens(b) for b in briefs) + sum(count_tokens(r) for r in reports)
    print(f"total tokens: single {totals['single']}, multi {totals['multi']} "
          f"({totals['multi'] / totals['single']:.2f}x); briefs plus reports: {comm} tokens")
    compounding_table()
    print()
    parallel_exploration()

Code explained

  • In simple words: one manager agent that hands out two jobs to specialist agents, an itemized token bill for the whole team, the math of chained failures, and a test of whether three searchers beat one.
  • What happens:
    • AccountWorker, BillingWorker, Orchestrator are scripted policies (not models) for the three roles. Each reads only its own context: the billing worker never sees the orchestrator's conversation, only its brief.
    • run_multi_agent wraps each worker as a Tool (delegate_account, delegate_billing) whose function calls run_agent with the worker's system prompt and tool subset, then returns the worker's final line. The orchestrator is itself a run_agent loop with those two tools. The billing worker gets the same flaky get_invoice and the same approval gate; subagents do not get to skip safety.
    • static_prefix_tokens counts the system prompt, first user message, and tool definitions that are re-sent on every call of a context. duplicated_tokens counts how many tokens of the customer's ticket text were re-sent across all calls of a context. Together they measure context duplication.
    • Briefs and reports are the communication cost: text written by one agent only so another can read it.
    • compounding_table computes the chance that a chain of k steps (or agents) all succeed at per-step reliability p, analytically as p to the power k and by simulation, and then with a checker that catches 80 percent of failures and allows one retry.
    • parallel_exploration runs three retrieval "explorers" per ticket (full text, subject only, body only) over the 62 answerable tickets, in threads, and compares three ways of aggregating their answers. Each explorer sleeps 20 ms to stand in for a worker's network wait; that part is simulated. The hit counts are real KBSearch results.
    • wilson gives a 95 percent confidence interval for a proportion, because 62 tickets is a small sample.
  • Comes out: (ScriptedLLM-driven for the agents; retrieval hits are real)
text
  multi-agent final: Hi, sorry about the double charge. Our billing team confirmed it: REFUND RF-00001 288.0 USD on INV-2026-004512 ...

  context       calls  input output static prefix ticket body re-sent
  single agent      8   9362    377          6296                 264
  orchestrator      3    992    174           741                  99
  account worker     2    252     37           242                   0
  billing worker     5   3687    141          2505                 165
  total tokens: single 9739, multi 5283 (0.54x); briefs plus reports: 117 tokens

  P(all k steps succeed) = p ** k        (analytic | simulated, 20,000 trials)
     p    k=1     k=3     k=5     k=10
    0.99 0.990|0.990 0.970|0.970 0.951|0.952 0.904|0.906
    0.95 0.950|0.950 0.857|0.857 0.774|0.776 0.599|0.601
    0.90 0.900|0.900 0.729|0.736 0.590|0.591 0.349|0.351
  with a checker (recall 0.8) and one retry per step: p=0.9 becomes 0.972; k=5 chain: 0.590 -> 0.868

  sequential wall time for 62 tickets: 4.37 s
  parallel wall time for 62 tickets: 1.43 s
    full_text        top-1 52/62 = 0.84  (95% CI 0.73 to 0.91)
    subject          top-1 42/62 = 0.68  (95% CI 0.55 to 0.78)
    body             top-1 49/62 = 0.79  (95% CI 0.67 to 0.87)
    fused_rrf        top-1 48/62 = 0.77  (95% CI 0.66 to 0.86)
    fused_max_score  top-1 52/62 = 0.84  (95% CI 0.73 to 0.91)
    any_explorer     top-1 53/62 = 0.85  (95% CI 0.75 to 0.92)

Communication cost and context duplication

Read the token table carefully, because it runs against the usual story. The orchestrator plus two workers used 5,283 tokens against 9,739 for the single agent (0.54x). Where did the saving come from? The single agent resends all seven tool definitions and the whole growing history on every call (6,296 tokens of static prefix alone). The workers each carry only their own two or four tools, and the billing worker's context never contains the account lookup. Briefs plus reports cost only 117 tokens. Splitting a task can shrink per-call context when the pieces are genuinely independent.

The bill moved elsewhere. Count the sequential model calls: 3 orchestrator calls plus 2 account calls plus 5 billing calls is 10 round trips on the critical path, against 8 for the single agent. With the Part 4 latency assumptions that is slower, not faster, unless the workers run in parallel. And the ticket text was still re-sent: 99 tokens across the orchestrator's 3 calls and 165 across the billing worker's 5 (the single agent re-sent it 8 times, 264 tokens). Every context that needs the facts pays for them again, on every call.

Now compare with a published data point, measured on real research tasks rather than a scripted refund. Anthropic's "How we built our multi-agent research system" (June 13, 2025) reports that agents "typically use about 4x more tokens than chat interactions, and multi-agent systems use about 15x more tokens than chats", and that a multi-agent setup with a Claude Opus 4 lead and Claude Sonnet 4 subagents "outperformed single-agent Claude Opus 4 by 90.2%" on their internal research evaluation. On BrowseComp, they found "token usage by itself explains 80% of the variance" in performance. The lesson is not that multi-agent is cheap or expensive in general. It is that multi-agent is a way to spend more tokens in parallel, and it pays when more exploration buys more quality. Our refund task has nothing to explore, so the only effect we could see was the smaller contexts.

Parallel exploration and aggregation

The retrieval experiment is the honest version of "run several explorers and pick the best". Three explorers per ticket, 62 answerable tickets:

  • Full text alone finds the gold article at rank 1 for 52 of 62 (0.84, 95 percent CI 0.73 to 0.91).
  • Fusing the three with reciprocal rank fusion scores 48 of 62: lower, most likely because the weaker subject-only explorer (42 of 62) pulls wrong articles up in the fused ranking.
  • Taking the single highest-scoring hit scores 52, the same as full text alone.
  • Even an oracle that picks whichever explorer was right ("any explorer") reaches only 53 of 62. The explorers make the same mistakes (mostly the non-English tickets, Module 7), so there is almost nothing for aggregation to recover.

Running them in threads cut wall time from about 4 s to about 1.4 s (timings vary from run to run), but that speedup comes from the simulated 20 ms waits overlapping. Real LLM workers are network-bound, so the same overlap applies; CPU-bound Python work would not speed up this way. So parallel exploration here tripled the work for no measurable gain: the aggregated variants land between 48 and 53 of 62, and all of them sit inside the confidence interval of full text alone. It pays only when the explorers are diverse (they fail on different inputs) and the aggregator can verify which answer is right.

Failure compounding

If each step of a chain succeeds with probability p and failures are independent, the whole chain of k steps succeeds with probability p to the power k. The table above shows how fast that falls: at 95 percent per step, a 10-step chain succeeds 60 percent of the time; at 90 percent, 35 percent. The simulation matches the formula to within sampling noise (for example 0.349 analytic against 0.351 simulated), which is a useful check that your intuition about "95 percent is pretty good" is wrong for chains.

Multi-agent systems add links to the chain: every brief can drop a fact, every report can compress away a caveat, and the orchestrator can combine two reports that made incompatible assumptions. The fix is the same as for a single agent: put checks between steps. With a checker that catches 80 percent of failures and one retry, per-step reliability rises from 0.90 to 0.972 and a 5-step chain from 0.590 to 0.868. Checks that use external signals (did get_invoice actually show two charges?) are what make that recall possible; Module 5 showed that self-critique without external signal often is not.

When multi-agent is genuinely better

Cognition's "Don't Build Multi-Agents" (Walden Yan, June 12, 2025) argued for single-threaded agents with two principles: "Share context, and share full agent traces, not just individual messages", and "Actions carry implicit decisions, and conflicting decisions carry bad results". Its example: parallel subagents asked to build pieces of a Flappy Bird clone produce a Super Mario style background and a mismatched bird, which the final agent cannot reconcile. Anthropic's post agrees on the boundary: domains "that require all agents to share the same context or involve many dependencies between agents are not a good fit for multi-agent systems today", noting that most coding tasks have fewer truly parallelizable parts than research.

Cognition's follow-up, "Multi-Agents: What's Actually Working" (April 22, 2026), narrows the claim rather than reversing it: multi-agent systems "work best today when writes stay single-threaded and the additional agents contribute intelligence rather than actions." What worked for them: review agents, read-only subagents that gather context, and a manager that delegates pieces and coordinates. What still did not: parallel writers making conflicting decisions.

For Brightlane, that points to a clear design: one agent owns every action on a ticket (refunds, notes, replies), and any extra agents are read-only helpers (search the help center in several languages, summarize the account history) whose reports it reads.

SituationUse thisWhy
Sequential task with dependent steps (most tickets)One agent loopShared context, one decision maker, fewer round trips
Broad research or search that overflows one context and splits cleanlyOrchestrator plus parallel read-only workersFresh contexts and parallel time; published results show quality gains for research
Several agents would take actions on the same objectDo not; keep writes single-threadedConflicting implicit decisions (Cognition)
A second opinion on a risky outputA reviewer agent that reads and reports, never actsAdds intelligence, not actions
A worker's tool set is large and unrelated to the main taskA subagent with a narrow tool listSmaller definitions and better tool selection