CourseLarge Language Models · Module 8: Agents · part 39 of 80
Part 39 · Module 8: Agents

Part 4: Is the agent worth it?

8 min read·22 Sept 2026

The same ticket, two designs

Now the measurement promised in Part 1. run_workflow handles T-1001 on a fixed path: regex for the email and invoice id, get_account and get_invoice through the same execute_tool (so it gets the same retries and approval gate), a code check for a duplicate charge, the refund, and then exactly one model call to write the reply as a DraftReply (Module 6's schema). The agent is run_agent from Part 2. Both use the same flaky billing API and the same approver.

examples/m08_workflow_vs_agent.py

python
"""Module 8: the same duplicate-charge ticket as a fixed workflow and as an agent.

Token counts are real counts (o200k_base) of the real prompts each design sends.
The replies come from ScriptedLLM policies, so the counts show what each DESIGN
costs, not how well a model performs. Latency is modeled from stated assumptions.
"""
from __future__ import annotations

import json
import random
import statistics
import time
from pathlib import Path

from m08_agent import (EMAIL_RE, INVOICE_RE, T1001, AgentConfig, Approver, BillingStore, ScriptedApprover,
                       SupportPolicy, Tracer, execute_tool, make_tools, run_agent)
from supportdesk.data import get_article
from supportdesk.llm import Usage
from supportdesk.pricing import cost_usd
from supportdesk.schemas import DraftReply
from supportdesk.stand_in import ScriptedLLM

WORK = Path("runs/m08/compare")
WORK.mkdir(parents=True, exist_ok=True)
for old in WORK.glob("*"):
    old.unlink()

DRAFT_SYSTEM = ("You write Brightlane support replies. Use only the facts and the policy given. "
                "Return JSON with keys reply, cited_articles, confidence.")


def run_workflow(ticket_text: str, task_id: str, tools: dict, chat_fn, approver: Approver,
                 tracer: Tracer) -> dict:
    """Fixed path: code does every lookup and check; the model writes only the customer reply."""
    config = AgentConfig(backoff_base_s=0.05)
    rng = random.Random(0)
    email, invoice_id = EMAIL_RE.search(ticket_text).group(0), INVOICE_RE.search(ticket_text).group(0)
    account = execute_tool(tools["get_account"], {"email": email}, approver, tracer, 1, config, rng)
    invoice = execute_tool(tools["get_invoice"], {"invoice_id": invoice_id}, approver, tracer, 2, config, rng)
    if not (account["ok"] and invoice["ok"]) or invoice["result"]["account_id"] != account["result"]["account_id"]:
        execute_tool(tools["escalate_to_human"], {"ticket_id": task_id, "reason": "lookup failed or mismatch"},
                     approver, tracer, 3, config, rng)
        return {"route": "escalated", "llm_calls": 0, "usage": Usage()}
    inv = invoice["result"]
    charges = inv["charges"]
    duplicate = [c for c in charges[1:] if (c["amount"], c["date"]) == (charges[0]["amount"], charges[0]["date"])]
    if not duplicate or inv["overpaid"] <= 0:
        execute_tool(tools["escalate_to_human"], {"ticket_id": task_id, "reason": "no duplicate found"},
                     approver, tracer, 3, config, rng)
        return {"route": "escalated", "llm_calls": 0, "usage": Usage()}
    refund = execute_tool(tools["issue_refund"], {"invoice_id": invoice_id, "charge_id": duplicate[0]["charge_id"],
                                                  "amount_usd": inv["overpaid"], "reason": "duplicate charge"},
                          approver, tracer, 3, config, rng)
    if not refund["ok"]:
        execute_tool(tools["escalate_to_human"], {"ticket_id": task_id, "reason": refund["error"]},
                     approver, tracer, 4, config, rng)
        return {"route": "escalated", "llm_calls": 0, "usage": Usage()}
    policy = get_article("billing-refunds")
    facts = {"refund_id": refund["result"]["refund_id"], "amount_usd": inv["overpaid"], "invoice_id": invoice_id}
    messages = [{"role": "system", "content": DRAFT_SYSTEM},
                {"role": "user", "content": f"Policy ({policy.id}):\n{policy.body}\n\nFacts: {json.dumps(facts)}\n\nTicket:\n{ticket_text}"}]
    result = chat_fn(messages, response_format={"type": "json_object"})
    draft = DraftReply.model_validate_json(result.text)
    execute_tool(tools["add_internal_note"], {"ticket_id": task_id, "note": f"Workflow refunded {facts}"},
                 approver, tracer, 5, config, rng)
    tracer.event("final", 6, reply=draft.reply)
    return {"route": "refunded", "llm_calls": 1, "usage": result.usage, "reply": draft.reply}


def draft_responder(messages, kwargs) -> str:
    facts = json.loads(messages[-1]["content"].split("Facts: ")[1].split("\n")[0])
    return json.dumps({"reply": f"Hi, sorry about the double charge. We refunded the duplicate payment of "
                                f"{facts['amount_usd']:.2f} USD on invoice {facts['invoice_id']} (refund {facts['refund_id']}). "
                                "It goes back to your original card within 5 to 10 business days.",
                       "cited_articles": ["billing-refunds"], "confidence": "high"})


# Latency model (ASSUMPTIONS, not measurements): replace with your own numbers from Module 3.
CALL_OVERHEAD_MS = 600      # network plus time to first token, per sequential LLM call
OUTPUT_TOK_PER_S = 150      # decode speed
PREFILL_TOK_PER_S = 5000    # input processing speed


def modeled_latency_ms(calls: int, input_tokens: int, output_tokens: int) -> float:
    return calls * CALL_OVERHEAD_MS + input_tokens / PREFILL_TOK_PER_S * 1000 + output_tokens / OUTPUT_TOK_PER_S * 1000


def row(name: str, calls: int, usage: Usage, wall_ms: float) -> dict:
    return {"design": name, "llm_calls": calls, "input": usage.input_tokens, "output": usage.output_tokens,
            "usd_oss120b": cost_usd(usage, "openai/gpt-oss-120b"), "usd_flash": cost_usd(usage, "gemini-3.5-flash"),
            "latency_model_s": modeled_latency_ms(calls, usage.input_tokens, usage.output_tokens) / 1000,
            "wall_ms": wall_ms}


if __name__ == "__main__":
    # 1) Fixed workflow
    from supportdesk.tokens import count_tokens
    count_tokens("warm up the tokenizer so its load time is not charged to either design")
    tools = make_tools(BillingStore(WORK / "workflow_billing.json"), flaky_failures=2)
    tracer = Tracer(WORK / "workflow_trace.jsonl", "T-1001")
    started = time.perf_counter()
    wf = run_workflow(T1001, "T-1001", tools, ScriptedLLM(responder=draft_responder),
                      ScriptedApprover([True]), tracer)
    wf_ms = (time.perf_counter() - started) * 1000
    print("workflow route:", wf["route"], "| reply:", wf["reply"][:70], "...")

    # 2) Agent, same ticket, same flaky tool
    tools = make_tools(BillingStore(WORK / "agent_billing.json"), flaky_failures=2)
    started = time.perf_counter()
    st = run_agent(T1001, "T-1001", tools, ScriptedLLM(responder=SupportPolicy()),
                   approver=ScriptedApprover([True]), trace_path=WORK / "agent_trace.jsonl")
    ag_ms = (time.perf_counter() - started) * 1000
    print("agent status:", st.status, "| reply:", st.final[:70], "...")

    rows = [row("workflow", wf["llm_calls"], wf["usage"], wf_ms),
            row("agent", st.llm_calls, Usage(st.input_tokens, st.output_tokens), ag_ms)]
    print(f"\n{'design':9} {'calls':>5} {'input':>6} {'output':>6} {'usd@oss-120b':>13} {'usd@3.5-flash':>14} "
          f"{'latency model s':>15} {'plumbing ms':>11}")
    for r in rows:
        print(f"{r['design']:9} {r['llm_calls']:>5} {r['input']:>6} {r['output']:>6} {r['usd_oss120b']:>13.6f} "
              f"{r['usd_flash']:>14.6f} {r['latency_model_s']:>15.1f} {r['wall_ms']:>11.0f}")
    ratio = (rows[1]["input"] + rows[1]["output"]) / (rows[0]["input"] + rows[0]["output"])
    print(f"agent/workflow tokens: {ratio:.1f}x | per 1,000 tickets at gemini-3.5-flash: "
          f"workflow {rows[0]['usd_flash'] * 1000:.2f} USD, agent {rows[1]['usd_flash'] * 1000:.2f} USD")

    # 3) Predictability: 50 agent runs with seeded detours (simulated variation, not a model)
    calls, tokens = [], []
    for seed in range(50):
        store = BillingStore(WORK / f"noisy_{seed}.json")
        s = run_agent(T1001, "T-1001", make_tools(store, flaky_failures=0), ScriptedLLM(responder=SupportPolicy(noise=0.3, seed=seed)),
                      approver=ScriptedApprover([True]), sleep=lambda x: None)
        assert s.status == "done", s.status
        calls.append(s.llm_calls)
        tokens.append(s.input_tokens + s.output_tokens)
    q = statistics.quantiles(tokens, n=20)
    print(f"\nagent over 50 noisy runs: LLM calls min {min(calls)} median {statistics.median(calls):.0f} max {max(calls)}; "
          f"tokens median {statistics.median(tokens):.0f}, p95 {q[18]:.0f}, max {max(tokens)}; "
          f"stdev {statistics.stdev(tokens):.0f}")
    print("workflow over any number of runs: LLM calls always 1, tokens always", rows[0]["input"] + rows[0]["output"])

Code explained

  • In simple words: run T-1001 both ways, count what each costs, then run the agent 50 more times with seeded detours to see how much its cost moves.
  • What happens:
    • run_workflow is a plain function: every decision is an if in code. The only model call is the reply, with the policy text and the facts (refund id, amount, invoice) in the prompt and response_format={"type": "json_object"}. Its reply is validated with DraftReply.model_validate_json. Any surprise (lookup failure, account mismatch, no duplicate, refund rejected) routes to escalate_to_human.
    • draft_responder is the scripted stand-in for that one call.
    • modeled_latency_ms is a latency model, not a measurement: 600 ms of overhead per sequential call, 5,000 input tokens per second of prefill, 150 output tokens per second of decode. These are assumptions; replace them with the time to first token and decode speed you measured for your provider in Module 3.
    • row prices each design with cost_usd at two price points from pricing.py.
    • The tokenizer is warmed up before timing so its load time is not charged to whichever design runs first. "Plumbing ms" is real wall time of the Python around the model calls; most of it is the two real backoff sleeps (about 0.06 s and 0.12 s).
    • The last block runs the agent 50 times with SupportPolicy(noise=0.3, seed=...). The detours are simulated variation, a stand-in for a real model sometimes searching again or re-reading. It shows the shape of the effect, not its size for any real model.
  • Comes out: (ScriptedLLM-driven; token counts real, latency modeled)
text
  workflow route: refunded | reply: Hi, sorry about the double charge. We refunded the duplicate payment o ...
  agent status: done | reply: Hi, sorry about the double charge. We refunded the duplicate payment o ...

  design    calls  input output  usd@oss-120b  usd@3.5-flash latency model s plumbing ms
  workflow      1    242     71      0.000079       0.001002             1.1         188
  agent         8   9362    377      0.001631       0.017436             9.2         226
  agent/workflow tokens: 31.1x | per 1,000 tickets at gemini-3.5-flash: workflow 1.00 USD, agent 17.44 USD

  agent over 50 noisy runs: LLM calls min 8 median 10 max 11; tokens median 13425, p95 15581, max 15668; stdev 1632
  workflow over any number of runs: LLM calls always 1, tokens always 313

Reading the comparison

MeasureWorkflowAgentWhere the number comes from
LLM calls18Counted by the loop
Tokens (input + output)3139,739o200k_base counts of the prompts actually sent
Cost per ticket, gpt-oss-120b0.000079 USD0.001631 USDpricing.py math
Cost per ticket, gemini-3.5-flash0.0010 USD0.0174 USDpricing.py math
Latencyabout 1.1 sabout 9.2 sModeled from stated assumptions
PredictabilitySame path, same tokens every run8 to 11 calls; tokens stdev 1,632 in the simulated runs50 seeded runs

The agent costs about 31 times the tokens for the same outcome. Two reasons, both visible in the data:

  1. Tool definitions ride along on every call. The seven tool schemas are 615 tokens. Over 8 calls that is 4,920 of the 9,362 input tokens, over half. Module 2's prompt caching would discount this stable prefix on providers that support it; it does not remove the round trips.
  2. Context grows every step. Call 1 sent 787 tokens; the last call sent the whole history. Total input grows roughly with the square of the number of steps.

Is the difference within noise? No: the workflow's number is exact and the agent's minimum across 50 runs (8 calls, about 9,700 tokens) is still about 30 times higher. What is uncertain is how a real model's path length varies, which is why you should rerun the 50-run block with chat and your key before you quote a p95 to anyone.

For duplicate charges, the workflow wins on every measure we have, and it is easier to certify: a compliance reviewer can read run_workflow in five minutes. The agent earns its cost only on tickets where you cannot write the path down, and even then the right move is often to turn the paths the agent discovers in production (read them from the traces) into workflows, keeping the agent for the long tail.

SituationUse thisWhy
High-volume ticket type with a known procedure (duplicate charges, invoice copies)Workflow31x fewer tokens here, fixed latency, auditable
Mixed queue where most tickets fit known proceduresRouter to workflows, agent as fallback for the restPays agent prices only on the long tail
Rare, open-ended investigations (why did automations stop?)Agent with guards and a budgetThe path is unknowable in advance; cap the cost instead
You are not sure whichLog agent trajectories for a few weeks, then promote common paths into workflows