CourseLarge Language Models · Module 5: Reasoning and Advanced Prompting · part 24 of 80
Part 24 · Module 5: Reasoning and Advanced Prompting

Part C: Decomposition

19 min read·22 Sept 2026

Least-to-most: break the problem down, then solve in order

Least-to-most prompting (Zhou et al., "Least-to-Most Prompting Enables Complex Reasoning in Large Language Models", ICLR 2023) runs in two stages: first ask the model to reduce the problem to simpler sub-questions, then answer them one at a time, feeding each answer into the next prompt. On the SCAN compositional-generalization benchmark, the authors report at least 99 percent accuracy with 14 exemplars using code-davinci-002, against 16 percent for chain-of-thought prompting. The gain comes from problems harder than the examples: each sub-question is easy even when the whole is not.

For refunds, the natural sub-questions are the policy's own checks. Because the policy is ordered, some early answers already settle the case (a member cannot request a refund, whatever else is true), so the loop can stop early and save calls.

examples/m05_least_to_most.py:

python
"""Part C: least-to-most prompting. First ask for the sub-questions, then answer them in order.

The model calls go through a scripted stand-in (plumbing only): it returns a
fixed decomposition and answers each sub-question from the billing record, so
you can see the message flow and count calls without an API key.
"""
import re

from examples.m05_helpers import CASES, TODAY, case_from_messages, decide, policy_text
from supportdesk.stand_in import ScriptedLLM

DECOMPOSE = ("To decide this refund request by the policy, what simpler questions must be answered first? "
             "List them in the order they should be answered, one per line, numbered.\n\n"
             "<policy>\n{policy}\n</policy>\n<ticket>\n{ticket}\n</ticket>")
SUBQUESTION = ("<policy>\n{policy}\n</policy>\n<ticket>\n{ticket}\n</ticket>\nToday is {today}.\n"
               "Answers so far:\n{solved}\nNow answer only this question in one short line: {question}")


def least_to_most(chat_fn, ticket: str, stop_rule=None) -> tuple[list[tuple[str, str]], int]:
    """Stage 1 decomposes; stage 2 answers each sub-question with all earlier answers in context.

    stop_rule(solved) may return True to end early once the answers settle the outcome.
    Returns the (question, answer) pairs and the number of model calls made.
    """
    reply = chat_fn([{"role": "user", "content": DECOMPOSE.format(policy=policy_text(), ticket=ticket)}])
    questions = [re.sub(r"^\d+[.)]\s*", "", line).strip() for line in reply.text.splitlines() if re.match(r"^\d+[.)]", line)]
    solved: list[tuple[str, str]] = []
    calls = 1
    for question in questions:
        context = "\n".join(f"Q: {q}\nA: {a}" for q, a in solved) or "(none yet)"
        prompt = SUBQUESTION.format(policy=policy_text(), ticket=ticket, today=TODAY, solved=context, question=question)
        answer = chat_fn([{"role": "user", "content": prompt}]).text.strip()
        calls += 1
        solved.append((question, answer))
        if stop_rule and stop_rule(solved):
            break
    return solved, calls


SUBQUESTIONS = ["Is the requester a workspace owner or billing admin?", "Is this a duplicate charge?",
                "Is the plan billed annually?", "How many days ago was the purchase or renewal?"]


def stand_in(messages, kwargs):
    content = messages[-1]["content"]
    if content.startswith("To decide"):
        return "\n".join(f"{i}. {q}" for i, q in enumerate(SUBQUESTIONS, 1))
    f = case_from_messages(messages).facts
    question = content.rsplit("question in one short line: ", 1)[1]
    return {SUBQUESTIONS[0]: "yes" if f.role in ("owner", "billing_admin") else "no",
            SUBQUESTIONS[1]: "yes" if f.duplicate else "no",
            SUBQUESTIONS[2]: "yes" if f.billing == "annual" else "no",
            SUBQUESTIONS[3]: str((TODAY - f.charged_on).days)}[question]


def settled(solved) -> bool:
    """The policy is ordered, so some early answers already decide the case."""
    answers = dict(solved)
    return (answers.get(SUBQUESTIONS[0]) == "no" or answers.get(SUBQUESTIONS[1]) == "yes"
            or answers.get(SUBQUESTIONS[2]) == "no")


llm = ScriptedLLM(responder=stand_in)
solved, calls = least_to_most(llm, CASES[3].text)
print(f"{CASES[3].id}: {CASES[3].text}")
for q, a in solved:
    print(f"  Q: {q}\n  A: {a}")
print(f"  calls: {calls}; the last prompt carried {len(solved) - 1} earlier answers\n")

for rule_name, rule in (("answer every sub-question", None), ("stop once settled", settled)):
    total = sum(least_to_most(llm, case.text, rule)[1] for case in CASES)
    print(f"{rule_name:<26} {total} calls for {len(CASES)} cases ({total / len(CASES):.2f} per case)")
print("gold decision for", CASES[3].id, "=", decide(CASES[3].facts))

Code explained

  • In simple words: one call writes the list of questions; then one call per question answers it, seeing all earlier answers.
  • What happens: least_to_most sends the decomposition prompt, parses the numbered lines into sub-questions, and answers each with a prompt that carries the policy, the ticket, today's date, and every earlier question and answer. stop_rule can end the loop once the outcome is settled. The model here is a ScriptedLLM stand-in (plumbing only): it returns the four policy checks as the decomposition and answers each from the billing record. settled encodes when the ordered policy has already decided: not an owner or billing admin, a duplicate, or not annual.
  • Comes out: for R04 (a member reporting a duplicate) the full loop makes 5 calls, and the last prompt carries 3 earlier answers. Over all 20 scenarios, answering every sub-question costs 100 calls; stopping once settled costs 82. The accuracy question (does a real model answer sub-questions more reliably than the whole?) needs a real model; the call and context growth shown here is what you pay for it.

Notice what least-to-most costs: 4 to 5 calls per case instead of 1, and each later prompt repeats the policy and the ticket. It is the right tool when problems are harder than your examples and a single prompt keeps skipping a step; it is the wrong tool when the sub-questions are fixed and known in advance, because then you can write them into a scaffold (Part A) or a chain (next) and skip the decomposition call.

Prompt chaining with typed intermediate results

A prompt chain is a fixed pipeline of specialized calls, each with one job, where the output of one step is the input of the next. The important word is typed: every intermediate result is validated against a schema before the next step sees it, so a bad step stops the chain instead of passing garbage downstream. For refunds the chain is:

Typed Refund Pipeline

Step 2 is not a model call at all. Once the facts are typed, the policy is a function, and a function applies rules in order perfectly every time. That is the most useful decomposition lesson in this module: split the task so that the model does the parts only a model can do (reading free text in five languages, writing a friendly reply) and code does the rest.

examples/m05_chain.py

python
"""Part C: a prompt chain with typed intermediate results: extract facts, apply policy, draft reply.

Stage 1 and stage 3 are model calls; here they run on scripted stand-ins
(plumbing only). The stage-1 stand-in uses the regular-expression extractor, so
its successes and failures on the 20 cases are real for that extractor, not for
any model. Stage 2 is plain code, because the policy is a rule.
"""
import json
from datetime import date
from typing import Literal

from pydantic import BaseModel, ConfigDict, ValidationError

from examples.m05_helpers import CASES, Facts, RefundCase, case_from_messages, decide, extract_facts_regex, policy_text
from supportdesk.schemas import DraftReply
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages


class RefundFacts(BaseModel):
    """Stage 1 output: the facts the policy needs, read from the ticket."""
    model_config = ConfigDict(extra="forbid")
    role: Literal["owner", "billing_admin", "admin", "member"]
    billing: Literal["monthly", "annual"]
    charged_on: date
    duplicate: bool


class RefundDecision(BaseModel):
    """Stage 2 output: the decision and the rule that produced it."""
    decision: Literal["refund_duplicate", "refund_full", "no_refund", "needs_owner"]
    facts: RefundFacts


EXTRACT = ("Read the ticket and return JSON with keys role (owner, billing_admin, admin, member), "
           "billing (monthly, annual), charged_on (YYYY-MM-DD), duplicate (true/false). "
           "Use null for anything the ticket does not state.\n<ticket>\n{ticket}\n</ticket>")
DRAFT = ("Write a short, friendly reply in the customer's language. The decision is final: {decision}. "
         "Explain it using only this policy:\n<policy>\n{policy}\n</policy>\n<ticket>\n{ticket}\n</ticket>\n"
         'Return JSON: {{"reply": str, "cited_articles": ["billing-refunds"], "confidence": "low"|"medium"|"high"}}')


def extract_stand_in(messages, kwargs):
    facts = extract_facts_regex(case_from_messages(messages).text)
    if facts is None:
        return json.dumps({"role": None, "billing": None, "charged_on": None, "duplicate": None})
    return json.dumps({"role": facts.role, "billing": facts.billing, "charged_on": str(facts.charged_on),
                       "duplicate": facts.duplicate})


REPLIES = {"refund_duplicate": "We are refunding the duplicate charge in full; expect it in 5 to 10 business days.",
           "refund_full": "Your annual plan is within the 14-day window, so we will refund it in full.",
           "no_refund": "This charge is outside what our refund policy covers, so we cannot refund it.",
           "needs_owner": "Only a workspace owner or billing admin can request refunds. Please ask them to contact us."}


def draft_stand_in(messages, kwargs):
    decision = messages[-1]["content"].split("The decision is final: ", 1)[1].split(".", 1)[0]
    return json.dumps({"reply": REPLIES[decision], "cited_articles": ["billing-refunds"], "confidence": "high"})


def run_chain(case: RefundCase, extract_llm, draft_llm) -> dict:
    """Run the three stages; any stage that fails validation stops the chain and escalates."""
    trace = {"case": case.id, "calls": 0}
    raw = extract_llm([{"role": "user", "content": EXTRACT.format(ticket=case.text)}])
    trace["calls"] += 1
    try:
        facts = RefundFacts.model_validate_json(raw.text)
    except ValidationError as err:
        trace.update(outcome="escalate", stage="extract", error=f"{err.error_count()} invalid fields")
        return trace
    decision = RefundDecision(decision=decide(Facts(**facts.model_dump())), facts=facts)
    raw = draft_llm([{"role": "user", "content": DRAFT.format(decision=decision.decision, policy=policy_text(),
                                                                ticket=case.text)}])
    trace["calls"] += 1
    try:
        draft = DraftReply.model_validate_json(raw.text)
    except ValidationError as err:
        trace.update(outcome="escalate", stage="draft", error=f"{err.error_count()} invalid fields")
        return trace
    trace.update(outcome="draft", decision=decision.decision, reply=draft.reply)
    return trace


if __name__ == "__main__":
    extract_llm, draft_llm = ScriptedLLM(responder=extract_stand_in), ScriptedLLM(responder=draft_stand_in)
    traces = [run_chain(case, extract_llm, draft_llm) for case in CASES]
    for t in traces[:3] + [traces[11]]:
        print(t)
    drafted = [t for t in traces if t["outcome"] == "draft"]
    right = sum(t["decision"] == c.gold for t, c in zip(traces, CASES) if t["outcome"] == "draft")
    print(f"\ndrafted {len(drafted)}, escalated {len(traces) - len(drafted)}; "
          f"decisions right {right}/{len(drafted)} of drafted; model calls {sum(t['calls'] for t in traces)}")
    for name, llm in (("extract", extract_llm), ("draft", draft_llm)):
        tokens = sum(count_messages(call["messages"]) for call in llm.calls)
        print(f"  {name:<8} {len(llm.calls):>2} calls, {tokens:>5} input tokens")

    # A failure you will see with real models: a date that is not ISO format.
    bad = '{"role": "owner", "billing": "annual", "charged_on": "3 September", "duplicate": false}'
    try:
        RefundFacts.model_validate_json(bad)
    except ValidationError as err:
        first = err.errors()[0]
        print(f"\nvalidation caught: field {first['loc'][0]!r}: {first['msg']} (input {first['input']!r})")

Code explained

  • In simple words: read the facts, decide by rule, write the reply, and stop at the first step whose output does not validate.
  • What happens:
    • RefundFacts and RefundDecision are pydantic models: the typed hand-offs between steps. extra="forbid" rejects invented fields; Literal types reject unknown roles; date rejects "3 September".
    • EXTRACT asks for JSON with null for anything the ticket does not state, which gives the model an honest way to say "not in the ticket".
    • DRAFT gives the decision as final and asks for a DraftReply (the canonical schema from supportdesk.schemas, introduced in Module 6).
    • extract_stand_in and draft_stand_in are ScriptedLLM responders (plumbing only). The extractor uses extract_facts_regex from the helpers, so its successes and failures are real for that regex extractor, not for any model.
    • run_chain records a trace per case: calls made, outcome, and the stage and error count on failure.
  • Comes out: 14 cases get a draft, all 14 with the right decision; 6 escalate at the extract step with 4 invalid fields (the regex extractor could not read them: R01 never says monthly or annual, R12 is Spanish). Nothing wrong reached the draft step. The chain spent 34 calls, and the draft step's prompts are larger than the extract step's because they carry the policy. The last line is the failure you will meet with real models: a date written as "3 September" is rejected by the schema, not silently parsed into the wrong year.

When a real model runs step 1, read the escalations first. They cluster: all non-English, or all relative dates ("last Monday"), or all tickets that do not say the billing period. Each cluster is a specific fix (add today's date to the extract prompt, allow billing to be looked up from the billing system instead of read from the ticket) that you can measure on the same 20 cases.

SituationUse thisWhy
Steps are known and fixed; some are pure rulesPrompt chain with typed hand-offs, rules in codeEach step is testable; code applies rules exactly
Steps depend on the problem and are not known aheadLeast-to-most, or an agent (Module 8)The model has to plan the breakdown
Everything fits one prompt and one call is accurate enoughSingle call with a scaffoldFewer calls, lower latency, one place to debug
One step needs a stronger model than the othersChain, with a different model per stepPay for the expensive model only where it matters

Router prompts and conditional paths

A router is a cheap first call that picks which path a request takes. Different paths can have different prompts, models, and costs, and one path can be "a person handles this". For Brightlane, three routes cover the week: refund requests go down the refund chain (2 more model calls), help-center questions get a retrieval-grounded answer (1 more call, Module 7), and legal, security, or unanswerable tickets go to a person (no more calls).

.

examples/m05_router.py:

python
"""Part C: a router prompt that sends each ticket down one of three paths.

The router call runs on a scripted stand-in that applies keyword rules (plumbing
only). Its accuracy below is the accuracy of those rules against hand-assigned
routes, not of any model; rerun with --live to measure a real router.
"""
import json
import re
import sys
from collections import Counter

from examples.m05_helpers import wilson
from supportdesk.data import load_tickets
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages

ROUTES = {
    "refund_chain": "the customer asks for money back or reports being charged twice",
    "kb_answer": "a question the help center can answer (plans, invoices, login, exports, bugs, how-to)",
    "human": "legal, security, account ownership, anger, or anything the help center cannot answer",
}
ROUTER = ("You route Brightlane support tickets. Choose exactly one route:\n"
          + "\n".join(f"- {name}: {desc}" for name, desc in ROUTES.items())
          + '\nReturn JSON: {"route": "<route>", "reason": "<five words or fewer>"}')
CALLS_PER_PATH = {"refund_chain": 2, "kb_answer": 1, "human": 0}  # model calls after routing
REFUND_ASKS = {"T-1001", "T-1004", "T-1005", "T-1031", "T-1060"}  # hand-labelled: asks for money back


def gold_route(ticket) -> str:
    if ticket.id in REFUND_ASKS:
        return "refund_chain"
    return "kb_answer" if ticket.gold["answerable"] else "human"


def rule_router(messages, kwargs):
    text = messages[-1]["content"].lower()
    if re.search(r"refund|money back|charged .*twice|reembolso|dos veces", text):
        return json.dumps({"route": "refund_chain", "reason": "refund request"})
    if re.search(r"legal|gdpr|attack|owner left|garbage|sap|declined|don't recognize", text):
        return json.dumps({"route": "human", "reason": "needs a person"})
    return json.dumps({"route": "kb_answer", "reason": "help center question"})


def route(chat_fn, ticket) -> str:
    reply = chat_fn([{"role": "system", "content": ROUTER}, {"role": "user", "content": ticket.text}],
                    temperature=0.0, response_format={"type": "json_object"})
    try:
        choice = json.loads(reply.text).get("route")
    except json.JSONDecodeError:
        choice = None
    return choice if choice in ROUTES else "human"  # anything unexpected goes to a person


if __name__ == "__main__":
    if "--live" in sys.argv:
        from supportdesk.llm import chat as router_llm
    else:
        router_llm = ScriptedLLM(responder=rule_router)
    tickets = load_tickets()
    pairs = [(route(router_llm, t), gold_route(t)) for t in tickets]
    k = sum(p == g for p, g in pairs)
    lo, hi = wilson(k, len(pairs))
    print(f"routing agreement {k}/{len(pairs)}  95% CI {lo:.2f}-{hi:.2f}")
    print("confusion (predicted, gold): count")
    for (p, g), c in sorted(Counter(pairs).items()):
        print(f"  {p:<13} {g:<13} {c}")
    predicted = Counter(p for p, _ in pairs)
    downstream = sum(CALLS_PER_PATH[p] * c for p, c in predicted.items())
    router_tokens = sum(count_messages([{"role": "system", "content": ROUTER}, {"role": "user", "content": t.text}])
                        for t in tickets)
    print(f"paths taken: {dict(predicted)}")
    print(f"model calls: {len(tickets)} router + {downstream} downstream = {len(tickets) + downstream}; "
          f"router input tokens {router_tokens}")
    missed = [t.id for t, (p, g) in zip(tickets, pairs) if g == "refund_chain" and p != g]
    print(f"refund requests the router missed: {missed or 'none'}")

Code explained

  • In simple words: one short call per ticket picks a path; anything the code does not recognize goes to a person.
  • What happens: ROUTER lists the routes with one-line definitions and asks for JSON. route asks for JSON mode, parses the reply, and falls back to human on bad JSON or an unknown route: the safe default for a support desk. gold_route is a hand labelling for measurement: the five tickets that ask for money back go to the refund chain; otherwise the dataset's answerable flag decides. Offline, rule_router stands in for the model (plumbing only), so the agreement number is the accuracy of those keyword rules, not of a model. --live routes with the real model.
  • Comes out: the rules agree with the hand labels on 67 of 72 tickets. The confusion table shows where: one help-center question was sent to the refund chain, one to a person, and three tickets that need a person were sent to the knowledge base. No refund request was missed, which is the error that matters most here. The whole week costs 72 router calls plus 70 downstream calls, and the router prompts total 9,141 input tokens.

Two design points carry over to real routers. First, give every path a cost and count it: the router's job is partly to spend less, and the call count by path makes that visible. Second, decide which errors are expensive. Sending a refund request to the knowledge-base path gives a customer a generic answer about money; sending a question to a person costs a few minutes. Bias the router (and its fallback) toward the cheap mistake.

Map-reduce over long inputs

Map-reduce handles an input too long, or too varied, for one call: a map step processes each piece independently (one call per ticket), and a reduce step combines the small results. For Maya's weekly report, map turns each ticket into a short validated note, code counts the notes, and one reduce call writes prose over the counts and one-line issues. Counting is code, not a model, for the same reason step 2 of the chain is code.

examples/m05_map_reduce.py:

python
"""Part C: map-reduce over all 72 tickets to write a weekly report, counting calls and tokens.

Map and reduce calls run on scripted stand-ins (plumbing only): the map stand-in
copies the gold labels, so the counts below are right by construction. What is
real: the number of calls, the token counts of every prompt, and the cost math.
"""
import json
from collections import Counter

from pydantic import BaseModel

from supportdesk.data import load_tickets
from supportdesk.llm import Usage
from supportdesk.pricing import cost_usd
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages


class TicketNote(BaseModel):
    id: str
    category: str
    priority: str
    language: str
    issue: str


MAP = ("Summarize this support ticket as JSON with keys id, category, priority, language, "
       "issue (at most 12 words, English).\n<ticket id=\"{id}\">\n{text}\n</ticket>")
REDUCE = ("Here are this week's ticket counts (computed by code, do not recount) and one-line issues.\n"
          "Counts: {counts}\nIssues:\n{issues}\n"
          "Write three short paragraphs for the support lead: main themes, anything urgent, suggested fixes.")
SINGLE = ("Here are all of this week's support tickets. Write a weekly report for the support lead with counts "
          "by category and priority, main themes, anything urgent, and suggested fixes.\n\n{all_tickets}")


def map_stand_in(messages, kwargs):
    ticket_id = messages[-1]["content"].split('<ticket id="', 1)[1].split('"', 1)[0]
    t = BY_ID[ticket_id]
    return json.dumps({"id": t.id, "category": t.gold["category"], "priority": t.gold["priority"],
                       "language": t.language, "issue": t.subject})


def reduce_stand_in(messages, kwargs):
    lines = messages[-1]["content"].split("Issues:\n", 1)[1].split("\nWrite three", 1)[0].splitlines()
    return f"Themes across {len(lines)} issues: billing and login dominate. (scripted stand-in text)"


tickets = load_tickets()
BY_ID = {t.id: t for t in tickets}
mapper, reducer = ScriptedLLM(responder=map_stand_in), ScriptedLLM(responder=reduce_stand_in)


def usage_of(llm: ScriptedLLM) -> Usage:
    total = Usage()
    for call in llm.calls:
        total.input_tokens += count_messages(call["messages"])
    return total


# Map: one small call per ticket, validated.
notes = [TicketNote.model_validate_json(mapper([{"role": "user", "content": MAP.format(id=t.id, text=t.text)}]).text)
         for t in tickets]

# Reduce: code does the counting; the model only writes prose over the compact notes.
counts = {"category": Counter(n.category for n in notes), "priority": Counter(n.priority for n in notes),
          "language": Counter(n.language for n in notes)}
issues = "\n".join(f"- [{n.priority}] {n.category}: {n.issue}" for n in notes)
report = reducer([{"role": "user", "content": REDUCE.format(counts=json.dumps({k: dict(v) for k, v in counts.items()}),
                                                            issues=issues)}])
print("counts by category:", dict(counts["category"].most_common()))
print("urgent or high:", sum(counts["priority"][p] for p in ("urgent", "high")), "| report:", report.text)

# Hierarchical reduce: summarize 24 notes at a time, then combine the three partial summaries.
partial = ScriptedLLM(responder=reduce_stand_in)
for start in range(0, len(notes), 24):
    chunk = notes[start:start + 24]
    partial([{"role": "user", "content": REDUCE.format(counts="(see final pass)",
                                                       issues="\n".join(f"- {n.category}: {n.issue}" for n in chunk))}])

single_prompt = [{"role": "user", "content": SINGLE.format(all_tickets="\n\n".join(f"[{t.id}]\n{t.text}" for t in tickets))}]
single_in = count_messages(single_prompt)

OUT = {"map": 40, "reduce": 400, "single": 600}  # assumed output tokens per call; replace with measured usage
map_u, reduce_u = usage_of(mapper), usage_of(reducer)
map_u.output_tokens = OUT["map"] * len(tickets)
reduce_u.output_tokens = OUT["reduce"]
rows = [("single long call", 1, single_in, OUT["single"], single_in),
        ("map (72 calls)", len(mapper.calls), map_u.input_tokens, map_u.output_tokens,
         max(count_messages(c["messages"]) for c in mapper.calls)),
        ("reduce (1 call)", 1, reduce_u.input_tokens, reduce_u.output_tokens, reduce_u.input_tokens),
        ("reduce, 3 partial passes", len(partial.calls), usage_of(partial).input_tokens, 3 * OUT["reduce"],
         max(count_messages(c["messages"]) for c in partial.calls))]
print(f"\n{'step':<26} {'calls':>5} {'input tok':>9} {'output tok':>10} {'largest prompt':>14} "
      f"{'USD gpt-oss-120b':>16} {'USD gemini-3.5-flash':>20}")
for name, calls, tin, tout, largest in rows:
    u = Usage(input_tokens=tin, output_tokens=tout)
    print(f"{name:<26} {calls:>5} {tin:>9} {tout:>10} {largest:>14} "
          f"{cost_usd(u, 'openai/gpt-oss-120b'):>16.5f} {cost_usd(u, 'gemini-3.5-flash'):>20.5f}")

per_ticket = (single_in - count_messages([{"role": "user", "content": SINGLE.format(all_tickets="")}])) / len(tickets)
print(f"\nat {per_ticket:.1f} tokens per ticket, a 5,000-ticket week is about {per_ticket * 5000:,.0f} input tokens "
      f"in one call; each map call stays near {map_u.input_tokens / len(tickets):.0f}")

Code explained

  • In simple words: summarize each ticket in a small call, count with code, and write the report from the summaries; then compare with one giant call.
  • What happens: TicketNote validates each map result. The map stand-in copies gold labels (so the counts are right by construction, plumbing only) and the reduce stand-in returns a fixed sentence. What is real: the number of calls, every prompt's token count, and the cost math. The script also runs a hierarchical reduce (three partial summaries of 24 notes each) and builds the single-call alternative that pastes all 72 tickets into one prompt. Output lengths are assumptions (40 tokens per map, 400 per reduce, 600 for the single call).
  • Comes out: 72 map calls use 5,109 input tokens and no single prompt exceeds 88 tokens; the reduce prompt is 956 tokens. The single long call is 2,126 input tokens and, at 72 tickets, is actually cheaper than map-reduce (0.00068 vs about 0.0029 USD on gpt-oss-120b), because map-reduce repeats instructions 72 times and produces more output. The last line shows why map-reduce still wins at scale: at 29 tokens per ticket, a 5,000-ticket week is about 145,000 input tokens in one prompt, where long-context accuracy drops (Module 2's position effects) and the model has to count 5,000 items, which models do badly; each map call stays near 71 tokens and code does the counting exactly.

SituationUse thisWhy
The whole input fits comfortably and you need prose onlyOne callCheapest, simplest, as measured at 72 tickets
You need exact counts or per-item fieldsMap to typed notes, count in code, reduce for proseModels miscount; code does not
The input exceeds the context window or degrades in the middleMap-reduce, hierarchical reduce if the notes are longEvery call stays short; failures are per item and retryable
Items depend on each other (a thread of replies)Map over whole threads, not single messagesMap calls cannot see other items