Part A: Eliciting reasoning
Module 5: Reasoning and Advanced Prompting
By the end of this module, you'll have:
- A refund-eligibility task with 20 labelled scenarios whose gold answers come from a rule function, a harness that scores four ways of eliciting reasoning (direct, zero-shot chain-of-thought, demonstrated, structured scaffold), and a measured token bill for each.
- Real math and simulation for self-consistency: the binomial formula for majority voting, what correlated errors do to it, when voting hurts, a real majority vote over TinyLM samples, and a comparison with best-of-n under a verifier.
- A cost curve for test-time compute built with
pricing.py: samples, early stopping, prompt caching, and reasoning effort as settings on one dial. - Decomposition pipelines with typed intermediate results: least-to-most questions, an extract, decide, draft chain, a router with conditional paths, and a map-reduce weekly report over all 72 tickets with every call and token counted.
- A critique-and-revise loop with an iteration ceiling that shows, with
ScriptedLLM, how a loop without an external signal turns right answers wrong, and how a verifier makes only checked fixes stick. - Prompts as versioned code, an automatic prompt search against the dev set that overfits in a way you can see on the test set, holdout early stopping that catches it, and a meta-prompt that proposes new versions behind a validator.
Prerequisites: Module 1 (llm.chat, ScriptedLLM, TinyLM), Module 2 (token counting and pricing.py), Module 3 (temperature, sampling, reasoning effort and thinking tokens), and Module 4 (prompt anatomy, versioned prompts, frozen dev and test sets, confidence intervals). Python: dataclasses, re, and a little numpy.
Where we are: Module 4 made one prompt as good as careful editing could make it and proved changes with a frozen test set. This module goes past the single call: it asks the model to reason, samples it several times, splits work across calls, checks answers against something outside the model, and finally lets code search for the prompt. Every one of these buys accuracy with calls and tokens, so every section measures the bill.
How this module is organized
| Part | What it covers |
|---|---|
| Part A: Eliciting reasoning | The refund task and its rule verifier; chain-of-thought and where it helps; zero-shot vs demonstrated reasoning; structured scaffolds; reasoning models vs prompted reasoning |
| Part B: Sampling-based improvement | Majority voting math with independent and correlated errors; self-consistency on TinyLM; best-of-n with a verifier or scorer; ensembling; test-time compute as a dial with a cost curve |
| Part C: Decomposition | Least-to-most; prompt chains with typed intermediate results; a router with conditional paths; map-reduce over 72 tickets |
| Part D: Self-correction | Critique-and-revise loops; verifiers and external checks; why self-correction without a signal fails; iteration ceilings and diminishing returns |
| Part E: Programmatic prompting | Prompts as versioned code; automatic prompt search against an eval set; overfitting and holdout early stopping; meta-prompting |
Examples run in order in one repository. Each script is a complete file you run from the repository root with PYTHONPATH=.. Later scripts import examples/m05_helpers.py (shown in full in Part A) and, in a few places, an earlier Module 5 script. Everything runs offline; scripts that can call a real model say so and take a --live flag.
A word on honesty before we start, because this module is about techniques whose value depends on a real model. No API key was available when this module was built. So the numbers you will see come from four sources, and every result says which one it is:
| Source | What it can measure honestly | Example in this module |
|---|---|---|
| Deterministic code (rules, regex, scikit-learn, BM25) | Its own accuracy on our labelled data | The one-step shortcut vs extract-then-rule baselines |
| TinyLM (a real 1.07M-parameter model) | Real sampling behavior of a tiny model | Majority vote over TinyLM samples |
| Simulation with stated assumptions | What a technique does to a model with a given error pattern | Voting with correlated errors; best-of-n |
ScriptedLLM | Plumbing: calls, tokens, parsing, loop control, routing | Chains, routers, critique loops |
ScriptedLLM output is never model output. When a section needs a real model to answer its question, it gives you the harness and the command.
Part A: Eliciting reasoning
Setting up the working copy
source .venv/bin/activate # the environment from Module 1
cd supportdesk
pip install scikit-learn==1.9.1 # used by the ensembling example in Part B
mkdir -p runs
PYTHONPATH=. python examples/m05_baselines.py
Code explained
- In simple words: activate the course environment, add one library, make a folder for run logs, and run the first script.
- What happens: every script imports
supportdesk.*andexamples.m05_helpers, soPYTHONPATH=.must point at the repository root.scikit-learn1.9.1 is the version used to produce the outputs below.runs/receives a JSONL log from the prompt search in Part E. On a shared or small machine, the TinyLM script in Part B callstorch.set_num_threads(1), which keeps timings steadier. - Comes out: the table shown in the next section.
The task: a refund decision that needs several steps
Triage (Module 4) asks for one label that a keyword often gives away. To see reasoning techniques earn their keep, we need a task where the answer depends on several facts combined in a fixed order. Brightlane has one: refund eligibility. The billing-refunds help-center article says:
- Monthly plans are not refunded pro rata.
- Annual plans cancelled within 14 days of purchase or renewal get a full refund; after 14 days, no refund.
- Duplicate charges are always refunded in full.
- Only workspace owners and billing admins can cancel or request refunds.
A correct decision needs four facts (who is asking, whether the charge is a duplicate, monthly or annual, how many days ago) and the rules applied in order: role first, then duplicate, then the 14-day window. A customer writing "we were charged twice" who is a regular member does not get a refund from support; the owner must ask. That ordering is exactly what a one-glance answer gets wrong.
We encode the policy as a rule function, decide(facts). It gives us two things at once. First, gold labels: each of the 20 scenarios carries the facts a billing system would hold about that customer, and its gold answer is decide(facts), not a label someone typed. Second, a verifier: an external check we will use in Parts B and D to accept or reject a model's answer. Writing the rule also forced one decision the article leaves open: is day 14 "within 14 days"? decide says yes (<= 14); scenario R06 sits exactly on the boundary and R07 one day past it. When your rule function has to pick, ask the policy owner (here Maya, the support lead) and write the answer down.
The scenarios are in several languages, like real tickets: R12 is Spanish, R13 German, R19 Japanese. Today is fixed at Monday 21 September 2026 so that "10 days ago" and "last Monday" have one right reading.
examples/m05_helpers.py
"""Shared helpers for Module 5: the refund-decision task, its rule verifier, prompts, and scoring.
The refund policy comes from the billing-refunds help-center article. Each
scenario is a customer message plus the facts a billing system would hold
about that customer. The gold decision is computed from those facts by
`decide`, which is also the external verifier used later in the module.
"""
from __future__ import annotations
import math
import re
from collections.abc import Callable
from dataclasses import dataclass
from datetime import date
from supportdesk.data import CATEGORIES, get_article
from supportdesk.llm import Usage
TODAY = date(2026, 9, 21) # a Monday; every scenario is judged as of this day
LABELS = ("refund_duplicate", "refund_full", "no_refund", "needs_owner")
REFUND_ROLES = ("owner", "billing_admin")
@dataclass(frozen=True)
class Facts:
"""What the billing system knows about the charge in question."""
role: str # owner, billing_admin, admin, member
billing: str # monthly or annual
charged_on: date # date of the purchase or renewal being disputed
duplicate: bool # True if the same invoice was charged twice
@dataclass(frozen=True)
class RefundCase:
id: str
text: str
facts: Facts
@property
def gold(self) -> str:
return decide(self.facts)
def decide(f: Facts, today: date = TODAY) -> str:
"""The refund policy as code. Rules apply in order; the first match wins."""
if f.role not in REFUND_ROLES:
return "needs_owner" # only owners and billing admins may request refunds
if f.duplicate:
return "refund_duplicate" # duplicate charges are always refunded
if f.billing == "annual" and (today - f.charged_on).days <= 14:
return "refund_full" # annual plans within 14 days of purchase or renewal
return "no_refund" # monthly plans, or annual after 14 days
def _c(i: int, text: str, role: str, billing: str, y: int, m: int, d: int, dup: bool = False) -> RefundCase:
return RefundCase(f"R{i:02d}", text, Facts(role, billing, date(y, m, d), dup))
CASES: list[RefundCase] = [
_c(1, "I'm the workspace owner. Our card was charged twice for the Team plan on 3 September, same invoice number. Please refund one.", "owner", "monthly", 2026, 9, 3, True),
_c(2, "Owner here. We renewed our annual Business plan on 16 September by mistake. Can we get our money back?", "owner", "annual", 2026, 9, 16),
_c(3, "I own the workspace. We bought the annual Team plan on 10 July and want to stop now. Refund for the rest of the year?", "owner", "annual", 2026, 7, 10),
_c(4, "I'm a regular member of our workspace and I noticed we were charged twice this month. Can you refund the duplicate?", "member", "monthly", 2026, 9, 3, True),
_c(5, "Billing admin here. Our monthly Team plan charged us on 18 September but we barely use it. Refund please, it was only 3 days ago.", "billing_admin", "monthly", 2026, 9, 18),
_c(6, "As the workspace owner: our annual Business plan renewed on 7 September. We'd like to cancel and get a refund.", "owner", "annual", 2026, 9, 7),
_c(7, "Workspace owner. Our annual plan renewed on 6 September and we want a refund.", "owner", "annual", 2026, 9, 6),
_c(8, "I'm the billing admin. We started the annual Team plan on 30 August. It is not working for us, can we be refunded?", "billing_admin", "annual", 2026, 8, 30),
_c(9, "I'm the owner. We first subscribed in September 2025 and the annual plan auto-renewed on 12 September. Please refund the renewal.", "owner", "annual", 2026, 9, 12),
_c(10, "I'm a project manager in our workspace (not the owner). Our annual plan renewed 3 days ago and we want a refund.", "member", "annual", 2026, 9, 18),
_c(11, "Owner here. When our annual plan renewed on 12 August the card was charged twice for the same invoice. I know it is past 14 days, but can we get the duplicate back?", "owner", "annual", 2026, 8, 12, True),
_c(12, "Soy la propietaria del espacio de trabajo. Renovamos el plan anual hace 20 días. ¿Podemos cancelar y recibir el reembolso?", "owner", "annual", 2026, 9, 1),
_c(13, "Ich bin Billing-Admin. Unser Monatsplan wurde am 3. September doppelt abgebucht (gleiche Rechnung). Bitte erstatten Sie die Doppelbuchung.", "billing_admin", "monthly", 2026, 9, 3, True),
_c(14, "I'm the owner. We pay monthly for Business, charged on 1 September. We're cancelling today; please refund the unused part of the month.", "owner", "monthly", 2026, 9, 1),
_c(15, "Owner of the workspace. I bought the annual Team plan last Monday and my team hates it. Can I get a refund?", "owner", "annual", 2026, 9, 14),
_c(16, "I'm a workspace admin (I manage members and boards). Our annual plan renewed 5 days ago; please refund it.", "admin", "annual", 2026, 9, 16),
_c(17, "Billing admin. Our annual Team plan renewed on 8 September. We want to cancel and get our money back.", "billing_admin", "annual", 2026, 9, 8),
_c(18, "I'm the owner. I saw two charges of 144 USD on 3 September and thought it was a duplicate, but they are invoices for our two separate workspaces. Can I still get one refunded? We pay monthly.", "owner", "monthly", 2026, 9, 3),
_c(19, "ワークスペースのオーナーです。10日前に年間プランが更新されました。返金してもらえますか?", "owner", "annual", 2026, 9, 11),
_c(20, "Billing admin here. Our annual Team plan renewed on 1 September and nobody has logged in since. Refund?", "billing_admin", "annual", 2026, 9, 1),
]
# Prompts ---------------------------------------------------------------------
def policy_text() -> str:
return get_article("billing-refunds").body
INSTRUCTION = (
"You decide refund eligibility for Brightlane support, strictly by the policy below.\n"
f"Today is {TODAY:%A %d %B %Y}.\n"
"Possible decisions: refund_duplicate, refund_full, no_refund, needs_owner "
"(needs_owner means the requester is not a workspace owner or billing admin).\n\n"
"<policy>\n{policy}\n</policy>\n\n<ticket>\n{ticket}\n</ticket>\n"
)
DECISION_LINE = "End with one line exactly like: DECISION: <decision>"
DEMOS = (
"Example ticket: I'm the billing admin. Our annual Business plan renewed on 10 September. Refund please.\n"
"Reasoning: The requester is a billing admin, so they may request refunds. No duplicate charge is mentioned. "
"The plan is annual and renewed 11 days before today, which is within 14 days, so a full refund applies.\n"
"DECISION: refund_full\n\n"
"Example ticket: I'm a member of the workspace. We were charged twice on 2 September.\n"
"Reasoning: The requester is a regular member, not an owner or billing admin, so they cannot request a refund, "
"even for a duplicate charge. The owner must ask.\n"
"DECISION: needs_owner\n"
)
SCAFFOLD = (
"Work through these steps in order and fill in every line:\n"
"ROLE: <owner | billing_admin | admin | member>\n"
"DUPLICATE: <yes | no>\n"
"BILLING: <monthly | annual>\n"
"CHARGED_ON: <YYYY-MM-DD>\n"
"DAYS_SINCE: <whole days from CHARGED_ON to today>\n"
"DECISION: <decision> (apply the rules in this order: role, duplicate, annual within 14 days, otherwise no_refund)\n"
)
def _user(case: RefundCase, tail: str) -> list[dict]:
body = INSTRUCTION.format(policy=policy_text(), ticket=case.text)
return [{"role": "user", "content": body + "\n" + tail}]
def prompt_direct(case: RefundCase) -> list[dict]:
return _user(case, "Answer with the decision only. " + DECISION_LINE)
def prompt_zero_shot_cot(case: RefundCase) -> list[dict]:
return _user(case, "Think step by step before you answer. " + DECISION_LINE)
def prompt_demonstrated(case: RefundCase) -> list[dict]:
return _user(case, DEMOS + "\nNow reason the same way about the ticket above. " + DECISION_LINE)
def prompt_scaffold(case: RefundCase) -> list[dict]:
return _user(case, SCAFFOLD)
PROMPTS: dict[str, Callable[[RefundCase], list[dict]]] = {
"direct": prompt_direct, "zero_shot_cot": prompt_zero_shot_cot,
"demonstrated": prompt_demonstrated, "scaffold": prompt_scaffold,
}
# Parsing and scoring ------------------------------------------------------------
DECISION_RE = re.compile(r"^\s*\**DECISION\**\s*:\s*\**\s*([a-z_]+)", re.IGNORECASE | re.MULTILINE)
def parse_decision(text: str) -> str | None:
"""The last 'DECISION: <label>' line, if it names a known label."""
found = DECISION_RE.findall(text or "")
label = found[-1].lower() if found else None
return label if label in LABELS else None
def wilson(k: int, n: int, z: float = 1.96) -> tuple[float, float]:
"""95 percent Wilson score interval for k successes out of n."""
if n == 0:
return (0.0, 1.0)
p = k / n
centre = (p + z * z / (2 * n)) / (1 + z * z / n)
half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
return (max(0.0, centre - half), min(1.0, centre + half))
@dataclass
class EvalResult:
correct: int
n: int
unparsed: int
usage: Usage
wrong: list[tuple[str, str | None, str]] # (case id, predicted, gold)
def line(self, name: str) -> str:
lo, hi = wilson(self.correct, self.n)
return (f"{name:<14} {self.correct:>2}/{self.n} 95% CI {lo:.2f}-{hi:.2f} unparsed {self.unparsed} "
f"tokens in/out {self.usage.input_tokens}/{self.usage.output_tokens}")
def evaluate(chat_fn: Callable, build: Callable[[RefundCase], list[dict]], cases: list[RefundCase] = CASES,
**chat_kwargs) -> EvalResult:
"""Run one prompt style over the cases and score the parsed decision against the rule."""
usage, correct, unparsed, wrong = Usage(), 0, 0, []
for case in cases:
result = chat_fn(build(case), **chat_kwargs)
usage.input_tokens += result.usage.input_tokens
usage.output_tokens += result.usage.output_tokens
usage.reasoning_tokens += result.usage.reasoning_tokens
label = parse_decision(result.text)
unparsed += label is None
if label == case.gold:
correct += 1
else:
wrong.append((case.id, label, case.gold))
return EvalResult(correct, len(cases), unparsed, usage, wrong)
def case_from_messages(messages: list[dict]) -> RefundCase:
"""Find which scenario a prompt is about (used by scripted stand-ins, never by real code)."""
content = messages[-1]["content"]
for case in CASES:
if case.text in content:
return case
raise KeyError("no scenario text found in prompt")
# Classical baselines --------------------------------------------------------------
def shortcut_baseline(text: str) -> str:
"""One step, no intermediate facts: react to the most salient keyword."""
t = text.lower()
if re.search(r"twice|duplicate|doppel|duplicado", t):
return "refund_duplicate"
if re.search(r"annual|anual|年間", t):
return "refund_full"
return "no_refund"
MONTHS = {m: i for i, m in enumerate(["january", "february", "march", "april", "may", "june", "july",
"august", "september", "october", "november", "december"], 1)}
def extract_facts_regex(text: str) -> Facts | None:
"""Pull the four facts out of English text with regular expressions; None if any is missing."""
t = text.lower()
if re.search(r"billing[- ]admin", t):
role = "billing_admin"
elif re.search(r"\bnot the owner\b|\bmember\b|project manager", t):
role = "member"
elif re.search(r"\bworkspace admin\b", t):
role = "admin"
elif re.search(r"\bowner\b|\bi own\b", t):
role = "owner"
else:
return None
billing = "annual" if "annual" in t else "monthly" if re.search(r"monthly|month", t) else None
duplicate = bool(re.search(r"twice|duplicate", t)) and "separate" not in t
charged_on = None
if m := re.search(r"(\d{1,2}) (" + "|".join(MONTHS) + r")", t):
charged_on = date(TODAY.year, MONTHS[m.group(2)], int(m.group(1)))
elif m := re.search(r"(" + "|".join(MONTHS) + r") (\d{1,2})\b", t):
charged_on = date(TODAY.year, MONTHS[m.group(1)], int(m.group(2)))
elif m := re.search(r"(\d+) days ago", t):
charged_on = date.fromordinal(TODAY.toordinal() - int(m.group(1)))
if billing is None or charged_on is None:
return None
return Facts(role, billing, charged_on, duplicate)
# Ticket triage by keyword definitions read from a prompt -----------------------------
DEF_LINE = re.compile(r"^- ([a-z_]+): (.+)$", re.MULTILINE)
FALLBACK = re.compile(r"If no category fits, answer: ([a-z_]+)")
def keyword_reader(prompt: str, ticket_text: str) -> str:
"""A deterministic stand-in 'model' that follows a triage prompt's category definitions.
It counts how many of each category's listed keywords appear in the ticket,
picks the highest count (ties go to the category listed first), and uses the
prompt's fallback when nothing matches. It is not a language model; it lets
prompt search run for real without an API key.
"""
text = ticket_text.lower()
best, best_hits = None, 0
for name, words in DEF_LINE.findall(prompt):
hits = sum(1 for w in words.split(", ") if w.strip() and w.strip() in text)
if hits > best_hits:
best, best_hits = name, hits
if best is None:
m = FALLBACK.search(prompt)
best = m.group(1) if m else "how_to"
return best
TRIAGE_KEYWORDS: dict[str, list[str]] = {
"billing": ["charge", "invoice", "price", "cost", "vat", "pay", "tax", "discount"],
"cancellation": ["cancel", "refund", "money back"],
"account_access": ["password", "locked", "2fa", "sign in", "saml", "gdpr", "region"],
"bug": ["not loading", "stopped", "fails", "error", "bug", "down", "lost my changes"],
"how_to": ["how do", "how many", "where", "export", "can i", "will the"],
"feature_request": ["please add", "integration", "roadmap", "when will", "support"],
}
def triage_prompt(keywords: dict[str, list[str]], fallback: str = "how_to") -> str:
"""A triage prompt whose category definitions are keyword lists (order matters for ties)."""
lines = "\n".join(f"- {name}: {', '.join(words)}" for name, words in keywords.items())
return ("Classify the Brightlane support ticket into exactly one category.\n"
f"Categories and the words that signal them:\n{lines}\n"
f"If no category fits, answer: {fallback}\n"
"Reply with the category name only.")
def triage_messages(prompt: str, ticket_text: str) -> list[dict]:
return [{"role": "system", "content": prompt}, {"role": "user", "content": ticket_text}]
def keyword_responder(messages: list[dict], kwargs: dict) -> str:
"""ScriptedLLM responder: apply keyword_reader to the system prompt and the ticket."""
return keyword_reader(messages[0]["content"], messages[-1]["content"])
def parse_category(text: str) -> str | None:
word = (text or "").strip().strip(".`'\"").lower()
return word if word in CATEGORIES else None
Code explained
- In simple words: one file holds the policy as code, the 20 scenarios, four prompt styles, a parser, a scorer, and two small stand-ins used later.
- What happens:
FactsandRefundCase: the billing record for one scenario and the customer's message.RefundCase.goldis computed, so the labels can never drift from the rule.decide: the policy as ordered rules; the first match wins. This is both the answer key and the verifier.CASES: 20 scenarios, built with the small_cconstructor. Several are traps on purpose: a member reporting a duplicate (R04), a monthly plan charged "only 3 days ago" (R05, 14-day rule is annual only), two charges that turn out to be separate workspaces (R18), a renewal 15 days ago (R07).policy_text: reads the real article body withget_article, so the prompt and the help center cannot disagree.INSTRUCTION,DEMOS,SCAFFOLDand the fourprompt_*builders: the same task asked four ways (Part A compares them). All four end with aDECISION: <label>line so one parser serves them all.parse_decision: takes the lastDECISION:line (tolerating Markdown bold) and returns it only if it is one of the four labels. Anything else is "unparsed", which is scored as wrong, never guessed.wilson: the 95 percent Wilson interval for k out of n, used on every accuracy in this module because n is small.EvalResultandevaluate: run one prompt style over the cases with any function shaped likellm.chat, sum usage, and record every miss as (case, predicted, gold).case_from_messages: lets a scripted stand-in find which scenario a prompt is about. Real code never needs it.shortcut_baselineandextract_facts_regex: two classical baselines, explained in the next section.keyword_reader,TRIAGE_KEYWORDS,triage_prompt,triage_messages,keyword_responder,parse_category: a deterministic stand-in that follows a triage prompt's keyword definitions literally. It is not a language model; it lets the ensembling and prompt-search examples run for real without a key.
- Comes out: nothing by itself; the other scripts import it.
Chain-of-thought, and where it genuinely helps
Chain-of-thought (CoT) means getting the model to write intermediate steps before the final answer. Two facts explain why it can help. A model produces each token with a fixed amount of computation, so a problem that needs several dependent steps cannot always be squeezed into the single token that states the answer. And each written step becomes context for the next one: once the model has written "the requester is a member", the rest of the answer is conditioned on that fact.
The research record is specific about where this pays:
- Wei et al., "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (NeurIPS 2022), showed that eight worked examples with reasoning let a 540B-parameter model reach state-of-the-art accuracy on GSM8K math word problems at the time, and that the benefit appeared only in large models.
- Sprague et al., "To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning" (ICLR 2025), looked at over 100 papers and ran 20 datasets on 14 models. CoT gave strong gains mainly on math and logic and much smaller gains elsewhere; on MMLU, answers with and without CoT were nearly identical except for questions involving symbolic operations.
So the rule of thumb: CoT helps when the answer depends on several facts combined by rules, dates, or arithmetic (our refund task), and mostly adds cost when the answer is recognition (most triage labels). You can see the same shape without any language model. Below, two classical solvers attack the refund task. The shortcut answers in one step from the most salient keyword ("twice" means refund the duplicate, "annual" means full refund). The extract-then-rule solver first writes down the four facts with regular expressions, then applies decide, and abstains when it cannot find a fact.
examples/m05_baselines.py:
"""Part A: the refund task, its gold labels, and two classical baselines (one step vs several steps)."""
from collections import Counter
from examples.m05_helpers import CASES, TODAY, decide, extract_facts_regex, shortcut_baseline, wilson
print(f"{len(CASES)} scenarios judged as of {TODAY}; gold labels: {dict(Counter(c.gold for c in CASES))}\n")
print(f"{'id':<4} {'gold':<17} {'shortcut':<17} {'extract+rule':<17} days")
scores = {"shortcut": 0, "extract+rule": 0}
abstained = 0
for case in CASES:
one_step = shortcut_baseline(case.text)
facts = extract_facts_regex(case.text)
stepwise = decide(facts) if facts else "(no facts)"
abstained += facts is None
scores["shortcut"] += one_step == case.gold
scores["extract+rule"] += stepwise == case.gold
days = (TODAY - case.facts.charged_on).days
print(f"{case.id:<4} {case.gold:<17} {one_step:<17} {stepwise:<17} {days}")
print()
for name, k in scores.items():
lo, hi = wilson(k, len(CASES))
print(f"{name:<13} {k:>2}/{len(CASES)} correct 95% CI {lo:.2f}-{hi:.2f}")
print(f"extract+rule abstained on {abstained} and was wrong on "
f"{len(CASES) - scores['extract+rule'] - abstained} of the cases it answered")
Code explained
- In simple words: compare answering at a glance with writing the facts down first and applying the rules.
- What happens: for each scenario the script prints the gold label, the shortcut's answer, the extract-then-rule answer (or
(no facts)when a regex could not find all four facts), and how many days ago the charge was. Then it scores both with Wilson intervals. - Comes out: the shortcut gets 11/20. Its misses are exactly the multi-step traps: it refunds R04's duplicate for a member, refunds annual plans past 14 days (R03, R07, R08, R12, R20), treats R18's two separate workspace charges as a duplicate, and refunds annual plans for a member and a workspace admin (R10, R16) who may not ask. Extract-then-rule answers 14 and gets all 14 right; it abstains on 6: three are in Spanish, German, or Japanese (R12, R13, R19), two give relative dates ("this month", "last Monday" in R04 and R15), and R01 never says whether the plan is monthly or annual. The intervals overlap (0.34 to 0.74 vs 0.48 to 0.85), so with n=20 this is suggestive, not proof. The lesson carries over to models: writing the intermediate facts turns silent errors into visible gaps you can route to a person.
Zero-shot vs demonstrated reasoning, and structured scaffolds
There are three common ways to ask for reasoning, and they differ in what they cost and how much they constrain the model:
- Zero-shot CoT: add an instruction such as "Think step by step." Kojima et al., "Large Language Models are Zero-Shot Reasoners" (NeurIPS 2022), found that adding "Let's think step by step" took text-davinci-002 from 17.7 to 78.7 percent on MultiArith and from 10.4 to 40.7 percent on GSM8K.
- Demonstrated (few-shot) CoT: show one or more worked examples whose reasoning follows the path you want. This is Wei et al.'s original method. It costs the example tokens on every call and teaches the order of checks by example.
- Structured scaffold: give the model a form to fill in, one line per intermediate fact, then the decision. It is the cheapest way to get reasoning you can check line by line, because each field can be compared against the billing record.
All four prompt builders in the helpers end with the same DECISION: line. The script below measures their input size with the real o200k tokenizer, prices them with pricing.py, and then checks the scoring harness with two scripted stand-ins: one that fills the scaffold perfectly and one whose output format drifts.
examples/m05_cot_prompts.py:
"""Part A: four ways to ask for the refund decision, their token cost, and a harness to score them.
Run without arguments to check the harness with scripted stand-ins (plumbing
only, no model quality). Run with --live to score a real model:
LLM_PROVIDER=groq GROQ_API_KEY=... PYTHONPATH=. python examples/m05_cot_prompts.py --live
"""
import sys
from examples.m05_helpers import CASES, PROMPTS, TODAY, case_from_messages, decide, evaluate
from supportdesk.llm import Usage
from supportdesk.pricing import cost_usd
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages
# 1. Prompt size per style (input tokens are measured; output lengths are assumptions to replace).
ASSUMED_OUTPUT = {"direct": 8, "zero_shot_cot": 180, "demonstrated": 120, "scaffold": 60}
print(f"{'style':<14} {'input tok':>9} {'assumed out':>11} {'USD/1k gpt-oss-120b':>20} {'USD/1k gemini-3.5-flash':>24}")
for name, build in PROMPTS.items():
tokens_in = count_messages(build(CASES[0]))
usage = Usage(input_tokens=tokens_in, output_tokens=ASSUMED_OUTPUT[name])
print(f"{name:<14} {tokens_in:>9} {ASSUMED_OUTPUT[name]:>11} "
f"{1000 * cost_usd(usage, 'openai/gpt-oss-120b'):>20.3f} {1000 * cost_usd(usage, 'gemini-3.5-flash'):>24.3f}")
# 2. Plumbing check: a scripted responder that fills the scaffold from the true facts.
def oracle(messages, kwargs):
f = case_from_messages(messages).facts
return (f"ROLE: {f.role}\nDUPLICATE: {'yes' if f.duplicate else 'no'}\nBILLING: {f.billing}\n"
f"CHARGED_ON: {f.charged_on}\nDAYS_SINCE: {(TODAY - f.charged_on).days}\nDECISION: {decide(f)}")
# 3. Format drift: the same answers, but a third of them written in a way the parser must reject.
def drifting(messages, kwargs):
text = oracle(messages, kwargs)
case = case_from_messages(messages)
if int(case.id[1:]) % 3 == 0:
label = text.rsplit("DECISION: ", 1)[1]
return text.rsplit("\n", 1)[0] + f"\nDecision - {label.replace('_', ' ').title()}"
if int(case.id[1:]) % 3 == 1:
return text.replace("DECISION:", "**DECISION:**")
return text
print()
for name, responder in (("oracle", oracle), ("drifting", drifting)):
print(evaluate(ScriptedLLM(responder=responder), PROMPTS["scaffold"]).line(name))
# 4. The real measurement, when a key is set.
if "--live" in sys.argv:
from supportdesk.llm import chat
for name, build in PROMPTS.items():
result = evaluate(chat, build, temperature=0.0, max_tokens=2000)
print(result.line(name), " wrong:", result.wrong)
Code explained
- In simple words: price four ways of asking the same question, then prove the scoring harness catches format drift before trusting it with a real model.
- What happens:
- Part 1 counts real input tokens for scenario R01 under each style. Output lengths are assumptions (8 tokens for a bare label, 180 for free-form reasoning, 120 for reasoning after examples, 60 for the six-line scaffold); replace them with your measured
usage.output_tokensonce you run--live. - Part 2's
oraclestand-in fills the scaffold from the true facts. It must score 20/20, which proves the parser, the scorer, andcase_from_messagesagree. - Part 3's
driftingstand-in gives the same right answers but writes a third of them asDecision - Refund Fulland another third in Markdown bold. The parser accepts bold and rejects the other form. - Part 4 runs each style against a real model when you pass
--live.
- Part 1 counts real input tokens for scenario R01 under each style. Output lengths are assumptions (8 tokens for a bare label, 180 for free-form reasoning, 120 for reasoning after examples, 60 for the six-line scaffold); replace them with your measured
- Comes out: zero-shot CoT adds only 2 input tokens to the direct prompt, but its output is where the money goes: at the assumed lengths it costs 3.6 times the direct prompt on gpt-oss-120b. The demonstrations add 138 input tokens to every call. The scaffold is the cheapest way to get reasoning. The drifting stand-in drops from 20/20 to 14/20 with 6 unparsed, all right answers: format drift looks exactly like a reasoning failure unless you count unparsed replies separately, which
EvalResultdoes.
To get your own accuracy numbers, run LLM_PROVIDER=groq GROQ_API_KEY=... PYTHONPATH=. python examples/m05_cot_prompts.py --live. It prints one line per style with the count correct, a 95 percent interval, unparsed replies, tokens, and the list of misses. With n=20, a difference of 2 or 3 cases is within noise; look at which cases moved (the traps R04, R05, R07, R18 are the informative ones) rather than at the totals.
Here is what one scaffold reply from a real model looks like. This is an illustrative sample (not captured in this build; produced for teaching). Your output will differ.
ROLE: member
DUPLICATE: yes
BILLING: monthly
CHARGED_ON: 2026-09-03
DAYS_SINCE: 18
DECISION: needs_owner
Code explained
- In simple words: the shape of a good scaffold reply for R04, a member reporting a duplicate charge.
- What happens: each line is a fact you can check against the billing record before trusting the decision: role, duplicate flag, billing period, date, and day count. The decision follows the rule order (role first), so the duplicate does not matter.
- Comes out:
parse_decisionreadsneeds_owner, which matches the gold label. If ROLE had saidowner, you would know the error was in reading the ticket, not in applying the policy.
Which style to reach for:
| Situation | Use this | Why |
|---|---|---|
| Recognition tasks (category, language, sentiment) | Direct answer, no CoT | Sprague et al. found little gain outside math and logic; CoT only adds output tokens |
| Multi-step rules, dates, arithmetic, on a non-reasoning model | Structured scaffold | Cheapest reasoning, and each line is a checkable fact |
| The model keeps applying rules in the wrong order | Demonstrated CoT with one or two examples that show the order | Examples teach the path; the cost is fixed input tokens per call |
| Exploring a new task before you know its steps | Zero-shot CoT, then read the traces | Two words of prompt; the traces show you what scaffold to write |
| A reasoning model (gpt-oss, Gemini thinking, Qwen3 thinking) | Plain direct prompt, no "think step by step" | It already reasons internally; see the next section |
Reasoning models vs prompted reasoning
A reasoning model is trained to produce a long hidden or visible chain of thought before answering (OpenAI's o1 in September 2024 and DeepSeek-R1 in January 2025 made this mainstream; Module 3 covered thinking tokens and effort settings). Our default Groq model, openai/gpt-oss-120b, is one; it takes reasoning_effort of low, medium, or high.
Prompted reasoning and trained reasoning overlap, so do not stack them blindly. OpenAI's reasoning best-practices guide says that "prompting them to 'think step by step' or 'explain your reasoning' is unnecessary", that reasoning models "often don't need few-shot examples", and to "try to write prompts without examples first". The guide still recommends delimiters (Markdown, XML tags, section titles) to separate parts of the input, which our <policy> and <ticket> tags do.
For the refund task, that means: on gpt-oss-120b, start with prompt_direct and turn the effort dial; on a non-reasoning model, use the scaffold. What you compare is not "which is smarter" but accuracy per dollar, and a reasoning model's hidden tokens bill as output tokens. examples/m05_effort_dial.py (Part B) has a --live mode that measures exactly this: the direct prompt at low, medium, and high effort over the 20 scenarios, with reasoning tokens and dollars per run.
| Situation | Use this | Why |
|---|---|---|
| Hard multi-step decisions, low volume, latency tolerant | Reasoning model at medium or high effort, direct prompt | Best accuracy per engineering hour; you pay in output tokens and seconds |
| High volume, the steps are known and fixed | Non-reasoning model with a structured scaffold, or code for the rule | Short, checkable output; much cheaper per call |
| You must show the steps to a reviewer | Scaffold output (your own fields) | Hidden reasoning is not returned by every provider and is not a reliable explanation (Module 3) |
| The rule is fully known and the facts are in a database | Code (decide), with a model only to read the ticket | Part C: a model should not do what a function can do exactly |