Part B: Sampling-based improvement
Self-consistency and majority voting: the math
Self-consistency (Wang et al., "Self-Consistency Improves Chain of Thought Reasoning in Language Models", ICLR 2023) samples several reasoning paths at a temperature above zero and returns the most common final answer. The paper reports absolute gains over greedy chain-of-thought of 17.9 points on GSM8K, 11.0 on SVAMP, 12.2 on AQuA, 6.4 on StrategyQA, and 3.9 on ARC-challenge.
Why would voting help? Start with the simplest model. Each sample is right with probability p, samples are independent, and there are two outcomes, right or wrong. With n samples, the vote is right when more than half are right:
P(vote right) = sum over k > n/2 of C(n, k) p^k (1 - p)^(n - k)
(for even n, a tie counts as half right). That formula has one consequence that surprises people: voting pushes accuracy away from 0.5 in whichever direction p already lies. If p is above 0.5, more samples help. If p is below 0.5, more samples make things worse. And real samples from one model are not independent: some questions are hard for every sample. The script computes the formula, checks it by simulation, and then adds correlation.
examples/m05_voting_math.py:
"""Part B: exactly how majority voting changes accuracy, with the binomial formula and a simulation."""
from math import comb
import numpy as np
rng = np.random.default_rng(5)
def majority_exact(p: float, n: int) -> float:
"""P(more than half of n independent samples are correct); ties split 50/50."""
win = sum(comb(n, k) * p**k * (1 - p) ** (n - k) for k in range(n // 2 + 1, n + 1))
if n % 2 == 0:
win += 0.5 * comb(n, n // 2) * (p * (1 - p)) ** (n // 2)
return win
def majority_simulated(p: float, n: int, questions: int = 200_000) -> float:
correct = rng.random((questions, n)) < p
k = correct.sum(axis=1)
return float(np.mean((k > n / 2) + 0.5 * (k == n / 2)))
# 1. Independent errors: the formula and the simulation agree.
ns = (1, 3, 5, 9, 15, 31)
print("Independent samples, two outcomes (right or wrong). Rows: per-sample accuracy p.")
print("p " + "".join(f"n={n:<6}" for n in ns))
for p in (0.3, 0.5, 0.6, 0.7, 0.8, 0.9):
print(f"{p:<5} " + "".join(f"{majority_exact(p, n):<8.3f}" for n in ns))
print("simulated p=0.7, n=5:", round(majority_simulated(0.7, 5), 3), " exact:", round(majority_exact(0.7, 5), 3))
# 2. Correlated errors: each question has its own accuracy p_i drawn from a Beta with mean p.
# rho is the correlation between two samples' correctness on the same question.
def majority_correlated(p: float, n: int, rho: float, questions: int = 200_000) -> float:
if rho == 0:
return majority_exact(p, n)
a, b = p * (1 / rho - 1), (1 - p) * (1 / rho - 1)
p_i = rng.beta(a, b, size=questions)
k = rng.binomial(n, p_i)
return float(np.mean((k > n / 2) + 0.5 * (k == n / 2)))
print("\nCorrelated samples, p=0.7. Rows: correlation rho between samples on the same question.")
print("rho " + "".join(f"n={n:<6}" for n in ns))
for rho in (0.0, 0.1, 0.3, 0.6, 0.9):
print(f"{rho:<5} " + "".join(f"{majority_correlated(0.7, n, rho):<8.3f}" for n in ns))
for rho in (0.1, 0.3, 0.6):
a, b = 0.7 * (1 / rho - 1), 0.3 * (1 / rho - 1)
ceiling = float(np.mean(rng.beta(a, b, size=200_000) > 0.5))
print(f"ceiling as n grows, rho={rho}: {ceiling:.3f} (share of questions whose own p_i > 0.5)")
# 3. More than two answers: plurality voting when wrong answers spread out or pile up.
def plurality(probs: list[float], n: int, questions: int = 100_000) -> float:
"""Answer 0 is correct. Ties are broken uniformly at random among the tied answers."""
counts = rng.multinomial(n, probs, size=questions)
top = counts.max(axis=1, keepdims=True)
tied = counts == top
return float(np.mean(tied[:, 0] / tied.sum(axis=1)))
print("\nPlurality vote with four possible answers (correct answer has probability 0.40):")
for name, probs in (("wrong answers spread 0.20/0.20/0.20", [0.4, 0.2, 0.2, 0.2]),
("one tempting wrong answer 0.45/0.10/0.05", [0.4, 0.45, 0.1, 0.05])):
print(f" {name:<42}" + " ".join(f"n={n}: {plurality(probs, n):.3f}" for n in (1, 5, 15, 31)))
Code explained
- In simple words: how much a vote of n samples helps, first when samples are independent, then when they share mistakes, then when there are several wrong answers to vote for.
- What happens:
majority_exactis the binomial formula above;majority_simulatedchecks it by drawing 200,000 simulated questions.majority_correlatedgives every question its own accuracy p_i, drawn from a Beta distribution with mean p. The Beta's parameters are chosen so that the correlation between two samples' correctness on the same question is exactly rho (for a Beta-binomial, that correlation is 1 / (a + b + 1)). rho = 0 is the independent case; rho near 1 means each question is either always right or always wrong.- The "ceiling" lines compute the share of questions whose own p_i is above 0.5. With infinitely many samples, the vote gets exactly those questions right and no others.
pluralityhandles four possible answers (our refund labels), with ties broken at random.
- Comes out:
- Independent samples: at p = 0.7, five votes give 0.837 and 31 give 0.990. At p = 0.5 nothing changes. At p = 0.3, 31 votes drive accuracy down to 0.010. The simulation matches the formula (0.837).
- Correlated samples at p = 0.7: with rho = 0.3, 31 votes reach only 0.768, close to the ceiling of 0.774; with rho = 0.6, voting adds about one point. Most of the gain from voting comes in the first 5 to 9 samples, and correlation decides where it stops.
- Four answers: if the wrong answers are spread out, a correct answer with only 0.40 probability still wins the plurality more often as n grows (0.843 at n = 31). If one tempting wrong answer has 0.45, voting converges on the wrong answer (0.382 at n = 31). That is exactly the refund shortcut: "charged twice, so refund" is one tempting wrong answer, not noise.
Suggested file: m05-voting-correlation.png
When voting helps and when it hurts:
| Situation | Use this | Why |
|---|---|---|
| Per-sample accuracy clearly above 0.5, answers vary between samples | Majority vote over 5 to 9 samples | Most of the gain arrives by n = 9; beyond that, correlation caps it |
| Samples agree almost always (low temperature, easy task) | One sample | rho near 1: voting costs n times as much and changes nothing |
| One tempting wrong answer, or accuracy below 0.5 on a class of inputs | Fix the prompt or add a verifier, do not vote | Voting amplifies a systematic error |
| Free-form answers (replies, summaries) | Do not vote on raw text; extract a label or use a scorer | Votes need answers that can be compared for equality |
Self-consistency for real on TinyLM
TinyLM (1.07M parameters, trained on this project's support corpus) cannot follow a refund policy, but it can answer short fact prompts whose answer is a number: "Reset links expire after" should continue with 30. That is enough to run self-consistency for real: sample nine continuations at a temperature above zero, pull the first number out of each, and vote.
examples/m05_tinylm_vote.py:
"""Part B: self-consistency for real on TinyLM. Sample several completions, vote on the number they state."""
import re
import time
from collections import Counter
import torch
from supportdesk.tinylm import SamplingParams, generate, load
torch.set_num_threads(1) # this machine is shared; 1 thread keeps timings steadier
model, tokenizer = load()
NUMBER = re.compile(r"\d+(?:\.\d+)?")
# Prompt, and the number a correct continuation states (from the help-center facts).
FACT_PROMPTS = [
("Reset links expire after", "30"),
("After 5 failed attempts an account is locked for", "15"),
("Annual plans cancelled within", "14"),
("Business costs 24 USD per user per month, or", "20"),
("Free supports up to", "3"),
("Customer (Ana): Hi, how long until my locked account opens?\nAgent (Ben):", "15"),
("Customer (Ana): Hi, what does Business cost per user if we pay annually?\nAgent (Ben):", "20"),
("Customer (Ana): Hi, how many users can the Free plan have?\nAgent (Ben):", "3"),
]
N_SAMPLES = 9
def first_number(text: str) -> str | None:
match = NUMBER.search(text)
return match.group(0) if match else None
started = time.perf_counter()
for temperature in (1.0, 1.5):
greedy_ok = single_ok = vote_ok = 0
print(f"temperature {temperature}, {N_SAMPLES} samples per prompt")
for prompt, answer in FACT_PROMPTS:
greedy = first_number(generate(model, tokenizer, prompt, SamplingParams(max_new_tokens=16, temperature=0)).text)
samples = [first_number(generate(model, tokenizer, prompt,
SamplingParams(max_new_tokens=16, temperature=temperature, seed=s)).text)
for s in range(N_SAMPLES)]
votes = Counter(x for x in samples if x is not None)
top = votes.most_common(2)
tie = len(top) == 2 and top[0][1] == top[1][1]
winner = None if not top or tie else top[0][0] # a tie is not a majority: count it as no answer
p_hat = sum(x == answer for x in samples) / N_SAMPLES
greedy_ok += greedy == answer
single_ok += p_hat
vote_ok += winner == answer
label = prompt.splitlines()[0].replace("Customer (Ana): Hi, ", "Q: ")
print(f" want {answer:>3} greedy {str(greedy):>4} per-sample {p_hat:.2f} vote {'tie' if tie else str(winner):>4} "
f"{dict(votes.most_common(3))} | {label[:48]}")
k = len(FACT_PROMPTS)
print(f" accuracy over {k} prompts: greedy {greedy_ok}/{k}, one sample {single_ok:.2f}/{k} (average), "
f"majority of {N_SAMPLES} {vote_ok}/{k}\n")
print(f"{2 * len(FACT_PROMPTS) * (N_SAMPLES + 1)} generations in {time.perf_counter() - started:.0f} s")
Code explained
- In simple words: ask TinyLM eight fact questions, once greedily and nine times with sampling, and see whether the vote beats a single sample.
- What happens: the first five prompts are sentence starts copied from the help center; the last three are phrased as a customer question inside a chat transcript, which the corpus contains in different wording.
generateruns withmax_new_tokens=16(TinyLM's context is 128 tokens, andgeneratefails whenmax_new_tokensreaches 128). Each sample has its own seed, so the run is reproducible. A tie between the top two answers counts as no answer, because a tie is not a majority.torch.set_num_threads(1)keeps timing steady on a shared machine. - Comes out: at temperature 1.0 the vote, the greedy answer, and the average sample all score 4/8. The model knows the four help-center sentences and is wrong on the four others in a consistent way: for the Business annual price it says 12 in 8 of 9 samples (a memorized neighbor from the Team plan). At temperature 1.5 the average sample drops to 2.56/8, and the vote pulls it back to 4/8, but not above greedy. That is the correlated-error story from the previous section, measured: when a model is wrong the same way every time, the vote reproduces the error. The timing line varies by machine and load (3 s here; 128 s on a busier run of the same script).
Treat this as a mechanism demo on a tiny model, not evidence about large ones. For a real model, sample the refund scaffold with temperature=0.8 and vote over parse_decision results; the best-of-n script below is the harness for that.
Best-of-n with a verifier or a scorer
Best-of-n samples n answers and picks one with something other than a vote: a verifier (a check that can say whether an answer is right) or a scorer (a model or heuristic that rates answers but can be wrong). Brown et al., "Large Language Monkeys: Scaling Inference Compute with Repeated Sampling" (2024), measured coverage, the share of problems solved by at least one sample, and found it keeps growing over four orders of magnitude of samples: on SWE-bench Lite, DeepSeek-Coder-V2-Instruct went from 15.9 percent with one attempt to 56 percent with 250. They also found that without an automatic verifier, majority voting and reward models plateau beyond several hundred samples. Coverage is only worth something if you can recognize the right sample.
The script below compares three pickers on the 20 refund scenarios. The generator is simulated, with a stated error model: each sample is right with probability 0.9 on easy cases and 0.45 on trap cases (the ones the shortcut gets wrong), and when a trap sample is wrong it gives the shortcut's tempting answer 80 percent of the time. The policy, the scenarios, and the verifier are real.
examples/m05_best_of_n.py:
"""Part B: majority vote vs best-of-n with a verifier vs best-of-n with a noisy scorer, on the 20 refund cases.
The generator is SIMULATED: each sample is right with probability 0.9 on easy
cases and 0.45 on trap cases (cases the one-step shortcut gets wrong), and trap
mistakes usually repeat the shortcut's tempting answer. Only the policy, the
cases, and the verifier are real. Replace the simulator with real samples
(temperature above 0) to measure your model.
"""
import numpy as np
from examples.m05_helpers import CASES, LABELS, decide, shortcut_baseline
rng = np.random.default_rng(11)
REPEATS = 2000
def sample_label(case) -> str:
trap = shortcut_baseline(case.text) != case.gold
if rng.random() < (0.45 if trap else 0.9):
return case.gold
wrong = [label for label in LABELS if label != case.gold]
tempting = shortcut_baseline(case.text)
if trap and rng.random() < 0.8:
return tempting
return str(rng.choice(wrong))
def verifier(case, label: str) -> bool:
"""External check: recompute the decision from the billing record (the case's facts)."""
return label == decide(case.facts)
def noisy_score(case, label: str, sigma: float = 0.8) -> float:
"""A learned scorer: prefers right answers on average, but with noise."""
return float(label == case.gold) + rng.normal(0, sigma)
def run(n: int, repeats: int = REPEATS) -> dict[str, float]:
totals = dict(majority=0, verified=0, escalated=0, scored=0, coverage=0, calls_verified=0)
for _ in range(repeats):
for case in CASES:
samples = [sample_label(case) for _ in range(n)]
counts = {label: samples.count(label) for label in set(samples)}
top = max(counts.values())
winner = str(rng.choice(sorted(label for label, c in counts.items() if c == top)))
totals["majority"] += winner == case.gold
accepted = next((s for s in samples if verifier(case, s)), None)
totals["verified"] += accepted == case.gold
totals["escalated"] += accepted is None
# Sampling one at a time, you can stop at the first sample the verifier accepts.
totals["calls_verified"] += samples.index(accepted) + 1 if accepted is not None else n
totals["scored"] += max(samples, key=lambda s: noisy_score(case, s)) == case.gold
totals["coverage"] += case.gold in samples
return {k: v / (repeats * len(CASES)) for k, v in totals.items()}
if __name__ == "__main__":
print("Share of the 20 cases answered correctly (simulated generator, 2,000 repeats per cell)")
print(f"{'n':>3} {'majority':>9} {'best-of-n+verifier':>19} {'avg calls':>9} {'escalated':>10} {'best-of-n+scorer':>17} {'pass@n':>7}")
for n in (1, 3, 5, 9, 15):
r = run(n)
print(f"{n:>3} {r['majority']:>9.3f} {r['verified']:>19.3f} {r['calls_verified']:>9.2f} {r['escalated']:>10.3f} "
f"{r['scored']:>17.3f} {r['coverage']:>7.3f}")
Code explained
- In simple words: draw n simulated answers per case and pick one by vote, by an exact check against the billing record, or by a noisy score.
- What happens:
verifierrecomputes the decision from the case's billing facts.noisy_scoregives right answers a higher score on average but adds Gaussian noise (sigma 0.8), like a learned reward model. With the verifier, the loop can stop at the first accepted sample, socalls_verifiedcounts calls actually needed. If no sample passes, the case is escalated to a person instead of guessed.pass@n(coverage) is the share of cases where at least one sample was right. Each cell averages 2,000 repeats over the 20 cases. - Comes out: majority voting climbs from 0.695 to about 0.76 and stops: on trap cases the tempting wrong answer wins the vote. Best-of-n with the verifier equals coverage by construction (0.978 at n = 5), and with early stopping it needs only 1.56 calls per case on average, because most cases pass on the first sample. The noisy scorer lands between the two (0.887 at n = 5). The escalation column is the verifier's other gift: at n = 5, 2.2 percent of cases go to a person instead of getting a wrong answer.
Share of the 20 cases answered correctly (simulated generator, 2,000 repeats per cell)
n majority best-of-n+verifier avg calls escalated best-of-n+scorer pass@n
1 0.695 0.695 1.00 0.305 0.695 0.695
3 0.742 0.922 1.45 0.078 0.841 0.922
5 0.758 0.978 1.56 0.022 0.887 0.978
9 0.761 0.997 1.61 0.003 0.923 0.997
15 0.756 1.000 1.61 0.000 0.945 1.000
A fair question: if decide can compute the answer, why sample a model at all? In this toy, you should not; Part C does exactly that and gives the rule to code. In real systems the verifier is usually narrower than the generator. Unit tests can check code but cannot write it. A billing lookup can confirm "this charge is a duplicate" but cannot read a Japanese ticket. A schema can reject a malformed reply but not judge its tone. Best-of-n pays off whenever checking is cheaper or more reliable than generating.
| Situation | Use this | Why |
|---|---|---|
| An exact check exists (tests, a database lookup, a schema, a rule) | Best-of-n with the verifier, stop at the first pass, escalate on none | Accuracy approaches coverage at a fraction of n calls |
| Only a learned or LLM judge exists | Best-of-n with the scorer, small n, spot-check the judge | Gains are real but capped by the scorer's noise (Module 10 evaluates judges) |
| Answers are labels and no check exists | Majority vote | Needs no scorer, but fails on tempting wrong answers |
| Answers are free text and no check exists | One good sample, human review | Neither voting nor noisy scoring is reliable enough to automate |
Ensembling across prompts or models
An ensemble combines different predictors, for example several prompts, several models, or a model and a classical rule. It is majority voting where the voters differ by design, which is the whole point: different predictors are more likely to make different mistakes, and the previous section showed that error correlation, not the number of voters, sets the ceiling.
Three real, deterministic triage classifiers make this measurable on all 72 labelled tickets: the keyword prompt read by keyword_reader, BM25 retrieval of the best help-center article mapped to a category, and a character n-gram logistic regression scored with cross-validation.
examples/m05_ensemble.py:
"""Part B: ensembling three different triage classifiers, and why their error overlap decides the gain.
All three are real, deterministic classifiers (no LLM), so every number here is
a real measurement on the 72 labelled tickets.
"""
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold, cross_val_predict
from sklearn.pipeline import make_pipeline
from examples.m05_helpers import TRIAGE_KEYWORDS, keyword_reader, triage_prompt, wilson
from supportdesk.data import load_tickets
from supportdesk.kb_search import KBSearch
tickets = load_tickets()
gold = [t.gold["category"] for t in tickets]
# Member 1: keyword definitions read from a triage prompt.
prompt = triage_prompt(TRIAGE_KEYWORDS)
keywords = [keyword_reader(prompt, t.text) for t in tickets]
# Member 2: retrieve the best help-center article and map it to a category (hand-written map).
ARTICLE_TO_CATEGORY = {
"account-login": "account_access", "account-sso": "account_access", "data-privacy": "account_access",
"billing-invoices": "billing", "billing-plans": "billing", "billing-refunds": "cancellation",
"boards-automations": "how_to", "exports-data": "how_to", "mobile-app": "how_to",
"integrations-slack": "bug", "status-incidents": "bug", "feature-requests": "feature_request",
}
search = KBSearch()
retrieval = []
for t in tickets:
hits = search.search(t.text, k=1)
retrieval.append(ARTICLE_TO_CATEGORY[hits[0].article_id] if hits else "how_to")
# Member 3: character n-gram TF-IDF + logistic regression, scored with 6-fold cross-validation.
model = make_pipeline(TfidfVectorizer(analyzer="char_wb", ngram_range=(2, 4)), LogisticRegression(max_iter=2000, C=10))
folds = StratifiedKFold(n_splits=6, shuffle=True, random_state=0)
learned = list(cross_val_predict(model, [t.text for t in tickets], gold, cv=folds))
def vote(a: str, b: str, c: str) -> str:
"""Plurality of three; if all three disagree, trust the learned model."""
if a == b or a == c:
return a
if b == c:
return b
return c
members = {"keywords": keywords, "retrieval": retrieval, "learned": learned}
members["ensemble"] = [vote(*x) for x in zip(keywords, retrieval, learned)]
n = len(tickets)
for name, preds in members.items():
k = sum(p == g for p, g in zip(preds, gold))
lo, hi = wilson(k, n)
print(f"{name:<10} {k:>2}/{n} = {k / n:.3f} 95% CI {lo:.2f}-{hi:.2f}")
any_right = sum(g in (a, b, c) for g, a, b, c in zip(gold, keywords, retrieval, learned))
print(f"at least one member right: {any_right}/{n} (the ceiling for any way of combining them)")
print("\nError overlap (tickets both members got wrong) and correlation of their error indicators:")
errors = {name: np.array([p != g for p, g in zip(members[name], gold)]) for name in ("keywords", "retrieval", "learned")}
names = list(errors)
for i in range(3):
for j in range(i + 1, 3):
a, b = errors[names[i]], errors[names[j]]
both = int((a & b).sum())
corr = float(np.corrcoef(a, b)[0, 1])
print(f" {names[i]:<9} & {names[j]:<9} wrong together on {both:>2} (errors {a.sum()} and {b.sum()}), correlation {corr:+.2f}")
Code explained
- In simple words: three different triage methods vote, and we measure how often they fail on the same tickets.
- What happens: member 1 follows the keyword definitions in a triage prompt. Member 2 runs
KBSearch(Module 7's BM25) and maps the top article to a category with a hand-written table. Member 3 is TF-IDF over character 2- to 4-grams plus logistic regression, scored with 6-fold stratified cross-validation so each ticket is predicted by a model that never saw it.votetakes the plurality and trusts the learned model when all three disagree. The last block counts tickets that each pair gets wrong together and the correlation between their error indicators. - Comes out: the members score 50, 49, and 43 of 72; the ensemble scores 51/72. A one-ticket gain is well within noise (the intervals overlap almost completely). The reason is in the overlap table: the error correlations are +0.25 to +0.47, and retrieval and the learned model share 17 errors. The ceiling line matters more than the vote: at least one member is right on 62/72, so a better combiner (for example, routing by language or by confidence) has 11 tickets of room that plain voting cannot reach.
For LLM ensembles, the recipe is the same: run two or three prompt versions or models over your dev set, compute each one's accuracy and their pairwise error overlap, and only then decide. Two prompts on the same model usually have highly correlated errors; a different model family or a classical baseline is often the more useful second voter.
Test-time compute as a tunable dial
Test-time compute is the extra computation you spend per request at inference time instead of on a bigger model: more samples, a verifier, longer reasoning. Snell et al., "Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters" (2024), found that allocating test-time compute according to prompt difficulty was more than 4 times as efficient as a best-of-N baseline, and that on problems where a smaller model already has a reasonable success rate, test-time compute could beat a 14 times larger model in a FLOPs-matched comparison.
For an engineer the practical form is a cost curve: for each setting of the dial, what accuracy do you get and what does 1,000 tickets cost? The script combines measured prompt sizes, real prices from pricing.py, and the simulated accuracies from the best-of-n script.
examples/m05_effort_dial.py:
"""Part B: test-time compute as a dial. What each extra sample or reasoning level costs per 1,000 tickets.
Offline it combines measured prompt sizes and real prices with the SIMULATED
accuracies from m05_best_of_n.py. With --live it measures a reasoning model at
three effort levels on the 20 refund cases:
LLM_PROVIDER=groq GROQ_API_KEY=... PYTHONPATH=. python examples/m05_effort_dial.py --live
"""
import sys
from examples.m05_best_of_n import run
from examples.m05_helpers import CASES, evaluate, prompt_direct, prompt_scaffold
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.tokens import count_messages
MODEL = "openai/gpt-oss-120b"
tokens_in = round(sum(count_messages(prompt_scaffold(c)) for c in CASES) / len(CASES))
OUT_PER_SAMPLE = 60 # assumption: the scaffold's six short lines; replace with your measured usage
print(f"scaffold prompt: {tokens_in} input tokens on average (measured, o200k_base); "
f"{OUT_PER_SAMPLE} output tokens per sample (assumed)\n")
print(f"{'strategy':<24} {'calls':>5} {'USD per 1k':>11} {'accuracy':>9} {'USD per 1k correct':>19}")
for n in (1, 3, 5, 9, 15):
r = run(n, repeats=400)
for name, calls, acc in ((f"majority n={n}", n, r["majority"]),
(f"verify, stop early n<={n}", r["calls_verified"], r["verified"])):
usage = Usage(input_tokens=round(calls * tokens_in), output_tokens=round(calls * OUT_PER_SAMPLE))
per_1k = 1000 * cost_usd(usage, MODEL)
print(f"{name:<24} {calls:>5.2f} {per_1k:>11.3f} {acc:>9.3f} {per_1k / acc:>19.3f}")
# Prompt caching: n samples of one ticket share the whole prompt. Groq's prompt-caching page lists
# 50 percent off cached input for gpt-oss; pricing.py (checked 21 Sep 2026) lists cached_input equal to
# input for these models, so this module applies Groq's documented rate locally instead of editing it.
GROQ_CACHED_RATE = PRICES[MODEL].input * 0.5
print("\nMajority of n when samples 2..n hit the prompt cache (Groq's documented 50 percent cached rate):")
for n in (1, 5, 15):
uncached = Usage(input_tokens=tokens_in, output_tokens=n * OUT_PER_SAMPLE)
cached_in = (n - 1) * tokens_in
per_1k = 1000 * (cost_usd(uncached, MODEL) + cached_in * GROQ_CACHED_RATE / 1_000_000)
full = 1000 * cost_usd(Usage(input_tokens=n * tokens_in, output_tokens=n * OUT_PER_SAMPLE), MODEL)
print(f" n={n:<3} USD per 1k tickets {per_1k:.3f} with cache hits, {full:.3f} without")
print("\nReasoning effort on one call (reasoning tokens are ASSUMED; they bill as output tokens):")
for effort, reasoning in (("low", 300), ("medium", 1200), ("high", 4000)):
usage = Usage(input_tokens=tokens_in, output_tokens=OUT_PER_SAMPLE + reasoning, reasoning_tokens=reasoning)
print(f" {effort:<7} ~{reasoning:>5} reasoning tokens USD per 1k tickets {1000 * cost_usd(usage, MODEL):.3f}")
if "--live" in sys.argv:
from supportdesk.llm import chat
for effort in ("low", "medium", "high"):
result = evaluate(chat, prompt_direct, reasoning_effort=effort, max_tokens=8000)
usage = result.usage
print(result.line(f"effort={effort}"), f" reasoning {usage.reasoning_tokens}",
f" USD {cost_usd(usage, MODEL):.4f}")
Code explained
- In simple words: turn each accuracy setting into dollars per 1,000 tickets and dollars per 1,000 correct answers.
- What happens: the scaffold prompt's average input size is measured with the o200k tokenizer (309 tokens); output per sample is an assumption (60 tokens). For each n it reruns the best-of-n simulation (400 repeats per cell, so accuracies wobble by about a point against the previous table) and prices majority voting (n calls every time) and verify-with-early-stop (the measured average calls). The caching block prices samples 2 to n as cache hits.
pricing.pylists gpt-oss's cached input price equal to its normal input price, while Groq's prompt-caching documentation lists a 50 percent discount for cached input onopenai/gpt-oss-20bandopenai/gpt-oss-120b; the script applies Groq's documented rate locally rather than editing the shared table. The effort block prices assumed reasoning-token counts, which bill as output.--livemeasures the three effort levels for real. - Comes out: majority voting's cost grows linearly (0.082 to 1.235 USD per 1,000 tickets from n = 1 to 15) while its accuracy stalls near 0.75, so the cost per correct answer rises 14-fold. Verify-and-stop costs 0.132 USD per 1,000 at any ceiling of 9 or more, because most cases stop after one or two calls. Cache hits cut majority-of-15 from 1.235 to 0.911, useful but not a change of shape. At the assumed reasoning lengths, high effort costs about 30 times a single scaffold call; only a measured accuracy gain can justify that.text
scaffold prompt: 309 input tokens on average (measured, o200k_base); 60 output tokens per sample (assumed) strategy calls USD per 1k accuracy USD per 1k correct majority n=1 1.00 0.082 0.701 0.118 verify, stop early n<=1 1.00 0.082 0.701 0.118 majority n=3 3.00 0.247 0.743 0.333 verify, stop early n<=3 1.45 0.119 0.928 0.129 majority n=5 5.00 0.412 0.756 0.545 verify, stop early n<=5 1.57 0.129 0.976 0.132 majority n=9 9.00 0.741 0.758 0.978 verify, stop early n<=9 1.60 0.132 0.999 0.132 majority n=15 15.00 1.235 0.751 1.644 verify, stop early n<=15 1.61 0.132 1.000 0.132 Majority of n when samples 2..n hit the prompt cache (Groq's documented 50 percent cached rate): n=1 USD per 1k tickets 0.082 with cache hits, 0.082 without n=5 USD per 1k tickets 0.319 with cache hits, 0.412 without n=15 USD per 1k tickets 0.911 with cache hits, 1.235 without Reasoning effort on one call (reasoning tokens are ASSUMED; they bill as output tokens): low ~ 300 reasoning tokens USD per 1k tickets 0.262 medium ~ 1200 reasoning tokens USD per 1k tickets 0.802 high ~ 4000 reasoning tokens USD per 1k tickets 2.482
Suggested file: m05-test-time-compute-curve.png
| Situation | Use this | Why |
|---|---|---|
| Most requests are easy, a few are hard | A cheap first pass, then more compute only where a check fails or confidence is low | Snell et al.'s point: allocate by difficulty; verify-and-stop is the simple version |
| No verifier, labels, p clearly above 0.5 | Majority of 5, cache the shared prompt | Most of the voting gain for a third of the n = 15 cost |
| A reasoning model is available | Measure low vs medium vs high on your dev set with --live | Effort changes cost several-fold; the accuracy gain is task-specific |
| Latency budget is tight | One call, parallel samples only if needed | Sequential samples and long reasoning both add seconds (Module 3) |