Part E: Iteration discipline
Changing one thing at a time
If v2 changes the priority definitions and the output format and adds examples, and the score moves, you do not know which change moved it, or whether two changes cancelled out. Each version in this module changes one thing and says so in its note. The diff proves it.
"""Version history of the triage prompt: diffs, size of each change, and token deltas."""
from supportdesk.data import load_tickets
from examples.m04_prompts import (ExampleSelector, arrange, count_prompt_tokens, diff_versions, list_versions,
load_prompt)
dev = load_tickets("dev")
ticket = {t.id: t for t in dev}["T-1001"]
examples = arrange(ExampleSelector(dev).balanced(ticket))
print(diff_versions("v1", "v2"))
versions = list_versions()
for old, new in zip(versions, versions[1:]):
lines = diff_versions(old, new).splitlines()
body = [ln for ln in lines if not ln.startswith(("---", "+++", "@@")) and "#! note:" not in ln]
added = sum(ln.startswith("+") for ln in body)
removed = sum(ln.startswith("-") for ln in body)
chunks = "\n".join(lines).split("\n@@")[1:] # one string per hunk
hunks = sum(any(ln[:1] in "+-" and "#! note:" not in ln for ln in c.splitlines()[1:]) for c in chunks)
a, b = load_prompt(old), load_prompt(new)
tok_a = count_prompt_tokens(a.render(ticket, examples if "examples" in a.placeholders else []))
tok_b = count_prompt_tokens(b.render(ticket, examples if "examples" in b.placeholders else []))
print(f"{old} -> {new}: +{added} -{removed} lines in {hunks} content hunk(s), tokens {tok_a} -> {tok_b} | {b.note}")
Code explained
- In simple words: print the real diff between v1 and v2, then summarize every version-to-version change in lines, content hunks, and tokens.
- What happens:
diff_versionsproduces a unified diff (-removed,+added,@@marks a hunk, one contiguous block of changes). The loop counts added and removed lines and the hunks that change content (the note line is excluded), and renders T-1001 with each version to measure tokens (with a balanced example set where the version has a slot). - Comes out: v1 to v2 and v2 to v3 each change one contiguous block: one idea per version. v3 to v4 touches three places, which is a sign it bundles several changes; it is the anti-pattern demo from Part F. The token column shows what each idea costs: definitions add about 100 tokens, six examples about 430.
For v3, the only change is the user section, which gains the examples slot:
=== user ===
Here are solved tickets with their correct labels:
{{examples}}
Now label this ticket:
{{ticket}}
Code explained
- In simple words: the user message now has a labeled block for examples, then the ticket.
- What happens: the renderer fills
{{examples}}with the<examples>block fromrender_examplesand{{ticket}}with the escaped ticket. If you pass no examples, the slot renders empty, so v3 with--shots noneis a fair zero-shot comparison against v2. - Comes out: a prompt about 430 tokens longer with six examples (the table above).
Versioning: every result points at the exact prompt
Each run file starts with a metadata line that includes the version name and its fingerprint. If someone edits v2.txt in place without making a v3, the fingerprint changes and the old runs no longer match it, which is visible in every summary line. Keep old versions forever; they are small, and you will want to rerun them on a new model (below).
Diagnosing a control that moved
Here is a real failure from building this module. The first time the lab ran v3 with the keyword backend, the control collapsed:
[4] v3 (6989d7242718) backend=rules shots=balanced k=6 split=dev n=48
parse failures: 0/48 mean input tokens: 671
category 4/48 acc 0.083 95% CI [0.033, 0.196]
priority 11/48 acc 0.229 95% CI [0.133, 0.365]
needs_human 13/48 acc 0.271 95% CI [0.166, 0.410]
Code explained
- In simple words: the keyword baseline, which should score the same whatever the prompt, dropped from 38/48 to 4/48 when examples were added.
- What happens: this is pasted output from that first run (before the fix below). The rules ignore instructions and examples, so only the extracted ticket text can change their answer.
- Comes out: a paired difference of -0.708 on category, clearly "worse". A tempting reading is "few-shot hurts". The correct reading is "the control moved, so the plumbing is broken".
Read the records before you theorize. The saved run showed the rules classifying the first example's subject, not the ticket. The extractor used a regex that started at the first <ticket tag in the prompt, which in v3 is the first example, and ran to the last </ticket>.
"""Diagnose a control that moved: the keyword baseline ignores examples, so v3 should score like v2."""
import re
from supportdesk.data import load_tickets
from examples.m04_eval import last_ticket_text
from examples.m04_prompts import ExampleSelector, arrange, load_prompt
dev = load_tickets("dev")
ticket = dev[0]
messages = load_prompt("v3").render(ticket, arrange(ExampleSelector(dev).balanced(ticket)))
FIRST_VERSION = re.compile(r"<ticket[^>]*>\n(.*?)\n</ticket>\s*$", re.S) # what the harness used at first
old = FIRST_VERSION.search(messages[-1]["content"]).group(1)
new = last_ticket_text(messages)
print(f"real ticket text : {len(ticket.text):>5} chars")
print(f"first extractor : {len(old):>5} chars, starts {old[:40]!r}")
print(f"fixed extractor : {len(new):>5} chars, equal to ticket: {new == ticket.text}")
Code explained
- In simple words: compare the first version of the ticket extractor with the fixed one on a v3 prompt.
- What happens:
FIRST_VERSIONis the original regex, searched from the start of the message.last_ticket_text(fixed) searches from the last<ticketonly. - Comes out: the first extractor returned 1,771 characters starting with an example's subject ("Gantt dependencies when?"); the fixed one returns the 161-character ticket exactly. After the fix, the rules score 38/48 with v3 as with v1 and v2 (see the lab), and
test_last_ticket_is_extracted_after_exampleskeeps it that way.TEXTCopy
textreal ticket text : 161 chars first extractor : 1771 chars, starts 'Subject: Gantt dependencies when?\n\nWhen ' fixed extractor : 161 chars, equal to ticket: True
Without the control, this bug would have shown up as "few-shot examples make the model worse", and you might have deleted a good idea. Keep a prompt-insensitive baseline in every comparison.
Knowing when a prompt has hit its ceiling
A prompt has hit its ceiling when changes stop moving the score by more than the noise. At that point more prompt work is wasted, and the next lever is something else: better examples, retrieval (Module 7), a stronger model, or fine-tuning (Module 9). The plateau function makes that call from data: it takes the last few versions and checks whether any is clearly below the best one, using the paired bootstrap.
"""Ceiling detection: when do further changes stop moving the score beyond noise?
The runs here use the copy-the-examples stand-in with more and more similar
examples, because it gives real, prompt-dependent scores without an API key.
With a key, pass your v1..vN runs from the llm backend instead.
"""
from examples.m04_eval import compare, copy_backend, plateau, run_eval
runs = []
for k in (1, 3, 5, 8, 10):
run = run_eval(copy_backend(), "v3", "dev", "similar", k, backend="copy")
run.meta["version"] = f"similar-k{k}"
runs.append(run)
print(f"{run.meta['version']:<12} " + " ".join(f"{label} {sum(run.correct(label)):>2}/48"
for label in ("category", "priority", "needs_human")))
print(compare(runs[0], runs[2]))
for label in ("category", "needs_human"):
done, why = plateau(runs, label, window=3)
print(f"plateau on {label}: {done} {why}")
Code explained
- In simple words: try five "versions" that show more and more similar examples, and ask whether the later ones are distinguishable from the best.
- What happens: each run uses the copy stand-in with k = 1, 3, 5, 8, 10 similar examples, because it gives real, prompt-dependent scores without a key.
compareshows k=5 against k=1 in detail;plateauchecks the last three runs against the best for two labels. - Comes out: category scores wander between 21 and 28 of 48. The drop from k=1 to k=5 is 7 tickets (-0.146), and its interval runs from -0.292 to exactly 0.000: right on the edge, so the harness calls it within noise. That is the honest answer at n=48: a 15-point drop is not provable with 48 tickets. The plateau check says the last three versions are within noise of the best for both labels, so this line of changes has hit its ceiling. With a real model, feed
plateauyour v1 to vN runs instead.
similar-k1 category 28/48 priority 17/48 needs_human 35/48
similar-k3 category 26/48 priority 16/48 needs_human 37/48
similar-k5 category 21/48 priority 16/48 needs_human 37/48
similar-k8 category 26/48 priority 16/48 needs_human 37/48
similar-k10 category 23/48 priority 15/48 needs_human 36/48
similar-k5/copy minus similar-k1/copy (paired, n=48)
category -0.146 95% CI [-0.292, +0.000] within noise
priority -0.021 95% CI [-0.146, +0.104] within noise
needs_human +0.042 95% CI [-0.062, +0.146] within noise
plateau on category: True last 3 versions within noise of best (similar-k1): [('similar-k5', 0.146, 0.0, 0.292), ('similar-k8', 0.042, -0.104, 0.188), ('similar-k10', 0.104, -0.042, 0.25)]
plateau on needs_human: True last 3 versions within noise of best (similar-k3): [('similar-k5', 0.0, -0.104, 0.104), ('similar-k8', 0.0, -0.104, 0.104), ('similar-k10', 0.021, -0.083, 0.125)]
How big a difference can n=48 detect? Roughly, the paired interval above has a half-width of about 0.15 for a disagreement of 7 tickets. To detect a 5-point improvement reliably you need several hundred labeled tickets. Until you have them, treat small wins as "not yet shown", and grow the labeled set from real traffic (Module 10).
| Situation | Use this | Why |
|---|---|---|
| Last 3 versions within noise of the best | Stop prompt work on this label; change lever | More wording changes are unlikely to show a real gain |
| A version is clearly better (interval above 0) | Keep it and continue one change at a time | The gain is larger than the noise |
| Differences are on the edge of the interval | Label more tickets before deciding | n is too small to separate the versions |
| Per-label scores disagree (one up, one down) | Decide which label matters more to the business | Averaging hides a tradeoff Maya needs to see |
Transferring prompts across models, and why they degrade
A prompt tuned on one model often scores lower on another. Models are trained on different data and instruction formats, so they react differently to the same wording, label names, example order, and formatting (the Sclar et al. result above is one measurement of how large this can be). Some mechanics differ too: prefill support (Part D), whether reasoning_effort is accepted, and the tokenizer, which changes cost and how close you are to limits. The tokenizer difference is measurable here:
"""Moving a prompt to another model family: the same text is a different number of tokens."""
from supportdesk.data import load_tickets
from supportdesk.tokens import ENCODINGS, count_messages
from examples.m04_prompts import ExampleSelector, arrange, load_prompt
dev = load_tickets("dev")
by_id = {t.id: t for t in dev}
selector = ExampleSelector(dev)
print(f"{'prompt':<22}" + "".join(f"{e:>15}" for e in ENCODINGS))
for version, ticket_id in [("v2", "T-1001"), ("v2", "T-1034"), ("v3", "T-1001"), ("v3", "T-1034")]:
t = load_prompt(version)
ticket = by_id[ticket_id]
examples = arrange(selector.balanced(ticket)) if "examples" in t.placeholders else []
messages = t.render(ticket, examples)
counts = [count_messages(messages, encoding=e) for e in ENCODINGS]
print(f"{version + ' ' + ticket_id + ' (' + ticket.language + ')':<22}" + "".join(f"{c:>15}" for c in counts))
Code explained
- In simple words: count the same rendered prompts with four real tokenizers from different model families.
- What happens: v2 and v3 are rendered for an English and a Japanese ticket (v3 with a balanced example set) and counted with each encoding in
supportdesk.tokens.ENCODINGS. - Comes out: the same v3 prompt is 707 tokens under o200k_base and 827 under p50k_base, 17% more. Budgets and cost estimates tuned on one tokenizer do not carry over. The accuracy side of transfer needs real models: rerun the harness per model.
prompt p50k_base cl100k_base o200k_base claude-legacy
v2 T-1001 (en) 308 278 281 305
v2 T-1034 (ja) 322 277 271 303
v3 T-1001 (en) 827 705 707 801
v3 T-1034 (ja) 804 670 668 764
The transfer procedure that follows from this module:
- Keep every version file and the frozen set unchanged.
- Run the current best version on the new model with
LLM_PROVIDERandLLM_MODELset, and also run the previous two versions. The best version on the old model is not necessarily best on the new one. - Compare per label with
compare, and read the failed records, especially non-English tickets. - Check provider mechanics: prefill, stop sequences, JSON mode, reasoning settings.
- Only then start a new line of versions for the new model, one change at a time.