CourseLarge Language Models · Module 4: Prompt Engineering Fundamentals · part 20 of 80
Part 20 · Module 4: Prompt Engineering Fundamentals

Part E: Iteration discipline

10 min read·22 Sept 2026

Changing one thing at a time

If v2 changes the priority definitions and the output format and adds examples, and the score moves, you do not know which change moved it, or whether two changes cancelled out. Each version in this module changes one thing and says so in its note. The diff proves it.

python
"""Version history of the triage prompt: diffs, size of each change, and token deltas."""
from supportdesk.data import load_tickets

from examples.m04_prompts import (ExampleSelector, arrange, count_prompt_tokens, diff_versions, list_versions,
                                  load_prompt)

dev = load_tickets("dev")
ticket = {t.id: t for t in dev}["T-1001"]
examples = arrange(ExampleSelector(dev).balanced(ticket))

print(diff_versions("v1", "v2"))

versions = list_versions()
for old, new in zip(versions, versions[1:]):
    lines = diff_versions(old, new).splitlines()
    body = [ln for ln in lines if not ln.startswith(("---", "+++", "@@")) and "#! note:" not in ln]
    added = sum(ln.startswith("+") for ln in body)
    removed = sum(ln.startswith("-") for ln in body)
    chunks = "\n".join(lines).split("\n@@")[1:]  # one string per hunk
    hunks = sum(any(ln[:1] in "+-" and "#! note:" not in ln for ln in c.splitlines()[1:]) for c in chunks)
    a, b = load_prompt(old), load_prompt(new)
    tok_a = count_prompt_tokens(a.render(ticket, examples if "examples" in a.placeholders else []))
    tok_b = count_prompt_tokens(b.render(ticket, examples if "examples" in b.placeholders else []))
    print(f"{old} -> {new}: +{added} -{removed} lines in {hunks} content hunk(s), tokens {tok_a} -> {tok_b}  | {b.note}")

Code explained

  • In simple words: print the real diff between v1 and v2, then summarize every version-to-version change in lines, content hunks, and tokens.
  • What happens: diff_versions produces a unified diff (- removed, + added, @@ marks a hunk, one contiguous block of changes). The loop counts added and removed lines and the hunks that change content (the note line is excluded), and renders T-1001 with each version to measure tokens (with a balanced example set where the version has a slot).
  • Comes out: v1 to v2 and v2 to v3 each change one contiguous block: one idea per version. v3 to v4 touches three places, which is a sign it bundles several changes; it is the anti-pattern demo from Part F. The token column shows what each idea costs: definitions add about 100 tokens, six examples about 430.

For v3, the only change is the user section, which gains the examples slot:

text
=== user ===
Here are solved tickets with their correct labels:
{{examples}}

Now label this ticket:
{{ticket}}

Code explained

  • In simple words: the user message now has a labeled block for examples, then the ticket.
  • What happens: the renderer fills {{examples}} with the <examples> block from render_examples and {{ticket}} with the escaped ticket. If you pass no examples, the slot renders empty, so v3 with --shots none is a fair zero-shot comparison against v2.
  • Comes out: a prompt about 430 tokens longer with six examples (the table above).

Versioning: every result points at the exact prompt

Each run file starts with a metadata line that includes the version name and its fingerprint. If someone edits v2.txt in place without making a v3, the fingerprint changes and the old runs no longer match it, which is visible in every summary line. Keep old versions forever; they are small, and you will want to rerun them on a new model (below).

Diagnosing a control that moved

Here is a real failure from building this module. The first time the lab ran v3 with the keyword backend, the control collapsed:

text
[4] v3 (6989d7242718) backend=rules shots=balanced k=6 split=dev n=48
  parse failures: 0/48   mean input tokens: 671
  category       4/48  acc 0.083  95% CI [0.033, 0.196]
  priority      11/48  acc 0.229  95% CI [0.133, 0.365]
  needs_human   13/48  acc 0.271  95% CI [0.166, 0.410]

Code explained

  • In simple words: the keyword baseline, which should score the same whatever the prompt, dropped from 38/48 to 4/48 when examples were added.
  • What happens: this is pasted output from that first run (before the fix below). The rules ignore instructions and examples, so only the extracted ticket text can change their answer.
  • Comes out: a paired difference of -0.708 on category, clearly "worse". A tempting reading is "few-shot hurts". The correct reading is "the control moved, so the plumbing is broken".

Read the records before you theorize. The saved run showed the rules classifying the first example's subject, not the ticket. The extractor used a regex that started at the first <ticket tag in the prompt, which in v3 is the first example, and ran to the last </ticket>.

python
"""Diagnose a control that moved: the keyword baseline ignores examples, so v3 should score like v2."""
import re

from supportdesk.data import load_tickets

from examples.m04_eval import last_ticket_text
from examples.m04_prompts import ExampleSelector, arrange, load_prompt

dev = load_tickets("dev")
ticket = dev[0]
messages = load_prompt("v3").render(ticket, arrange(ExampleSelector(dev).balanced(ticket)))

FIRST_VERSION = re.compile(r"<ticket[^>]*>\n(.*?)\n</ticket>\s*$", re.S)  # what the harness used at first
old = FIRST_VERSION.search(messages[-1]["content"]).group(1)
new = last_ticket_text(messages)
print(f"real ticket text : {len(ticket.text):>5} chars")
print(f"first extractor  : {len(old):>5} chars, starts {old[:40]!r}")
print(f"fixed extractor  : {len(new):>5} chars, equal to ticket: {new == ticket.text}")

Code explained

  • In simple words: compare the first version of the ticket extractor with the fixed one on a v3 prompt.
  • What happens: FIRST_VERSION is the original regex, searched from the start of the message. last_ticket_text (fixed) searches from the last <ticket only.
  • Comes out: the first extractor returned 1,771 characters starting with an example's subject ("Gantt dependencies when?"); the fixed one returns the 161-character ticket exactly. After the fix, the rules score 38/48 with v3 as with v1 and v2 (see the lab), and test_last_ticket_is_extracted_after_examples keeps it that way.

    TEXTCopy

    text
    real ticket text :   161 chars
    first extractor  :  1771 chars, starts 'Subject: Gantt dependencies when?\n\nWhen '
    fixed extractor  :   161 chars, equal to ticket: True
    

Without the control, this bug would have shown up as "few-shot examples make the model worse", and you might have deleted a good idea. Keep a prompt-insensitive baseline in every comparison.

Knowing when a prompt has hit its ceiling

A prompt has hit its ceiling when changes stop moving the score by more than the noise. At that point more prompt work is wasted, and the next lever is something else: better examples, retrieval (Module 7), a stronger model, or fine-tuning (Module 9). The plateau function makes that call from data: it takes the last few versions and checks whether any is clearly below the best one, using the paired bootstrap.

python
"""Ceiling detection: when do further changes stop moving the score beyond noise?

The runs here use the copy-the-examples stand-in with more and more similar
examples, because it gives real, prompt-dependent scores without an API key.
With a key, pass your v1..vN runs from the llm backend instead.
"""
from examples.m04_eval import compare, copy_backend, plateau, run_eval

runs = []
for k in (1, 3, 5, 8, 10):
    run = run_eval(copy_backend(), "v3", "dev", "similar", k, backend="copy")
    run.meta["version"] = f"similar-k{k}"
    runs.append(run)
    print(f"{run.meta['version']:<12} " + "  ".join(f"{label} {sum(run.correct(label)):>2}/48"
                                                  for label in ("category", "priority", "needs_human")))

print(compare(runs[0], runs[2]))
for label in ("category", "needs_human"):
    done, why = plateau(runs, label, window=3)
    print(f"plateau on {label}: {done}  {why}")

Code explained

  • In simple words: try five "versions" that show more and more similar examples, and ask whether the later ones are distinguishable from the best.
  • What happens: each run uses the copy stand-in with k = 1, 3, 5, 8, 10 similar examples, because it gives real, prompt-dependent scores without a key. compare shows k=5 against k=1 in detail; plateau checks the last three runs against the best for two labels.
  • Comes out: category scores wander between 21 and 28 of 48. The drop from k=1 to k=5 is 7 tickets (-0.146), and its interval runs from -0.292 to exactly 0.000: right on the edge, so the harness calls it within noise. That is the honest answer at n=48: a 15-point drop is not provable with 48 tickets. The plateau check says the last three versions are within noise of the best for both labels, so this line of changes has hit its ceiling. With a real model, feed plateau your v1 to vN runs instead.
text
similar-k1   category 28/48  priority 17/48  needs_human 35/48
similar-k3   category 26/48  priority 16/48  needs_human 37/48
similar-k5   category 21/48  priority 16/48  needs_human 37/48
similar-k8   category 26/48  priority 16/48  needs_human 37/48
similar-k10  category 23/48  priority 15/48  needs_human 36/48
similar-k5/copy minus similar-k1/copy (paired, n=48)
  category     -0.146  95% CI [-0.292, +0.000]  within noise
  priority     -0.021  95% CI [-0.146, +0.104]  within noise
  needs_human  +0.042  95% CI [-0.062, +0.146]  within noise
plateau on category: True  last 3 versions within noise of best (similar-k1): [('similar-k5', 0.146, 0.0, 0.292), ('similar-k8', 0.042, -0.104, 0.188), ('similar-k10', 0.104, -0.042, 0.25)]
plateau on needs_human: True  last 3 versions within noise of best (similar-k3): [('similar-k5', 0.0, -0.104, 0.104), ('similar-k8', 0.0, -0.104, 0.104), ('similar-k10', 0.021, -0.083, 0.125)]

How big a difference can n=48 detect? Roughly, the paired interval above has a half-width of about 0.15 for a disagreement of 7 tickets. To detect a 5-point improvement reliably you need several hundred labeled tickets. Until you have them, treat small wins as "not yet shown", and grow the labeled set from real traffic (Module 10).

SituationUse thisWhy
Last 3 versions within noise of the bestStop prompt work on this label; change leverMore wording changes are unlikely to show a real gain
A version is clearly better (interval above 0)Keep it and continue one change at a timeThe gain is larger than the noise
Differences are on the edge of the intervalLabel more tickets before decidingn is too small to separate the versions
Per-label scores disagree (one up, one down)Decide which label matters more to the businessAveraging hides a tradeoff Maya needs to see

Transferring prompts across models, and why they degrade

A prompt tuned on one model often scores lower on another. Models are trained on different data and instruction formats, so they react differently to the same wording, label names, example order, and formatting (the Sclar et al. result above is one measurement of how large this can be). Some mechanics differ too: prefill support (Part D), whether reasoning_effort is accepted, and the tokenizer, which changes cost and how close you are to limits. The tokenizer difference is measurable here:

python
"""Moving a prompt to another model family: the same text is a different number of tokens."""
from supportdesk.data import load_tickets
from supportdesk.tokens import ENCODINGS, count_messages

from examples.m04_prompts import ExampleSelector, arrange, load_prompt

dev = load_tickets("dev")
by_id = {t.id: t for t in dev}
selector = ExampleSelector(dev)

print(f"{'prompt':<22}" + "".join(f"{e:>15}" for e in ENCODINGS))
for version, ticket_id in [("v2", "T-1001"), ("v2", "T-1034"), ("v3", "T-1001"), ("v3", "T-1034")]:
    t = load_prompt(version)
    ticket = by_id[ticket_id]
    examples = arrange(selector.balanced(ticket)) if "examples" in t.placeholders else []
    messages = t.render(ticket, examples)
    counts = [count_messages(messages, encoding=e) for e in ENCODINGS]
    print(f"{version + ' ' + ticket_id + ' (' + ticket.language + ')':<22}" + "".join(f"{c:>15}" for c in counts))

Code explained

  • In simple words: count the same rendered prompts with four real tokenizers from different model families.
  • What happens: v2 and v3 are rendered for an English and a Japanese ticket (v3 with a balanced example set) and counted with each encoding in supportdesk.tokens.ENCODINGS.
  • Comes out: the same v3 prompt is 707 tokens under o200k_base and 827 under p50k_base, 17% more. Budgets and cost estimates tuned on one tokenizer do not carry over. The accuracy side of transfer needs real models: rerun the harness per model.
text
prompt                      p50k_base    cl100k_base     o200k_base  claude-legacy
v2 T-1001 (en)                    308            278            281            305
v2 T-1034 (ja)                    322            277            271            303
v3 T-1001 (en)                    827            705            707            801
v3 T-1034 (ja)                    804            670            668            764

The transfer procedure that follows from this module:

  1. Keep every version file and the frozen set unchanged.
  2. Run the current best version on the new model with LLM_PROVIDER and LLM_MODEL set, and also run the previous two versions. The best version on the old model is not necessarily best on the new one.
  3. Compare per label with compare, and read the failed records, especially non-English tickets.
  4. Check provider mechanics: prefill, stop sequences, JSON mode, reasoning settings.
  5. Only then start a new line of versions for the new model, one change at a time.