Part 7: Verification
Held-out evaluation designed before training
We already did the most important part in Part 2: the metric, the test split, the ship rule, and a fingerprint of the test ids were written to data/m09/eval_plan.json before any training. Three habits kept it honest afterwards:
- The test set was scored once per final model, after the sweep had chosen settings using cross-validation on dev only.
- The contamination check ran against the exact training rows (Part 3), not against "the dev split in general".
- The ship rule compares against the best cheap baseline, not against the untuned model. Beating the base model is easy; beating keyword rules is the real bar.
The one thing we could not fix is sample size. With 24 test tickets, a 95% interval is about 30 to 40 points wide. A real triage eval should have hundreds of tickets per language, and Module 10 is about building exactly that. Until then, compare systems on the same tickets with a paired test (below), which squeezes more out of a small set than comparing two intervals.
Regression testing general capability after narrow tuning
A fine-tune is trained to do one thing better. The question nobody asks until production is: what did it get worse at? Catastrophic forgetting is the name for a model losing abilities it had because training pushed its weights toward a narrow task. You notice it only if you measure the old abilities before and after. For TinyLM, "general capability" means three things we can measure:
- Perplexity on the corpus validation text: the held-out 5% of its pretraining data. This is its general language modeling.
- Perplexity on the KB articles: Brightlane's own writing.
- Fact recall: greedy replies to the 12 standard support questions from Part 5. Does it still state the correct policy?
"""Regression tests after narrow tuning: did the triage models forget how to be a support-domain LM?"""
from examples.m09_dpo import FACTS, fixed_prompt
from examples.m09_lora import add_lora, with_adapter
from examples.m09_setup import (MODELS, corpus_validation_text, load_rows, seed_everything, sft_train,
to_pairs)
from supportdesk.data import load_articles
from supportdesk.tinylm import SamplingParams, generate, load, perplexity
def fact_recall(model, tok) -> int:
"""How many of the 12 standard support questions still get the correct policy fact (greedy)."""
hits = 0
for i, fact in enumerate(FACTS):
g = generate(model, tok, fixed_prompt(i), SamplingParams(max_new_tokens=40, temperature=0, stop=["\n"]))
hits += fact[2].lower() in g.text.lower()
return hits
if __name__ == "__main__":
base, tok = load()
val_text = corpus_validation_text(tok)
kb_text = "\n".join(f"# {a.title}\n{a.body}" for a in load_articles())
seed_everything(0)
gentle, _ = load() # same LoRA recipe with a smaller total update: lr 1e-3 for 4 epochs (not saved)
add_lora(gentle, r=4, alpha=8)
sft_train(gentle, tok, to_pairs(tok, load_rows("dev_real") + load_rows("dev_aug")), epochs=4, lr=1e-3)
models = {"base": base, "triage full FT": load(MODELS / "m09-full")[0],
"triage LoRA": with_adapter("m09-lora-triage"), "triage LoRA gentle": gentle,
"reply DPO (LoRA)": with_adapter("m09-reply-dpo")}
ref = {}
print(f"{'model':<19} {'corpus val ppl':>15} {'KB ppl':>9} {'fact recall':>12}")
for name, m in models.items():
v, k, f = perplexity(m, tok, val_text), perplexity(m, tok, kb_text), fact_recall(m, tok)
ref = ref or {"v": v, "k": k}
print(f"{name:<19} {v:>8.2f} ({v / ref['v'] - 1:+5.0%}) {k:>6.2f} ({k / ref['k'] - 1:+4.0%}) {f:>8}/12")
probe = generate(models["triage LoRA"], tok, fixed_prompt(1), SamplingParams(max_new_tokens=30, temperature=0,
stop=["\n"]))
print("triage LoRA answering question 2:", repr(probe.text))
Code explained
- In simple words: run the same three checkups on the base model and every fine-tuned model, and compare.
- What happens:
fact_recallgenerates a greedy reply to each of the 12 questions and checks for the verified phrase. The main block builds five models: the base, the full fine-tune and LoRA adapter from Part 4, a "gentle" LoRA trained here with a smaller total update (lr 1e-3 for 4 epochs instead of 3e-3 for 12; not saved), and the reply DPO adapter. It prints each model's perplexities with the percentage change from base, and finally one reply from the triage LoRA to question 2 ("how do I cancel my Team subscription?"). Run it withpython -m examples.m09_forgetting(about 15 seconds). - Comes out:
model corpus val ppl KB ppl fact recall
base 1.65 ( +0%) 7.93 ( +0%) 12/12
triage full FT 1.71 ( +4%) 16.43 (+107%) 12/12
triage LoRA 3.12 ( +89%) 118.70 (+1397%) 4/12
triage LoRA gentle 1.81 ( +9%) 108.51 (+1269%) 12/12
reply DPO (LoRA) 2.10 ( +27%) 47.87 (+504%) 12/12
triage LoRA answering question 2: ' Invoices are emailed to'
- The triage LoRA forgot the most. Corpus perplexity is up 89%, KB perplexity is up about 15 times, and fact recall fell from 12/12 to 4/12. Asked how to cancel, it answers about invoices.
- The full fine-tune forgot the least here: +4% on the corpus and 12/12 facts. This is the opposite of the usual expectation. Biderman et al., "LoRA Learns Less and Forgets Less" (2024) found that LoRA generally forgets less than full fine-tuning. The explanation is in the settings, not the method: the chosen LoRA used a learning rate 30 times higher for 50% more epochs. The "gentle" LoRA row, with a smaller total update, is back to +9% on the corpus and 12/12 facts. Forgetting tracks how far you move the weights, and the method only sets the default distance. Measure it for your settings; do not assume it from the method name.
- KB perplexity is the most sensitive alarm. Even the gentle LoRA raises it by about 14 times, while its corpus perplexity barely moves. A single general metric can hide damage on the text you care most about.
- The reply DPO adapter shifted the style on purpose and paid +27% corpus perplexity for it, with facts intact. Whether that trade is acceptable is a decision for the eval plan, not for the training script.
When a regression shows up, find out where before deciding what to do. This script breaks the triage LoRA's extra loss on KB text down by position:
"""Where did the triage LoRA get worse on KB text? Sum the extra loss by the token that came before."""
from collections import Counter
import torch
import torch.nn.functional as F
from examples.m09_lora import with_adapter
from supportdesk.data import load_articles
from supportdesk.tinylm import load
base, tok = load()
tuned = with_adapter("m09-lora-triage")
ids = torch.tensor(tok.encode("\n".join(f"# {a.title}\n{a.body}" for a in load_articles())).ids)
extra, total = Counter(), 0.0
with torch.no_grad():
for start in range(0, len(ids) - 1, 128):
chunk = ids[start: start + 129]
x, y = chunk[:-1].unsqueeze(0), chunk[1:]
delta = F.cross_entropy(tuned(x)[0], y, reduction="none") - F.cross_entropy(base(x)[0], y, reduction="none")
for prev, d in zip(x[0].tolist(), delta.tolist()):
extra[tok.decode([prev])] += d
total += delta.sum().item()
print(f"extra loss over {len(ids) - 1} KB tokens: {total:.0f} nats ({total / (len(ids) - 1):.2f} per token)")
for prev, d in extra.most_common(6):
print(f" after {prev!r:<8} {d:6.0f} nats ({d / total:.0%})")
Code explained
- In simple words: for every token of KB text, compare how surprised the base and tuned models are, and add up the difference by the token that came just before.
- What happens: it runs the base and the triage LoRA over the concatenated KB articles in 128-token chunks, computes per-token loss for each, and accumulates
tuned - basein a counter keyed by the preceding token. Run it withpython -m examples.m09_forgetting_where. - Comes out:
extra loss over 1172 KB tokens: 3172 nats (2.71 per token)
after ',' 248 nats (8%)
after '\n' 198 nats (6%)
after ' the' 177 nats (6%)
after '.' 172 nats (5%)
after ' and' 105 nats (3%)
after ' to' 73 nats (2%)
The extra loss is 2.71 nats per token (exp of that is the roughly 15-times perplexity increase), and the biggest single contributor accounts for only 8%. It is spread across ordinary positions: after commas, newlines, "the", and "and". That is what forgetting looks like. There is no one broken phrase to patch; the model got a little worse at language everywhere. The remedies are the ones you have already seen: a smaller total update (lower learning rate, fewer epochs), replay of general data during training (Part 4), or a larger and more varied training set.
Deciding the fine-tune was not worth it, and saying so
The last step is a direct comparison on the same 24 tickets. Two systems that are each right about 60% of the time might be right on the same tickets or on completely different ones, so comparing their intervals wastes information. A paired test looks only at the tickets where they disagree. The exact McNemar test asks: among the tickets where exactly one system is right, is the split more lopsided than a fair coin would produce?
examples/m09_worth_it.py
"""Was the fine-tune worth it? Paired comparison against the cheapest alternatives, plus speed."""
import json
import time
from math import comb
from examples.m09_cheap_baselines import clf, gold, keyword_label, retrieval_label, test
from examples.m09_lora import with_adapter
from examples.m09_setup import MODELS, classify, prompt_for, wilson
from supportdesk.tinylm import load
def mcnemar_exact(a_right, b_right):
"""Two-sided exact McNemar test on paired right/wrong outcomes: only disagreements count."""
b = sum(x and not y for x, y in zip(a_right, b_right))
c = sum(y and not x for x, y in zip(a_right, b_right))
n = b + c
p = min(1.0, 2 * sum(comb(n, k) for k in range(min(b, c) + 1)) / 2 ** n) if n else 1.0
return b, c, p
if __name__ == "__main__":
preds = json.loads((MODELS / "m09-test-preds.json").read_text())
systems = {"keyword rules": [keyword_label(t.text) for t in test],
"tf-idf + logreg": list(clf.predict([t.text for t in test])),
"retrieval (KB top-1)": [retrieval_label(t.text) for t in test],
"TinyLM full FT": preds["full"], "TinyLM LoRA": preds["lora"]}
right = {name: [p == g for p, g in zip(ps, gold)] for name, ps in systems.items()}
for name, r in right.items():
lo, hi = wilson(sum(r), len(r))
print(f"{name:<22} {sum(r):>2}/24 (95% CI {lo:.0%} to {hi:.0%})")
for other in ("TinyLM LoRA", "TinyLM full FT", "tf-idf + logreg"):
b, c, p = mcnemar_exact(right["keyword rules"], right[other])
print(f"rules vs {other:<16}: rules alone right on {b}, {other} alone right on {c}, exact McNemar p = {p:.3f}")
model = with_adapter("m09-lora-triage")
tok = load()[1]
started = time.perf_counter()
for t in test:
keyword_label(t.text)
rules_ms = (time.perf_counter() - started) * 1000 / len(test)
started = time.perf_counter()
for t in test:
classify(model, tok, prompt_for(tok, t))
lm_ms = (time.perf_counter() - started) * 1000 / len(test)
print(f"latency per ticket: keyword rules {rules_ms:.3f} ms, TinyLM LoRA {lm_ms:.1f} ms")
Code explained
- In simple words: line up every system's answers ticket by ticket, test whether the keyword rules really beat each alternative, and time them.
- What happens:
mcnemar_exactcounts b (only system A right) and c (only system B right), then computes the two-sided binomial probability of a split at least that lopsided if both were equally good. It reuses the baselines fromm09_cheap_baselines.pyand the saved test predictions fromm09_train.py, and times the keyword rules and the LoRA model (log-probability scoring) per ticket. Run it withpython -m examples.m09_worth_it. - Comes out:
keyword rules 15/24 (95% CI 43% to 79%)
tf-idf + logreg 12/24 (95% CI 31% to 69%)
retrieval (KB top-1) 12/24 (95% CI 31% to 69%)
TinyLM full FT 5/24 (95% CI 9% to 40%)
TinyLM LoRA 6/24 (95% CI 12% to 45%)
rules vs TinyLM LoRA : rules alone right on 10, TinyLM LoRA alone right on 1, exact McNemar p = 0.012
rules vs TinyLM full FT : rules alone right on 14, TinyLM full FT alone right on 4, exact McNemar p = 0.031
rules vs tf-idf + logreg : rules alone right on 6, tf-idf + logreg alone right on 3, exact McNemar p = 0.508
latency per ticket: keyword rules 0.007 ms, TinyLM LoRA 10.8 ms
The verdict, with numbers:
- Keyword rules are right on 15/24; the LoRA fine-tune on 6/24. On the tickets where they disagree, the rules win 10 to 1 (p = 0.012). The full fine-tune loses 14 to 4 (p = 0.031). Both differences are real, not noise, and both go the wrong way for the fine-tune.
- Rules against TF-IDF (15 against 12, disagreements 6 to 3, p = 0.51) is noise. We cannot claim the rules are better than TF-IDF, only that neither is worse than the fine-tune.
- The rules run in about 0.01 ms per ticket, the LoRA model in about 11 ms, over 1,000 times slower, and TinyLM is about 100,000 times smaller than a small production model.
- And from the regression checks, the LoRA damaged the model's general ability (+89% corpus perplexity).
- So the decision, written the way Maya should read it: "We fine-tuned TinyLM for triage with LoRA and with full fine-tuning. Neither beat the keyword rules on the held-out tickets (6/24 and 5/24 against 15/24, p < 0.05 in favor of the rules), and the LoRA version damaged general ability. We are shipping the keyword rules, with TF-IDF as the next candidate. We will revisit fine-tuning when we have a stronger base model and at least a few hundred labeled tickets per language, and we will measure a prompted hosted model first." Writing that paragraph is part of the job. A fine-tune that does not beat the cheapest alternative on a pre-registered eval is a cost, not an asset, however much work went into it.
This result is about TinyLM and 48 tickets. With a real 7B base model and a few hundred tickets, a LoRA triage model very plausibly beats keyword rules, and with a strong hosted model, zero-shot prompting may beat both. The procedure is what you keep: the same plan, baselines, contamination check, paired test, and regression checks decide it either way.
Tests for the building blocks
The module's pieces have fast unit tests (no full training runs), so you can change the code without silently breaking the mechanics:
python -m pytest -q tests/test_m09_fine_tuning.py
Code explained
- In simple words: check the mechanics that every result above depends on.
- What happens: the tests check that a fresh adapter is a no-op and only LoRA parameters train (exactly 32,768 at rank 4), that merging preserves outputs, that adapters round-trip through disk, that the adapter server's padding is exact, that log-probability scores match the model's next-token distribution, that the greedy-mask trap really picks the lowest-id label, that the loss mask covers only the answer, the Wilson interval, the contamination and dedupe checks, that DPO's loss is log 2 when the policy equals the reference, that 8-bit and 4-bit quantization stay close to the original weights, and the reward, style, parsing, and log-check functions.
- Comes out:
.............. [100%]
14 passed in 3.06s
Module Lab
The lab runs the whole adaptation loop in one script, from the frozen plan to a written decision. It reuses every building block: the eval plan and fingerprint (Part 2), the data pipeline and contamination check (Part 3), LoRA and the sweep's chosen settings (Part 4), the paired test and regression checks (Part 7).
examples/m09_lab.py
"""Module 9 lab: the whole adaptation loop for ticket triage, ending in a written ship / no-ship decision.
Steps: check the frozen eval plan, build the SFT set from dev only, check contamination, train a LoRA
adapter, score it on test against the keyword baseline, run the regression checks, decide by the plan.
"""
import hashlib
import json
from examples.m09_cheap_baselines import keyword_label
from examples.m09_data import augment, contamination, dedupe, row, scrub
from examples.m09_forgetting import fact_recall
from examples.m09_lora import add_lora, save_adapter
from examples.m09_setup import (DATA_OUT, MODELS, corpus_validation_text, count_trainable, evaluate, seed_everything,
sft_train, to_pairs, wilson)
from examples.m09_worth_it import mcnemar_exact
from supportdesk.data import load_tickets
from supportdesk.tinylm import load, perplexity
# 1. The plan was written before any training. Refuse to run if the test set changed since.
plan = json.loads((DATA_OUT / "eval_plan.json").read_text())
dev, test = load_tickets("dev"), load_tickets("test")
fingerprint = hashlib.sha256(",".join(t.id for t in test).encode()).hexdigest()[:16]
assert fingerprint == plan["test_ids_sha256"], "test set changed since the eval plan was frozen"
print(f"1. eval plan {fingerprint} ok: {plan['metric']} on {plan['n']} test tickets")
# 2. Data: dev only, scrubbed, augmented, deduplicated, checked against test.
real = [row(scrub(t.subject)[0], scrub(t.body)[0], t.gold["category"], "real") for t in dev]
rows = dedupe(real + [a for r in real for a in augment(r)])
leaks = contamination(rows, test)
assert not leaks, leaks
print(f"2. training rows: {len(rows)} ({len(real)} real), contamination with test: none")
# 3. Train a LoRA adapter with the settings the cross-validated sweep chose.
cfg = json.loads((MODELS / "m09-sweep.json").read_text())["chosen"]["lora"]
seed_everything(0)
base, tok = load()
model, _ = load()
add_lora(model, r=cfg["rank"], alpha=2 * cfg["rank"])
info = sft_train(model, tok, to_pairs(tok, rows), epochs=cfg["best_epoch"], lr=cfg["lr"])
size = save_adapter(model, MODELS / "m09-lab" / "adapter.pt", {"r": cfg["rank"], "alpha": 2 * cfg["rank"]})
print(f"3. LoRA r={cfg['rank']}: {count_trainable(model):,} trainable parameters, {info['seconds']:.0f}s, "
f"adapter {size / 1024:.0f} KB")
# 4. Held-out evaluation against the cheapest alternative, paired.
res = evaluate(model, tok, test)
rules_right = [keyword_label(t.text) == t.gold["category"] for t in test]
lora_right = [p == t.gold["category"] for p, t in zip(res["preds"], test)]
b, c, p = mcnemar_exact(lora_right, rules_right)
lo, hi = res["ci95"]
rlo, rhi = wilson(sum(rules_right), len(test))
print(f"4. LoRA {res['correct']}/24 (CI {lo:.0%} to {hi:.0%}) vs keyword rules {sum(rules_right)}/24 "
f"(CI {rlo:.0%} to {rhi:.0%}); LoRA-only wins {b}, rules-only wins {c}, McNemar p = {p:.3f}")
# 5. Regression checks: general language modelling and the old support behaviour.
val_text = corpus_validation_text(tok)
ppl_base, ppl_new = perplexity(base, tok, val_text), perplexity(model, tok, val_text)
recall = fact_recall(model, tok)
print(f"5. corpus perplexity {ppl_base:.2f} -> {ppl_new:.2f} ({ppl_new / ppl_base - 1:+.0%}); "
f"fact recall {recall}/12 (base 12/12)")
# 6. Decide by the rule written in step 1, and record why.
beats_baseline = b > c and p < 0.05
keeps_ability = ppl_new <= 1.10 * ppl_base
decision = {"ship": beats_baseline and keeps_ability, "beats_cheapest_baseline": beats_baseline,
"keeps_general_ability": keeps_ability, "lora_correct": res["correct"],
"rules_correct": sum(rules_right), "mcnemar_p": round(p, 4),
"ppl_change": round(ppl_new / ppl_base - 1, 3), "adapter": "models/m09-lab/adapter.pt"}
(MODELS / "m09-lab" / "decision.json").write_text(json.dumps(decision, indent=2) + "\n")
print("6. decision:", "SHIP" if decision["ship"] else "DO NOT SHIP; keep the keyword rules",
json.dumps({k: decision[k] for k in ("beats_cheapest_baseline", "keeps_general_ability")}))
Code explained
- In simple words: check the plan, build clean data, train, test against the cheapest alternative, check for damage, and let the pre-written rule decide.
- What happens:
- Loads
eval_plan.jsonand recomputes the test fingerprint. If anyone changed the test set since the plan was frozen, the script stops. - Scrubs and augments all 48 dev tickets, deduplicates, and asserts zero contamination with the test set.
- Trains a LoRA adapter with the settings in
models/m09-sweep.jsonand saves it tomodels/m09-lab/adapter.pt. - Scores the test set and runs the exact McNemar test against the keyword rules.
- Measures corpus perplexity and fact recall against the base model.
- Applies the ship rule from the plan (beat the rules with p < 0.05 and keep perplexity within 10%) and writes
models/m09-lab/decision.json. Run it withpython -m examples.m09_lab(about 30 seconds).
- Loads
- Comes out:
1. eval plan 5a95381c9e48536b ok: exact-match accuracy of the label with the highest log-probability (six candidates) on 24 test tickets
2. training rows: 191 (48 real), contamination with test: none
3. LoRA r=4: 32,768 trainable parameters, 22s, adapter 137 KB
4. LoRA 4/24 (CI 7% to 36%) vs keyword rules 15/24 (CI 43% to 79%); LoRA-only wins 1, rules-only wins 12, McNemar p = 0.003
5. corpus perplexity 1.65 -> 2.73 (+65%); fact recall 7/12 (base 12/12)
6. decision: DO NOT SHIP; keep the keyword rules {"beats_cheapest_baseline": false, "keeps_general_ability": false}
Notice step 2: 191 rows here against 189 in m09_train.py, because the lab deduplicates real and augmented rows in one pass. Two extra near-duplicate rows moved test accuracy from 6/24 to 4/24. That is what "within noise" means in practice: a change you would never think about moves the result by two tickets. The decision does not depend on it, because the gap to the rules (15/24, p = 0.003) is far larger than that noise.
Extend the lab yourself:
- Add the TF-IDF classifier as a second baseline in step 4, and make the ship rule require beating the best baseline.
- Add replay (half of each batch from the corpus, as in
m09_cpt.py) tosft_trainand see whether the perplexity check passes. - If you have an API key, run
m09_hosted_baseline.py, add the hosted model's accuracy to the plan, and ask whether any fine-tune is needed.
Project Milestone
After this module, your supportdesk copy contains:
data/m09/eval_plan.json: the frozen evaluation plan with the test fingerprint and ship rule.data/m09/*.jsonl: the scrubbed, augmented, deduplicated SFT data (and, with a key, distilled labels with provenance).examples/m09_setup.py,m09_lora.py: the prompt format, log-probability classifier, SFT loop, and LoRA (add, save, load, merge), which later modules can reuse by the same pattern.models/m09-lora-triage,models/m09-full,models/m09-reply-sft,models/m09-reply-dpo,models/m09-cpt,models/m09-lab: every trained artifact, each with its settings recorded, and the base model untouched.examples/m09_serve.py: an adapter server that runs the triage and reply adapters over one base.models/m09-lab/decision.json: the written ship decision. The assistant's triage step stays on the keyword rules; the reply-style adapter (DPO) is a candidate for drafts, pending the evaluation work in Module 10.tests/test_m09_fine_tuning.py: 14 fast tests of the mechanics.
The assistant itself did not change in production this module, and that is the correct outcome: we measured a lever and decided not to pull it yet.
Interview Questions
1. A product manager says "the model keeps giving wrong prices, let's fine-tune it on our pricing page". What do you say? Wrong facts are a retrieval problem, not a fine-tuning problem. Fine-tuning bakes in a snapshot: when prices change, the model is wrong until you retrain, retest, and redeploy, and it is still free to make up prices it never saw. Retrieval puts the current pricing page in the context, can cite it, and updates when someone edits the page. I would first check whether the retriever finds the pricing article for those questions (Module 7's hit rate), fix that, and add a grounding check. Fine-tuning becomes relevant if the problem turns out to be format or style.
2. How do you decide between prompting, retrieval, and fine-tuning for a new problem? By the symptom and by cost. Missing or changing knowledge points to retrieval. An unclear instruction points to prompting, which is also the cheapest thing to try first. A format problem points to structured outputs. A consistent style or a narrow high-volume task where a big model is too slow or expensive, with hundreds of good examples and an eval ready, points to fine-tuning. I also price the options: in this module, dropping a 48-shot prompt would save about 25 USD a month at Brightlane's volume, far less than the cost of maintaining a fine-tune.
3. You have 48 labeled tickets. Is that enough to fine-tune? It depends on whether you are steering an ability the model already has or teaching a new one. For a style TinyLM could already produce, 8 clean examples gave 56 of 72 house-style replies. For triage, which TinyLM could not do at all, 48 tickets produced 47/48 on the training tickets and 5/24 on test: memorization. With a real pretrained model, a few hundred examples per class is a common starting point for classification. What matters most is data quality, a held-out test set, and a learning curve (train on 25%, 50%, 100%) to see whether more data is still helping.
4. What is LoRA, and what are its knobs? LoRA freezes the base weights and learns a low-rank update B times A for selected linear layers, scaled by alpha / r. B starts at zero, so training starts from the base model's behavior. The knobs are the rank r (capacity), alpha (scale; keep alpha / r fixed when changing r), which layers get adapters, the learning rate (much higher than for full fine-tuning), and epochs. The practical wins are memory during training and tiny adapter files: 137 KB against 4.2 MB for TinyLM, and megabytes against gigabytes for real models. It does not save compute per step. In our run it was slower than full fine-tuning.
5. How do you know your fine-tuning data is not contaminated with your test set? Split first, then augment, dedupe, and generate only inside the training split. Then check mechanically: word n-gram overlap (any shared 8-gram) and near-duplicate similarity (3-gram Jaccard) between every training row and every test item. In this module, augmenting before splitting leaked 13 of 13 validation tickets, and the correct order leaked none. Also fingerprint the test set when you freeze the plan, so it cannot drift.
6. Explain DPO to a backend engineer. You have pairs of replies to the same prompt, one preferred. DPO trains the model so that, relative to a frozen copy of itself (the reference, usually the SFT model), the preferred reply becomes more likely and the rejected one less likely. beta controls how far it may move from the reference. There is no reward model and no reinforcement learning loop, just a classification-style loss on pairs. Measure it on held-out pairs the starting model gets wrong: our easy pairs were at 96% before training, while the house-vs-plain pairs moved from 0% to 100%.
7. What is reward hacking, and how do you guard against it? The model finds the cheapest way to raise the reward, which often is not what you meant. Our "count polite words" reward drove correct answers from 12/12 to 2/12 in 30 steps while the reward went up ninefold. Guards: prefer verifiable rewards that check the thing itself, add a KL penalty to limit drift (it slowed the hack, 6/12, but did not stop it), track metrics the reward does not include during training, read samples, and stop early.
8. What is RLVR, and why did DeepSeek use rule-based rewards for R1? Reinforcement learning with verifiable rewards scores outputs with a program (is the math answer right, do tests pass, is the format right) instead of a learned reward model. DeepSeek-R1-Zero used accuracy and format rewards with GRPO, which scores a group of samples per question relative to each other. The paper says it avoided neural reward models because they are susceptible to reward hacking at scale. It works where correctness can be checked, and does not apply to subjective qualities like tone.
9. After fine-tuning, how do you check you did not break the model? Measure general capability before and after on held-out data the fine-tune did not target: perplexity on general text and on domain text, plus task checks the old model passed (fact recall here). Our triage LoRA looked fine on its own task format but its corpus perplexity rose 89% and fact recall fell to 4/12. Then find out where the damage is (ours was spread across ordinary tokens), and reduce the total update, add replay, or broaden the data.
10. Does LoRA forget less than full fine-tuning? Often, and there is published evidence for it (Biderman et al., 2024). It is not guaranteed. In this module the chosen LoRA used a 30 times higher learning rate and more epochs, and it forgot much more than the full fine-tune. A gentler LoRA did not. Forgetting follows how far the weights move, so measure it for your actual settings.
11. How would you serve 20 customer-specific fine-tunes? One base model in memory, one LoRA adapter per customer, swapped per request or batched with a multi-LoRA server such as vLLM. Our swap took about 0.1 ms against 23 ms to load a full TinyLM copy; for a 7B model the gap is milliseconds against seconds and gigabytes. Adapters must match the base they were trained on, so a base upgrade means retraining and re-evaluating all 20.
12. Your fine-tune scored 6/24 and keyword rules 15/24. Your manager asks "but did the fine-tune work?". How do you answer? It learned the format (valid labels went from 0% to 100%) and memorized its training tickets, but it did not learn triage: on held-out tickets it loses to the rules 10 disagreements to 1 (exact McNemar p = 0.012), it is 1,000 times slower, and it damaged the model's general ability. By the rule we wrote before training, it does not ship. I would say that plainly, recommend the rules, and state what would change the answer (a stronger base, more labeled data, a hosted-model baseline).
Other Tools and Providers
| Tool or provider | What it is | When to use it instead of what we did |
|---|---|---|
| Hugging Face PEFT | Library of LoRA, QLoRA, and other adapter methods for Transformers models | Any real open-weight model; replaces our hand-written m09_lora.py |
| Hugging Face TRL | Trainers for SFT, DPO, GRPO, and reward modeling | The production version of our sft_train, dpo_loss, and GRPO loop |
| bitsandbytes | 8-bit and 4-bit (NF4) quantization used by QLoRA | Real QLoRA on a GPU; replaces our QuantLinear |
| Unsloth, Axolotl, LLaMA-Factory | Config-driven fine-tuning toolkits built on the libraries above | Faster setup and memory-optimized training runs |
| vLLM (multi-LoRA serving), LoRAX | Inference servers that serve many adapters over one base | Production version of AdapterServer |
Ollama (ADAPTER in a Modelfile) | Local serving of a base model plus a LoRA adapter | Running your adapter locally with the course's Ollama provider |
| Google Vertex AI tuning | Managed supervised and preference tuning for Gemini and open models | Tuning Gemini (the Gemini API itself has offered no tunable model since May 2025) |
| OpenAI fine-tuning API | Managed SFT, DPO, and reinforcement fine-tuning | Only for existing customers: being wound down, with no new jobs for active customers from 6 January 2027 |
| Together AI, Fireworks AI, and similar hosts | Managed fine-tuning and serving of open-weight models | Hosted LoRA training on open models without running GPUs yourself |
| scikit-learn | Classic classifiers (our TF-IDF baseline) | Always as the baseline; often as the answer |
Check each provider's current model list, prices, and terms before committing. This area changed more in 2025 and 2026 than any other part of the stack.
Coming Up in Module 10
This module's hardest problem was not training: it was that 24 test tickets could barely tell systems apart, and that "is this reply better?" had no automatic answer beyond a phrase match. Module 10 is about evaluation. You will build a real golden dataset for the Brightlane assistant, with sourced and curated hard cases, deterministic checks, LLM-as-judge with calibrated rubrics, human agreement, and CI gates. Then you will learn how to size an eval so its noise floor is below the difference you need to detect. Every decision in this module ("do not ship", "DPO is a candidate") gets its proper test there.