CourseLarge Language Models · Module 9: Adaptation: Fine-Tuning and Customization · part 48 of 80
Part 48 · Module 9: Adaptation: Fine-Tuning and Customization

Part 5: Preference and reasoning training

26 min read·22 Sept 2026

Preference data collection

SFT needs the right answer. For many things Maya cares about there is no single right answer, only a better and a worse one: the polite reply over the curt one, the correct policy over the confident mistake, the short draft over the rambling one. Preference data captures that as triples: a prompt, a chosen response, and a rejected response.

Where preference pairs come from in a support team:

  • Agent edits. The assistant drafts, the agent edits before sending. Draft = rejected, sent version = chosen. This is free and plentiful, but only if you log both versions.
  • Side-by-side review. Show a reviewer two drafts for the same ticket and ask which is better. Slower, but you control the comparison.
  • Rules plus review. Pair a known-good reply with a known-bad one (curt, or contradicting the KB), which is what we do below.

Keep the pairs honest: both responses must answer the same prompt, the chosen one must be genuinely better on the dimension you care about, and the reviewers should agree with each other. Measure agreement on a sample (Module 10 covers inter-annotator agreement). Pairs where reviewers disagree teach noise.

DPO and direct methods, practically

The classic way to use preferences was RLHF (next section): train a separate reward model on the pairs, then optimize the language model against it with reinforcement learning. DPO (Direct Preference Optimization, Rafailov et al., 2023) showed you can skip both steps and train directly on the pairs with a simple loss:

loss = -log sigmoid( beta * [ (log pi(chosen) - log ref(chosen)) - (log pi(rejected) - log ref(rejected)) ] )

In words: pi is the model being trained (the policy), ref is a frozen reference model (normally the SFT model you start from). For each pair, measure how much more the policy likes the chosen reply than the reference does, and the same for the rejected reply. The loss pushes the first up and the second down. The difference, times beta, is the implicit reward margin. beta sets how strongly the policy is held near the reference: small beta lets it move further. Because everything is measured relative to the reference, the model is not rewarded for changing things the pairs do not talk about.

The recipe below is the standard one: SFT first (on the mixed-quality logs, as a LoRA adapter; this model is also the reference), then DPO on pairs, with two measurements: the win rate (how often the model gives the chosen reply a higher log-probability than the rejected one) and the styles of replies it actually samples.

examples/m09_dpo.py

python
"""Preference training for reply style: SFT on mixed-quality logs, then DPO on chosen/rejected pairs.

Reference model = the SFT model. Everything is a LoRA adapter over the same frozen TinyLM base.
"""
import random
from collections import Counter

import torch
import torch.nn.functional as F

from examples.m09_lora import add_lora, load_adapter, save_adapter
from examples.m09_setup import MODELS, encode_example, make_batch, seed_everything, sft_train
from supportdesk.tinylm import SamplingParams, generate, load

# question, correct answer, a phrase that proves the policy fact is right, a confident wrong answer
FACTS = [
    ("I was charged twice for the {plan} plan this month. Can you refund the duplicate?",
     "Duplicate charges are always refunded in full within 5 to 10 business days to the original payment method.",
     "5 to 10 business days", "Duplicate charges are kept as account credit and are not refunded to your card."),
    ("how do I cancel my {plan} subscription?",
     "You can cancel at any time from Settings > Billing > Cancel plan.", "Cancel plan",
     "Please email us to cancel. We need 30 days notice before any plan can end."),
    ("I forgot my password and the reset link expired.",
     "Reset links expire after 30 minutes and can be used once. Please request a new link from the sign-in page.",
     "30 minutes", "Reset links never expire, so just use the old link again."),
    ("my account is locked after too many attempts.",
     "After 5 failed attempts an account is locked for 15 minutes. Please wait and try again.", "15 minutes",
     "Send us your password and we will unlock the account right away."),
    ("does the {plan} plan include SSO?",
     "SSO is available on Business and Enterprise plans.", "Business and Enterprise",
     "SSO is included in every plan, including Free."),
    ("our automations stopped and there is a banner about a limit.",
     "Automations pause when the monthly run limit is reached. They resume on the first day of next month.",
     "first day of next month", "Automation limits never reset, so you must buy a new plan."),
    ("how do I export my {area} data?",
     "Export a board to CSV or JSON from the board menu > Export.", "CSV or JSON",
     "Exports are only available on the Enterprise plan."),
    ("Slack notifications stopped posting.",
     "The Slack token was usually revoked. Please reconnect the integration from Settings > Integrations > Slack.",
     "reconnect", "The Slack integration was removed last year."),
    ("can I get a refund on my annual {plan} plan?",
     "Annual plans cancelled within 14 days of purchase or renewal get a full refund.", "14 days",
     "Annual plans can be refunded at any time for any reason."),
    ("where can I find my invoices?",
     "Invoices are emailed to the billing contact and are available under Settings > Billing > Invoices.",
     "Billing > Invoices", "We do not issue invoices."),
    ("how much does the {plan} plan cost?",
     "Team costs 12 USD per user per month. Business costs 24 USD per user per month.", "12 USD",
     "Team costs 5 USD per user per month."),
    ("the {area} is not loading. Is there an outage?",
     "Please check status.brightlane.example for live updates.", "status.brightlane.example",
     "There is never an outage, the problem is on your side."),
]
HOUSE = "Thank you. "  # every token here is common in the pretraining corpus (see Part 4)
CURT = ["Read the help center.", "Not our problem.", "Figure it out yourself."]
NAMES = ["Ana", "Ben", "Chen", "Dara", "Eli", "Hana", "Jo", "Kofi", "Priya"]
PLANS, AREAS = ["Free", "Team", "Business"], ["board", "export", "mobile app"]
TRAIN_FACTS, HELD_OUT_FACTS = list(range(9)), [9, 10, 11]  # DPO pairs never use the last three questions
rng = random.Random(3)


def prompt(i: int) -> str:
    q = FACTS[i][0].format(plan=rng.choice(PLANS), area=rng.choice(AREAS))
    return f"Customer ({rng.choice(NAMES)}): {rng.choice(['Hi', 'Hello'])}, {q}\nAgent ({rng.choice(NAMES)}):"


def fixed_prompt(i: int) -> str:
    """One fixed phrasing per question, for side-by-side greedy outputs."""
    q = FACTS[i][0].format(plan="Team", area="board")
    return f"Customer (Ana): Hi, {q}\nAgent (Ben):"


def style_of(reply: str, i: int) -> str:
    """Label a reply: house (polite and correct), plain (correct), wrong, curt, or other."""
    correct = FACTS[i][2].lower() in reply.lower()
    if any(c.lower().rstrip(".") in reply.lower() for c in CURT):
        return "curt"
    if FACTS[i][3][:25].lower() in reply.lower():
        return "wrong"
    if correct:
        return "house" if reply.strip().startswith(HOUSE.strip()) else "plain"
    return "other"


def seq_logprob(model, tok, pairs):
    """Sum of log-probabilities of each reply given its prompt (one number per pair)."""
    x, y, m = make_batch([encode_example(tok, p, a) for p, a in pairs])
    logp = torch.log_softmax(model(x), dim=-1).gather(-1, y.unsqueeze(-1)).squeeze(-1)
    return (logp * m).sum(dim=1)


def dpo_loss(policy, ref, tok, batch, beta, nll_weight=0.0):
    """DPO: -log sigmoid(beta * [(log pi(c) - log ref(c)) - (log pi(r) - log ref(r))]).

    nll_weight > 0 adds the ordinary SFT loss on the chosen reply (per token), which keeps the
    chosen text likely while the margin grows.
    """
    chosen = [(p, c) for p, c, _ in batch]
    rejected = [(p, r) for p, _, r in batch]
    pc, pr = seq_logprob(policy, tok, chosen), seq_logprob(policy, tok, rejected)
    with torch.no_grad():
        rc, rr = seq_logprob(ref, tok, chosen), seq_logprob(ref, tok, rejected)
    margin = beta * ((pc - rc) - (pr - rr))
    loss = -F.logsigmoid(margin).mean()
    if nll_weight:
        n_tokens = sum(len(tok.encode(c).ids) for _, c in chosen)
        loss = loss + nll_weight * (-pc.sum() / n_tokens)
    return loss, margin.detach()


@torch.no_grad()
def preference_stats(model, tok, pairs):
    """Win rate: how often the model gives the chosen reply a higher log-probability than the rejected one."""
    diff = seq_logprob(model, tok, [(p, c) for p, c, _ in pairs]) - seq_logprob(model, tok, [(p, r) for p, _, r in pairs])
    return (diff > 0).float().mean().item(), diff.mean().item()


def sample_styles(model, tok, facts, n=10):
    styles = Counter()
    for i in facts:
        for s in range(n):
            g = generate(model, tok, prompt(i), SamplingParams(max_new_tokens=40, temperature=0.8, seed=s, stop=["\n"]))
            styles[style_of(g.text, i)] += 1
    return styles


def build_logs(per_fact=10):
    """Support 'logs': the same questions answered in mixed quality, as real agents do."""
    logs = []
    for i, (_, answer, _, wrong) in enumerate(FACTS):
        for _ in range(per_fact):
            style = rng.choices(["house", "plain", "curt", "wrong"], weights=[3, 4, 2, 1])[0]
            reply = {"house": HOUSE + answer, "plain": answer, "curt": rng.choice(CURT), "wrong": wrong}[style]
            logs.append((prompt(i), " " + reply + "\n", style))
    return logs


def build_pairs(facts, per_fact=4):
    """Chosen = polite and correct. Rejected = curt, or confidently wrong about policy."""
    pairs = []
    for i in facts:
        _, answer, _, wrong = FACTS[i]
        for _ in range(per_fact):
            p = prompt(i)
            pairs.append((p, " " + HOUSE + answer + "\n", " " + rng.choice(CURT) + "\n"))
            pairs.append((p, " " + HOUSE + answer + "\n", " " + wrong + "\n"))
    return pairs


def build_style_pairs(facts, per_fact=4):
    """A harder check: house style vs the SAME correct answer without it (never used for training)."""
    return [(p, " " + HOUSE + FACTS[i][1] + "\n", " " + FACTS[i][1] + "\n")
            for i in facts for p in [prompt(i) for _ in range(per_fact)]]


def report(name, model, tok, train_pairs, test_pairs):
    win_tr, m_tr = preference_stats(model, tok, train_pairs)
    win_te, m_te = preference_stats(model, tok, test_pairs)
    win_st, m_st = preference_stats(model, tok, STYLE_PAIRS)
    with torch.no_grad():
        chosen_lp = seq_logprob(model, tok, [(p, c) for p, c, _ in test_pairs]).mean().item()
    held = sample_styles(model, tok, HELD_OUT_FACTS)
    print(f"{name:<8} win rate train {win_tr:.0%}, held-out {win_te:.0%} (mean margin {m_te:+.1f} nats); "
          f"house-vs-plain held-out {win_st:.0%} (margin {m_st:+.2f}); held-out log p(chosen) {chosen_lp:.1f}")
    print(f"         30 sampled held-out replies (T=0.8): {dict(sorted(held.items()))}")


if __name__ == "__main__":
    seed_everything(0)
    base, tok = load()
    logs = build_logs()
    print("log styles:", dict(Counter(s for _, _, s in logs)), "rows:", len(logs))

    # Step 1: SFT on the logs as a LoRA adapter. This model is also the DPO reference.
    sft, _ = load()
    add_lora(sft, r=8, alpha=16)
    info = sft_train(sft, tok, [(p, a) for p, a, _ in logs], epochs=4, lr=3e-3)
    save_adapter(sft, MODELS / "m09-reply-sft" / "adapter.pt", {"r": 8, "alpha": 16, "task": "reply-sft"})
    print(f"SFT: {info['steps']} steps in {info['seconds']:.0f}s")

    train_pairs, test_pairs = build_pairs(TRAIN_FACTS), build_pairs(HELD_OUT_FACTS)
    STYLE_PAIRS = build_style_pairs(HELD_OUT_FACTS)
    print(f"preference pairs: train {len(train_pairs)} (questions 1-9), held-out {len(test_pairs)} (questions 10-12)")
    report("base", base, tok, train_pairs, test_pairs)
    report("SFT", sft, tok, train_pairs, test_pairs)
    for i in HELD_OUT_FACTS:
        g = generate(sft, tok, fixed_prompt(i), SamplingParams(max_new_tokens=40, temperature=0, stop=["\n"]))
        print(f"    greedy, question {i + 1}: {g.text.strip()!r}")

    # Step 2: DPO. The policy starts as a copy of the SFT model; the reference (SFT) stays frozen.
    for p in sft.parameters():
        p.requires_grad = False
    for name, lr, epochs, nll in (("DPO-hot", 1e-3, 4, 0.0), ("DPO+NLL", 3e-4, 3, 1.0)):
        seed_everything(0)
        policy, _ = load()
        add_lora(policy, r=8, alpha=16)
        load_adapter(policy, MODELS / "m09-reply-sft" / "adapter.pt")
        params = [p for p in policy.parameters() if p.requires_grad]
        opt = torch.optim.AdamW(params, lr=lr, weight_decay=0.0)
        beta, order = 0.1, list(range(len(train_pairs)))
        for epoch in range(epochs):
            random.Random(epoch).shuffle(order)
            policy.train()
            for k in range(0, len(order), 8):
                loss, margin = dpo_loss(policy, sft, tok, [train_pairs[j] for j in order[k:k + 8]], beta, nll)
                opt.zero_grad()
                loss.backward()
                torch.nn.utils.clip_grad_norm_(params, 1.0)
                opt.step()
            policy.eval()
        print(f"{name}: lr {lr:g}, {epochs} epochs, beta {beta}, nll weight {nll}, "
              f"last batch implicit reward margin {margin.mean():+.2f}")
        report(name, policy, tok, train_pairs, test_pairs)
        for i in HELD_OUT_FACTS:
            g = generate(policy, tok, fixed_prompt(i), SamplingParams(max_new_tokens=40, temperature=0, stop=["\n"]))
            print(f"    greedy, question {i + 1}: {g.text.strip()!r}")
        if nll:
            save_adapter(policy, MODELS / "m09-reply-dpo" / "adapter.pt",
                         {"r": 8, "alpha": 16, "task": "reply-dpo", "beta": beta, "nll_weight": nll})

Code explained

  • In simple words: teach TinyLM Brightlane's house style in two steps, imitation first and then preferences, and measure what each step changed.
  • What happens:
    • FACTS: 12 common support questions, each with the correct answer (taken from the KB), a phrase that proves the answer is right, and a confident wrong answer. prompt renders a question in TinyLM's corpus format (Customer (Ana): Hi, ...\nAgent (Ben):) with random names and plans. Questions 10 to 12 are never used for preference training.
    • style_of: labels a reply as house, plain, curt, wrong, or other.
    • seq_logprob: the total log-probability of each reply given its prompt, using the same masked batching as SFT.
    • dpo_loss: the formula above. The reference log-probabilities are computed without gradients. nll_weight optionally adds the ordinary SFT loss on the chosen reply, a common stabilizer that keeps the chosen text likely while the margin grows.
    • preference_stats: win rate and mean margin in nats (a nat is a unit of log-probability; +1 nat means about 2.7 times more likely).
    • build_logs: 120 mixed-quality log rows (house, plain, curt, wrong in proportions 3:4:2:1). build_pairs: for questions 1 to 9, the chosen reply is house style and correct, and the rejected reply is either curt or wrong. That gives 72 training pairs, and 24 held-out pairs from questions 10 to 12. build_style_pairs: a harder held-out check, house style against the same correct answer without "Thank you.". Nothing like it is in the training pairs.
    • Main: SFT a rank-8 LoRA on the logs (saved as models/m09-reply-sft), then run DPO twice from that adapter with the SFT model frozen as reference: "DPO-hot" (lr 1e-3, 4 epochs, no NLL) and "DPO+NLL" (lr 3e-4, 3 epochs, NLL weight 1.0, saved as models/m09-reply-dpo). beta is 0.1 in both. Run it with python -m examples.m09_dpo (about 20 seconds).
  • Comes out:
text
log styles: {'house': 28, 'wrong': 13, 'plain': 53, 'curt': 26} rows: 120
SFT: 32 steps in 3s
preference pairs: train 72 (questions 1-9), held-out 24 (questions 10-12)
base     win rate train 100%, held-out 96% (mean margin +65.9 nats); house-vs-plain held-out 0% (margin -14.51); held-out log p(chosen) -29.0
         30 sampled held-out replies (T=0.8): {'plain': 30}
SFT      win rate train 100%, held-out 100% (mean margin +56.4 nats); house-vs-plain held-out 0% (margin -2.48); held-out log p(chosen) -3.3
         30 sampled held-out replies (T=0.8): {'house': 8, 'other': 1, 'plain': 21}
    greedy, question 10: 'Invoices are emailed to the billing contact and are available under Settings > Billing > Invoices.'
    greedy, question 11: 'Team costs 12 USD per user per month. Business costs 24 USD per user per month.'
    greedy, question 12: 'Please check status.brightlane.example for live updates.'
DPO-hot: lr 0.001, 4 epochs, beta 0.1, nll weight 0.0, last batch implicit reward margin +7.01
DPO-hot  win rate train 100%, held-out 100% (mean margin +118.1 nats); house-vs-plain held-out 100% (margin +4.13); held-out log p(chosen) -1.9
         30 sampled held-out replies (T=0.8): {'house': 15, 'other': 4, 'plain': 11}
    greedy, question 10: 'Thank you. Invoices are emailed to the billing contact and are available under Settings > Billing > Invoices.'
    greedy, question 11: 'Thank you. Team costs 12 USD per user per month. Business costs 24 USD per user per month.'
    greedy, question 12: 'Thank you. Please check status.brightlane.example for live updates.'
DPO+NLL: lr 0.0003, 3 epochs, beta 0.1, nll weight 1.0, last batch implicit reward margin +5.94
DPO+NLL  win rate train 100%, held-out 100% (mean margin +96.8 nats); house-vs-plain held-out 100% (margin +2.08); held-out log p(chosen) -1.3
         30 sampled held-out replies (T=0.8): {'house': 17, 'other': 1, 'plain': 12}
    greedy, question 10: 'Thank you. Invoices are emailed to the billing contact and are available under Settings > Billing > Invoices.'
    greedy, question 11: 'Thank you. Team costs 12 USD per user per month. Business costs 24 USD per user per month.'
    greedy, question 12: 'Thank you. Please check status.brightlane.example for live updates.'
  • The easy win rate was saturated from the start. Even the base model prefers correct-and-polite over curt-or-wrong on 96% of held-out pairs, because it memorized the correct answers in pretraining. If this were your only preference metric you would conclude DPO did nothing. Always include pairs the starting model gets wrong.
  • The house-vs-plain win rate is the real signal: 0% for the base and SFT models, 100% after either DPO run, on questions DPO never trained on. The mean margin moves from -2.48 nats (SFT prefers the plain answer) to +2.08 or more (DPO prefers the house style).
  • Greedy outputs change exactly as intended: "Thank you." is added, and the facts are untouched on all three held-out questions.
  • Sampled outputs move less than the win rate suggests. At temperature 0.8, house style goes from 8 of 30 (SFT) to 15 (DPO-hot) and 17 (DPO+NLL). A pairwise win rate says the model ranks chosen above rejected; it does not say how often the model generates the chosen style. Report both.
  • DPO-hot drifts more. A larger learning rate with no NLL term inflates the margin further (+118 nats) and produces more "other" replies (4 of 30) than DPO+NLL (1 of 30). A bigger margin is not a better model. The NLL term and a lower learning rate kept the output distribution closer to fluent replies, which is why we saved that one.

Back to data: logs with review, and how few examples is enough

Part 3 promised a measurement of what unreviewed logs do. build_logs in m09_dpo.py above makes 120 simulated log rows for 12 common support questions, in four styles with realistic proportions: the house style ("Thank you." and the correct answer), a plain correct answer, a curt reply ("Not our problem."), and a confidently wrong policy answer. The next script trains on those logs with and without review, and then asks how few good examples are enough.

examples/m09_log_review.py

python
"""SFT data from support logs: automatic checks, a human review sample, and how few examples are enough.

The 'logs' are the mixed-quality agent replies from m09_dpo.build_logs (house style, plain, curt, wrong).
Training uses questions 1-9 only; questions 10-12 are held out to see whether the style generalizes.
"""
import math
import random
from collections import Counter

from examples.m09_dpo import FACTS, HELD_OUT_FACTS, TRAIN_FACTS, build_logs, sample_styles
from examples.m09_lora import add_lora
from examples.m09_setup import seed_everything, sft_train
from supportdesk.tinylm import load

PER_FACT = 10
STEPS = 40


def automatic_checks(reply: str, i: int) -> list[str]:
    """Problems a script can find without a person: too short, or missing the policy fact from the KB."""
    problems = []
    if len(reply.split()) < 6:
        problems.append("too short")
    if FACTS[i][2].lower() not in reply.lower():
        problems.append("policy fact missing")
    return problems


def train_and_measure(rows, seed=0):
    seed_everything(seed)
    model, tok = load()
    add_lora(model, r=8, alpha=16)
    batches_per_epoch = math.ceil(len(rows) / 16)
    epochs = max(1, round(STEPS / batches_per_epoch))  # same number of updates whatever the data size
    sft_train(model, tok, [(p, a) for p, a, _ in rows], epochs=epochs, lr=3e-3, seed=seed)
    seen = sample_styles(model, tok, TRAIN_FACTS, n=4)     # new phrasings of questions 1-9: 36 replies
    unseen = sample_styles(model, tok, HELD_OUT_FACTS)     # held-out questions 10-12: 30 replies
    return seen, unseen


if __name__ == "__main__":
    seed_everything(0)
    logs = [(p, a, s, n // PER_FACT) for n, (p, a, s) in enumerate(build_logs(PER_FACT))]
    logs = [r for r in logs if r[3] in TRAIN_FACTS]
    print(f"log rows for questions 1-9: {len(logs)}, styles {dict(Counter(r[2] for r in logs))}")

    kept, dropped = [], Counter()
    for p, a, s, i in logs:
        problems = automatic_checks(a, i)
        if problems:
            dropped[(s, problems[0])] += 1
        else:
            kept.append((p, a, s))
    print(f"automatic checks keep {len(kept)}; dropped {dict(dropped)}")
    for p, a, s in random.Random(1).sample(kept, 3):
        print(f"  review sample [{s}]: {p.splitlines()[-2][:48]!r} -> {a.strip()[:50]!r}")

    house = [(p, a, s) for p, a, s in kept if s == "house"]
    print("styles of sampled replies (T=0.8) after SFT on:")
    for name, rows in (("all logs, unreviewed", [(p, a, s) for p, a, s, _ in logs]),
                       ("checked logs", kept), ("house style only", house)):
        seen, unseen = train_and_measure(rows)
        print(f"  {name:<21} ({len(rows):>2} rows) questions 1-9: {dict(sorted(seen.items()))}")
        print(f"  {'':<31} questions 10-12: {dict(sorted(unseen.items()))}")

    print("how few is enough? house-style rows only, k examples:")
    for k in (2, 4, 8, len(house)):
        styles = Counter()
        for seed in (0, 1):  # two seeds: small k is noisy
            styles += train_and_measure(random.Random(seed).sample(house, k), seed)[0]
        print(f"  k={k:>2}: questions 1-9, 72 replies: house {styles['house']:>2}, plain {styles['plain']:>2}, "
              f"other {styles['other']:>2}, wrong {styles['wrong']}, curt {styles['curt']}")

Code explained

  • In simple words: filter the logs with automatic checks, show a reviewer a sample, then compare what the model learns from unreviewed logs, checked logs, and house-style rows only.
  • What happens:
    • Only questions 1 to 9 are used for training; 10 to 12 are held out to see whether the style generalizes to questions the model was not trained on.
    • automatic_checks: two checks a script can do. "Too short" (under 6 words) catches curt replies. "Policy fact missing" checks that the reply contains the phrase from the KB that makes it correct (for example "5 to 10 business days" for refunds), which catches wrong answers. Anything flagged is dropped. A random sample of the kept rows is printed for a human reviewer. In production, that reviewer also checks for personal data and for answers that are correct but off-policy.
    • train_and_measure: trains a rank-8 LoRA adapter for the same number of updates (STEPS = 40) whatever the data size, so small datasets are not also under-trained. Then it samples replies at temperature 0.8 and labels each one's style: 36 replies to new phrasings of questions 1 to 9, and 30 to the held-out questions.
    • The last loop trains on only k house-style rows (k = 2, 4, 8, 16), twice with different seeds, and counts house-style replies out of 72. Run it with python -m examples.m09_log_review (about 45 seconds).
  • Comes out:
text
log rows for questions 1-9: 90, styles {'house': 16, 'wrong': 11, 'plain': 40, 'curt': 23}
automatic checks keep 56; dropped {('wrong', 'policy fact missing'): 11, ('curt', 'too short'): 23}
  review sample [house]: 'Customer (Dara): Hello, how do I cancel my Team ' -> 'Thank you. You can cancel at any time from Setting'
  review sample [house]: 'Customer (Priya): Hello, how do I export my boar' -> 'Thank you. Export a board to CSV or JSON from the '
  review sample [plain]: 'Customer (Chen): Hi, can I get a refund on my an' -> 'Annual plans cancelled within 14 days of purchase '
styles of sampled replies (T=0.8) after SFT on:
  all logs, unreviewed  (90 rows) questions 1-9: {'other': 6, 'plain': 30}
                                  questions 10-12: {'house': 4, 'other': 12, 'plain': 14}
  checked logs          (56 rows) questions 1-9: {'house': 12, 'plain': 24}
                                  questions 10-12: {'house': 3, 'other': 12, 'plain': 15}
  house style only      (16 rows) questions 1-9: {'house': 33, 'other': 3}
                                  questions 10-12: {'house': 19, 'other': 10, 'plain': 1}
how few is enough? house-style rows only, k examples:
  k= 2: questions 1-9, 72 replies: house 21, plain  6, other 45, wrong 0, curt 0
  k= 4: questions 1-9, 72 replies: house 35, plain  1, other 36, wrong 0, curt 0
  k= 8: questions 1-9, 72 replies: house 56, plain  1, other 15, wrong 0, curt 0
  k=16: questions 1-9, 72 replies: house 66, plain  0, other  6, wrong 0, curt 0
  • The checks dropped all 23 curt and all 11 wrong replies and kept 56. The reviewer sample looks right.
  • Unreviewed logs: 30 of 36 replies are plain and none use the house style. The model learns the majority behavior in the data, and in these logs the house style is only 16 of 90 rows. Checked logs raise the house share to 16 of 56 rows, and 12 of 36 replies come out in house style. House-style rows only: 33 of 36.
  • An honest note: the unreviewed model produced no curt or wrong replies in these samples, because TinyLM's memorized plain answers dominate 23 curt rows. Do not read that as "bad rows are harmless". What you train on becomes what comes out, in proportion. With a real model and a larger share of bad rows, curt and wrong replies do come back.
  • Held-out questions 10 to 12 move less (19 of 30 house-style at best, with many "other" replies that mix fragments of several answers). A style learned on 9 questions transfers only partly to new ones in a model this small.
  • How few is enough: 2 examples give 21 of 72 house-style replies, 4 give 35, 8 give 56, and 16 give 66. For a style the base model can already produce (every token of "Thank you." is common in its corpus), a handful of clean examples does most of the work and quality beats quantity: 16 clean rows beat 90 unreviewed ones on the thing we wanted. The LIMA paper made the same point at scale, tuning a 65B model on 1,000 carefully chosen examples (Zhou et al., 2023). For a task the base model cannot do (triage, Part 4), 48 examples turned out to be far from enough.

RLHF and reinforcement learning with verifiable rewards, conceptually

RLHF (reinforcement learning from human feedback) is the three-step recipe behind the first instruction-following chat models (Ouyang et al., 2022, "Training language models to follow instructions with human feedback"):

  1. SFT on human-written demonstrations.
  2. Train a reward model: a network that reads a (prompt, response) pair and outputs a score, fitted to human preference pairs.
  3. Optimize the SFT model with reinforcement learning (PPO in that paper) to get high reward-model scores, with a KL penalty that charges the policy for drifting away from the SFT model.

In that paper, outputs from a 1.3B InstructGPT model were preferred by labelers over outputs from the 175B GPT-3. DPO later showed that for many uses, steps 2 and 3 can be folded into one loss.

RLVR (reinforcement learning with verifiable rewards) replaces the learned reward model with a program that checks the answer: does the math answer match, do the unit tests pass, does the JSON validate. DeepSeek used it at scale. Its DeepSeek-R1 paper (arXiv 2501.12948, published in Nature in 2025) trained DeepSeek-R1-Zero with pure reinforcement learning from a base model, with no supervised reasoning data. The rewards were rule-based: an accuracy reward (is the final answer right) and a format reward (is the reasoning inside the required tags). The authors explain that they did not use neural reward models because they "are susceptible to reward hacking during large-scale reinforcement learning". In the current version of the paper, pass@1 on the AIME 2024 math competition rose from 15.6% to 77.9% during training. The algorithm was GRPO (Group Relative Policy Optimization, introduced in DeepSeekMath, arXiv 2402.03300). It samples a group of answers per question (16 in R1), scores each one, and uses each answer's reward relative to the group average as its training signal. That removes the separate "critic" network PPO needs.

For Brightlane, RLVR is attractive exactly where a check exists: "the reply quotes the correct refund window from the KB", "the triage JSON validates and the category matches the agent's final category". It is not available for "the reply is warm and clear", which is where preference data and DPO fit.

SituationUse thisWhy
You can write the right answer for each inputSFTSimplest, most stable; learn by imitation
You can say which of two answers is better, not what the best one isDPO (or similar direct methods) on preference pairsNo reward model, no RL loop; works on pairs from agent edits
A program can check correctness (tests, schemas, exact answers)RLVR with a GRPO-style loopThe check cannot be flattered; reward hacking is much harder
Quality is subjective and you need to optimize it at scaleRLHF with a reward model, a KL penalty, and constant auditingMost flexible and most fragile; the reward model is itself a model that can be gamed

Reward hacking as a predictable outcome

Reward hacking is what happens when a model is optimized against a reward that only approximates what you want: it finds the cheapest way to raise the number, which is usually not the thing you wanted. It is not a rare accident. It is what optimization does. The script below optimizes the reply SFT adapter three ways with a bare-bones GRPO-style loop: sample 8 replies per prompt, score them, and push up the ones that scored above the group average.

examples/m09_reward_hacking.py

python
"""Reward hacking, measured: optimize TinyLM against a flawed reward and against a verifiable one.

The optimizer is a bare-bones GRPO-style policy gradient: sample a group of replies per prompt,
score them, and push up the log-probability of replies that scored above the group average.
"""
import random
import re

import torch

from examples.m09_dpo import FACTS, TRAIN_FACTS, fixed_prompt, prompt
from examples.m09_lora import with_adapter
from examples.m09_setup import seed_everything
from supportdesk.tinylm import SamplingParams, generate

POLITE = re.compile(r"\b(thank|thanks|sorry|please|happy|glad|welcome)\b", re.IGNORECASE)


def polite_reward(reply: str, i: int) -> float:
    """The flawed reward someone might write for 'be polite': count polite words."""
    return float(len(POLITE.findall(reply)))


def verified_reward(reply: str, i: int) -> float:
    """A verifiable reward: 1 if the reply states the correct policy fact, else 0."""
    return float(FACTS[i][2].lower() in reply.lower())


def reply_logprob(model, prompt_ids, reply_ids):
    x = torch.tensor([prompt_ids + reply_ids[:-1]])
    logp = torch.log_softmax(model(x)[0, len(prompt_ids) - 1:], dim=-1)
    return logp.gather(1, torch.tensor(reply_ids).unsqueeze(1)).sum()


def train(reward_fn, steps=30, group=8, lr=1e-3, kl_weight=0.0, seed=0):
    seed_everything(seed)
    policy, ref = with_adapter("m09-reply-sft"), with_adapter("m09-reply-sft")
    tok = load_tokenizer()
    params = [p for p in policy.parameters() if p.requires_grad]
    opt = torch.optim.AdamW(params, lr=lr, weight_decay=0.0)
    rng = random.Random(seed)
    history = []
    for step in range(1, steps + 1):
        i = rng.choice(TRAIN_FACTS)
        p = prompt(i)
        p_ids = tok.encode(p).ids
        samples = [generate(policy, tok, p, SamplingParams(max_new_tokens=30, temperature=1.0, seed=step * 100 + k,
                                                           stop=["\n"])) for k in range(group)]
        rewards = torch.tensor([reward_fn(s.text, i) for s in samples])
        adv = (rewards - rewards.mean()) / (rewards.std() + 1e-6)  # group-relative advantage
        policy.train()
        loss = 0.0
        for s, a in zip(samples, adv):
            if not s.token_ids:
                continue
            lp = reply_logprob(policy, p_ids, s.token_ids)
            loss = loss - a * lp / len(s.token_ids)
            if kl_weight:
                with torch.no_grad():
                    ref_lp = reply_logprob(ref, p_ids, s.token_ids)
                loss = loss + kl_weight * (lp - ref_lp) / len(s.token_ids)
        opt.zero_grad()
        (loss / group).backward()
        torch.nn.utils.clip_grad_norm_(params, 1.0)
        opt.step()
        policy.eval()
        history.append(rewards.mean().item())
        if step % 10 == 0:
            report(step, policy, tok, history)
    return policy


def report(step, policy, tok, history):
    replies = [generate(policy, tok, fixed_prompt(i), SamplingParams(max_new_tokens=30, temperature=0, stop=["\n"])).text
               for i in range(len(FACTS))]
    polite = sum(polite_reward(r, i) for i, r in enumerate(replies)) / len(replies)
    correct = sum(verified_reward(r, i) for i, r in enumerate(replies))
    trained = f"{sum(history[-10:]) / 10:4.2f}" if history else " n/a"
    print(f"  step {step:>2}: train reward (last 10) {trained} | greedy on 12 questions: "
          f"polite words/reply {polite:4.2f}, correct fact {correct:>2.0f}/12 | e.g. {replies[10].strip()[:60]!r}")


def load_tokenizer():
    from supportdesk.tinylm import load
    return load()[1]


if __name__ == "__main__":
    tok = load_tokenizer()
    print("starting point (reply SFT adapter):")
    report(0, with_adapter("m09-reply-sft"), tok, [])
    print("A) reward = number of polite words")
    train(polite_reward)
    print("B) same flawed reward, plus a KL penalty that keeps the policy near the SFT model")
    train(polite_reward, kl_weight=1.0)
    print("C) verifiable reward = states the correct policy fact")
    train(verified_reward)

Code explained

  • In simple words: tell the model "be polite" with a lazy reward (count polite words), then with the same reward plus a leash, then with a reward that checks the facts, and watch what each one produces.
  • What happens:
    • polite_reward: the number of words like "thank", "please", "sorry" in the reply. Someone might really write this to encourage politeness.
    • verified_reward: 1 if the reply contains the verified policy phrase for that question, else 0. This is a verifiable reward.
    • train: for 30 steps, pick a training question, sample 8 replies at temperature 1.0, compute rewards, standardize them within the group (the GRPO advantage), and take a gradient step that raises the log-probability of above-average replies and lowers the rest. With kl_weight it also charges the policy for moving its log-probabilities away from the frozen SFT reference, which is RLHF's KL penalty in miniature.
    • report: every 10 steps, the average training reward and greedy replies to all 12 questions: polite words per reply, and how many still state the correct fact. Run it with python -m examples.m09_reward_hacking (about a minute).
  • Comes out:
text
starting point (reply SFT adapter):
  step  0: train reward (last 10)  n/a | greedy on 12 questions: polite words/reply 0.33, correct fact 12/12 | e.g. 'Team costs 12 USD per user per month. Business costs 24 USD '
A) reward = number of polite words
  step 10: train reward (last 10) 0.41 | greedy on 12 questions: polite words/reply 1.08, correct fact 12/12 | e.g. 'Thank you. Team costs 12 USD per user per month. Business co'
  step 20: train reward (last 10) 1.11 | greedy on 12 questions: polite words/reply 1.50, correct fact 10/12 | e.g. 'Thank you. Team costs 12 USD per user per month. Business co'
  step 30: train reward (last 10) 3.61 | greedy on 12 questions: polite words/reply 12.50, correct fact  2/12 | e.g. 'Thank Thank Thank Thank Thank Thank Thank Thank Thank Thank '
B) same flawed reward, plus a KL penalty that keeps the policy near the SFT model
  step 10: train reward (last 10) 0.35 | greedy on 12 questions: polite words/reply 0.75, correct fact 12/12 | e.g. 'Team costs 12 USD per user per month. Business costs 24 USD '
  step 20: train reward (last 10) 0.47 | greedy on 12 questions: polite words/reply 1.25, correct fact 10/12 | e.g. 'Team costs 12 USD per user per month. Business costs 24 USD '
  step 30: train reward (last 10) 1.10 | greedy on 12 questions: polite words/reply 2.83, correct fact  6/12 | e.g. 'Thank you. Team costs 12 USD per user per user per month. Bu'
C) verifiable reward = states the correct policy fact
  step 10: train reward (last 10) 0.62 | greedy on 12 questions: polite words/reply 0.33, correct fact 12/12 | e.g. 'Team costs 12 USD per user per month. Business costs 24 USD '
  step 20: train reward (last 10) 0.85 | greedy on 12 questions: polite words/reply 0.33, correct fact 12/12 | e.g. 'Team costs 12 USD per user per month. Business costs 24 USD '
  step 30: train reward (last 10) 0.89 | greedy on 12 questions: polite words/reply 0.33, correct fact 12/12 | e.g. 'Team costs 12 USD per user per month. Business costs 24 USD '
  • A) The flawed reward is hacked in 30 steps. Training reward climbs from 0.41 to 3.61 and polite words per greedy reply from 0.33 to 12.50, while correct answers fall from 12/12 to 2/12. The typical reply becomes "Thank Thank Thank Thank...". By the number, this is the best model of the three. By Maya's standards, it is useless.
  • B) A KL penalty slows the hack but does not stop it. After 30 steps correct answers are 6/12 and degeneration ("per user per user") has started. The leash buys time, not safety.
  • C) The verifiable reward cannot be gamed this way. The only way to score is to state the correct fact, so correctness stays at 12/12 and the sampled training reward rises from 0.62 to 0.89 (at temperature 1.0 the model now states the fact more reliably). Politeness does not change, because nothing rewarded it. The lessons generalize: every reward is a proxy; watch a metric the reward does not include (here, fact accuracy) during training; read samples, not just the curve; and prefer rewards that check the thing itself.

The shape of runs A and C side by side is the picture to remember:

.