CourseLarge Language Models · Module 9: Adaptation: Fine-Tuning and Customization · part 46 of 80
Part 46 · Module 9: Adaptation: Fine-Tuning and Customization

Part 3: Data

12 min read·22 Sept 2026

Building a supervised fine-tuning dataset

An SFT dataset is a list of (prompt, completion) pairs in exactly the format you will use at inference time. Three rules matter more than any hyperparameter:

  1. Same format as production. If the live system sends Ticket: ...\nCategory:, train on that, character for character. A model tuned on one template and served with another loses much of what it learned.
  2. Loss on the completion only. encode_example in the setup file masks the prompt, so the model is graded on the label, not on reproducing the ticket.
  3. Split first, then do everything else. Augmentation, deduplication, and synthetic generation all happen inside the training split. Otherwise copies of validation or test items leak into training and your eval measures memory.

Our only labeled data is the 48 dev tickets. The 24 test tickets are sealed by the eval plan. The data script splits dev into train and validation (for early experiments), scrubs identifiers, adds variants, generates synthetic tickets, deduplicates, and checks for contamination.

examples/m09_data.py

python
"""Build the SFT dataset for ticket triage: split, scrub, augment, synthesize, dedupe, check contamination."""
import json
import random
import re
from collections import Counter

from examples.m09_setup import DATA_OUT
from supportdesk.data import CATEGORIES, load_tickets

rng = random.Random(9)
WORDS = re.compile(r"\w+", re.UNICODE)


def row(subject, body, category, source):
    return {"subject": subject, "body": body, "category": category, "source": source}


# 1. Split the dev tickets into train and validation BEFORE anything else touches them.
def stratified_split(tickets, val_share=0.25):
    train, val = [], []
    for c in CATEGORIES:
        group = [t for t in tickets if t.gold["category"] == c]
        rng.shuffle(group)
        n_val = max(1, round(len(group) * val_share))
        val += group[:n_val]
        train += group[n_val:]
    return train, val


# 2. Scrub identifiers the way you would scrub production logs.
PII = [(re.compile(r"INV-\d{4}-\d+"), "INV-XXXX"), (re.compile(r"\b[A-Z]{2}\d{9}\b"), "VATID"),
       (re.compile(r"[\w.+-]+@[\w-]+\.[\w.]+"), "EMAIL"), (re.compile(r"\b\d{3,}\s?USD\b"), "AMOUNT USD")]


def scrub(text):
    hits = 0
    for pattern, placeholder in PII:
        text, n = pattern.subn(placeholder, text)
        hits += n
    return text, hits


# 3. Label-preserving variants of real tickets (templated augmentation).
GREETINGS = ["Hi team, ", "Hello, ", "Hey Brightlane, ", ""]
SIGN_OFFS = [" Thanks.", " Thanks, Sam", " Regards, Priya", ""]


def augment(r):
    out = [row(r["subject"], r["body"].lower(), r["category"], "aug:lower"),
           row(r["subject"], rng.choice(GREETINGS) + r["body"] + rng.choice(SIGN_OFFS), r["category"], "aug:wrap"),
           row(r["body"][:60], r["body"], r["category"], "aug:subject-from-body"),
           row(r["subject"], r["subject"] + ".", r["category"], "aug:subject-only")]
    swapped = r["body"].replace("Team", "__B__").replace("Business", "Team").replace("__B__", "Business")
    if swapped != r["body"]:
        out.append(row(r["subject"], swapped, r["category"], "aug:plan-swap"))
    return out


# 4. Synthetic tickets from templates: cheap, clean-looking, and narrow.
TEMPLATES = {
    "billing": ["I was charged twice for the {plan} plan.", "Where do I download my {plan} invoices?",
                "How much does {plan} cost per user?", "Can I pay for {plan} by bank transfer?"],
    "cancellation": ["How do I cancel my {plan} subscription?", "Can I get a refund if I cancel {plan} now?",
                     "Please cancel our {plan} plan at the end of the term."],
    "account_access": ["I cannot log in, my account is locked.", "The password reset link expired.",
                       "I lost my 2FA device and cannot sign in.", "SSO sign-in fails for our {plan} workspace."],
    "bug": ["The {area} is broken and shows an error.", "The {area} stopped working since yesterday.",
            "The {area} crashes when I open it."],
    "how_to": ["How do I export my {area} data?", "How do I set up the {area}?", "Does {plan} include the {area}?"],
    "feature_request": ["Please add {feature}.", "When will you support {feature}?", "We would love {feature}."],
}
SLOTS = {"plan": ["Free", "Team", "Business", "Enterprise"],
         "area": ["board", "automation", "export", "Slack integration", "mobile app", "calendar view"],
         "feature": ["a Gantt chart", "dark mode", "a Jira integration", "offline editing", "custom fields"]}


def synthesize(per_category=20):
    rows = []
    for c, templates in TEMPLATES.items():
        for _ in range(per_category):
            text = rng.choice(templates).format(**{k: rng.choice(v) for k, v in SLOTS.items()})
            rows.append(row(text[:40], text, c, "synthetic"))
    return rows


# 5. Deduplication and contamination checks, both on word n-grams.
def ngrams(text, n):
    words = [w.lower() for w in WORDS.findall(text)]
    return {tuple(words[i:i + n]) for i in range(len(words) - n + 1)}


def jaccard(a, b):
    return len(a & b) / max(len(a | b), 1)


def dedupe(rows, threshold=0.8):
    kept, seen = [], []
    for r in rows:
        grams = ngrams(r["subject"] + " " + r["body"], 3) or {(r["body"].lower(),)}
        if any(jaccard(grams, s) >= threshold for s in seen):
            continue
        kept.append(r)
        seen.append(grams)
    return kept


def contamination(train_rows, eval_tickets, n=8, near=0.5):
    """Eval tickets that share any word n-gram with training text, or are near-duplicates of it."""
    train_long = set().union(*(ngrams(r["subject"] + " " + r["body"], n) for r in train_rows))
    train_tri = [ngrams(r["subject"] + " " + r["body"], 3) for r in train_rows]
    flagged = []
    for t in eval_tickets:
        text = t["subject"] + " " + t["body"] if isinstance(t, dict) else t.subject + " " + t.body
        shared = ngrams(text, n) & train_long
        best = max((jaccard(ngrams(text, 3), g) for g in train_tri), default=0)
        if shared or best >= near:
            flagged.append((t["subject"] if isinstance(t, dict) else t.id, len(shared), round(best, 2)))
    return flagged


def as_row(t, source):
    return row(t.subject, t.body, t.gold["category"], source)


def save(name, rows):
    with (DATA_OUT / f"{name}.jsonl").open("w", encoding="utf-8") as f:
        for r in rows:
            f.write(json.dumps(r, ensure_ascii=False) + "\n")


if __name__ == "__main__":
    DATA_OUT.mkdir(parents=True, exist_ok=True)
    dev, test = load_tickets("dev"), load_tickets("test")
    train_t, val_t = stratified_split(dev)
    print(f"dev {len(dev)} -> train {len(train_t)}, val {len(val_t)}; test {len(test)} untouched")

    real, pii = [], 0
    for t in train_t:
        subject, h1 = scrub(t.subject)
        body, h2 = scrub(t.body)
        pii += h1 + h2
        real.append(row(subject, body, t.gold["category"], "real"))
    print(f"scrubbed {pii} identifiers from training tickets")

    aug = [a for r in real for a in augment(r)]
    synth = synthesize()
    before = len(real) + len(aug) + len(synth)
    real_d = dedupe(real)
    aug_d = dedupe(real_d + aug)[len(real_d):]
    synth_d = dedupe(synth)
    print(f"rows before dedupe {before}: real {len(real)}, augmented {len(aug)}, synthetic {len(synth)}")
    print(f"rows after dedupe  {len(real_d) + len(aug_d) + len(synth_d)}: real {len(real_d)}, "
          f"augmented {len(aug_d)}, synthetic {len(synth_d)}")
    print("synthetic label counts:", dict(Counter(r["category"] for r in synth_d)))

    def diversity(rows):
        grams = [g for r in rows for g in ngrams(r["body"], 2)]
        return len(set(grams)) / max(len(grams), 1)

    print(f"distinct-bigram ratio: real {diversity(real_d):.2f}, synthetic {diversity(synth_d):.2f}")

    everything = real_d + aug_d + synth_d
    val_rows = [as_row(t, "val") for t in val_t]
    print("contamination train vs test:", contamination(everything, test))
    print("contamination train vs val: ", contamination(everything, val_rows))
    closest = sorted(((round(max(jaccard(ngrams(t.subject + " " + t.body, 3), ngrams(r["subject"] + " " + r["body"], 3))
                                 for r in everything), 2), t.id) for t in test), reverse=True)[:3]
    print("closest test tickets to any training row (3-gram Jaccard):", closest)

    # The classic mistake: augment all of dev first, split afterwards.
    leaky = [a for t in dev for a in augment(as_row(t, "real"))]
    print("if you augment BEFORE splitting, val tickets leaking into train:",
          len(contamination(leaky, val_rows)), "of", len(val_rows))

    # For the final model (after hyperparameters are chosen) we train on ALL 48 dev tickets.
    dev_real = [row(scrub(t.subject)[0], scrub(t.body)[0], t.gold["category"], "real") for t in dev]
    dev_aug = dedupe(dev_real + [a for r in dev_real for a in augment(r)])[len(dev_real):]
    print("contamination all-dev training rows vs test:", contamination(dev_real + dev_aug, test))
    save("dev_real", dev_real)
    save("dev_aug", dev_aug)
    save("train_real", real_d)
    save("train_aug", aug_d)
    save("train_synth", synth_d)
    save("val", val_rows)
    print("wrote", sorted(p.name for p in DATA_OUT.glob("*.jsonl")))

Code explained

  • In simple words: turn 48 labeled tickets into a clean training set, and prove none of it overlaps the tickets we test on.
  • What happens:
    1. stratified_split: 25% of each category goes to validation, the rest to train, so rare categories appear in both.
    2. scrub: replaces invoice numbers, VAT ids, emails, and amounts with placeholders, the same thing you must do with production logs before they go anywhere near a training job.
    3. augment: label-preserving variants of each real ticket: lower-cased, wrapped in a greeting and sign-off, a subject made from the body, subject only, and Team/Business swapped. Cheap and useful, but every variant is still that ticket.
    4. synthesize: 20 synthetic tickets per category from fill-in-the-blank templates.
    5. ngrams, jaccard, dedupe: word n-grams are runs of n consecutive words. Two texts are near-duplicates if their sets of 3-grams overlap heavily (Jaccard similarity, shared over total, at or above 0.8). dedupe keeps the first of each near-duplicate group.
    6. contamination: flags an evaluation ticket if it shares any 8-word run with the training text, or if its 3-gram Jaccard with any training row is 0.5 or more. This is the same check large labs run between pretraining data and benchmarks, at a smaller scale.
    7. The main block prints counts, label balance, a diversity score (distinct bigrams over all bigrams), the contamination checks, and the classic mistake (augment all of dev first, split afterwards). Finally it writes dev_real and dev_aug (all 48 dev tickets plus variants) for the final models, and the 35/13 split files for experiments. Run it with python -m examples.m09_data.
  • Comes out:
text
dev 48 -> train 35, val 13; test 24 untouched
scrubbed 2 identifiers from training tickets
rows before dedupe 306: real 35, augmented 151, synthetic 120
rows after dedupe  206: real 35, augmented 103, synthetic 68
synthetic label counts: {'billing': 12, 'cancellation': 9, 'account_access': 6, 'bug': 13, 'how_to': 16, 'feature_request': 12}
distinct-bigram ratio: real 0.94, synthetic 0.42
contamination train vs test: []
contamination train vs val:  []
closest test tickets to any training row (3-gram Jaccard): [(0.12, 'T-1009'), (0.09, 'T-1006'), (0.06, 'T-1030')]
if you augment BEFORE splitting, val tickets leaking into train: 13 of 13
contamination all-dev training rows vs test: []
wrote ['dev_aug.jsonl', 'dev_real.jsonl', 'train_aug.jsonl', 'train_real.jsonl', 'train_synth.jsonl', 'val.jsonl']
  • Scrubbing found 2 identifiers (an invoice number and an amount) in 35 training tickets. In real logs expect far more, including names and addresses a regex will not catch, which is why log data also needs human review (below).
  • Dedupe removed 100 of 306 rows. All 35 lower-cased variants went, because our n-grams are lower-cased words, so they are exact duplicates at the word level. Near-identical synthetic tickets went too (120 down to 68). The templates only produce so many distinct sentences.
  • The distinct-bigram ratio is 0.94 for real tickets and 0.42 for synthetic ones: the synthetic set is less than half as varied. That number is the degradation risk in one line (next section).
  • No training row overlaps the test or validation sets. The closest test ticket to anything in training has a Jaccard of 0.12, far from the 0.5 flag.
  • The leak demo: if you augment all 48 dev tickets before splitting, 13 of 13 validation tickets have a variant in the training set. Your validation accuracy would then be a memory test. It is the most common way small fine-tuning projects fool themselves.

Synthetic data and its degradation risks

Synthetic data is training data produced by a program or a model rather than collected. It is tempting because it is free and perfectly labeled. It has three failure modes you can measure:

  • Narrowness. Our template tickets have a diversity ratio of 0.42 against 0.94 for real ones. A model trained mostly on them learns the templates ("I was charged twice for the {plan} plan") and not the messy ways customers actually write, in five languages.
  • Label noise you did not intend. "Can I get a refund if I cancel Team now?" is labeled cancellation by the template, but a similar real ticket might be billing. Templates encode the author's assumptions.
  • Collapse over generations. When models are trained on model-generated text, then generate more text for the next round, the rare cases (the tails of the distribution) vanish first and quality degrades over generations (Shumailov et al., "AI models collapse when trained on recursively generated data", Nature, 2024).

The practical rules: keep real data as the anchor, cap the synthetic share, measure diversity, and always evaluate on real held-out data only. Part 4's sweep runs the model with and without the synthetic rows so you can see whether they help at all.

Data from production logs, with review

Brightlane's best source of training examples is its own support history: every ticket an agent answered is a (question, reply) pair. It is also the most dangerous source, because logs contain whatever agents actually did, including curt replies and wrong answers, plus customer personal data. Logs need three filters before training: scrubbing (as scrub does above, plus human review for what a regex misses), automatic quality checks, and a human review of a sample. Because measuring what bad log rows do to a model needs the reply-writing setup from Part 5, the log-review experiment (and the "how few examples is enough" curve) runs there, right after DPO.

Distillation from a stronger model, and licence constraints

Distillation means using a stronger model's outputs as training data for a smaller one: label your tickets with a big model, then fine-tune a small, cheap, fast model on those labels. m09_hosted_baseline.py above writes exactly that file when run with a key, with the teacher's name and the date on every row. Two things decide whether you may do it.

The provider's terms. Most commercial API terms restrict using outputs to build competing models. As checked on 21 September 2026:

ProviderWhat the terms saySource
OpenAI (API)You may not "use Output to develop artificial intelligence models that compete with OpenAI's products and services". The agreement defines a "Permitted Exception" that includes classifier-type models that are not distributed commercially, and fine-tuning through OpenAI's own services.OpenAI Services Agreement, section 3.3(e), effective 1 January 2026
Anthropic (API)Customers may not access the Services "to build a competing product or service, including to train competing AI models", except as expressly approved.Anthropic Commercial Terms, section D.4, effective 17 June 2025
Google (Gemini API)"You may not use the Services to develop models that compete with the Services (e.g., Gemini API or Google AI Studio)."Gemini API Additional Terms, last updated 28 April 2026
Meta Llama 3.1 (open weights)Using Llama outputs to train or improve another model is allowed, but a distributed model must have "Llama" at the beginning of its name. Companies with over 700 million monthly active users need a separate licence.Llama 3.1 Community License, sections 1.b.i and 2
OpenAI gpt-oss (open weights, the course's Groq default)Released under Apache 2.0 plus a gpt-oss usage policy. When you call it through a host such as Groq, the host's terms apply to the service as well.OpenAI help center on gpt-oss

A triage classifier used only inside Brightlane is not obviously "a competing model", but "obviously" is a lawyer's call, not yours. Read the current terms of the exact service you call, record the URL and version in the dataset (the terms_checked field), and ask legal before distilling anything you will distribute.

The data itself. Distilled labels are only as good as the teacher: compare them with human labels on a sample (the script prints the agreement on the 48 dev tickets). And customer tickets sent to a third-party API are a data-processing question under your privacy agreements (Module 11 covers PII handling; a fine-tuned model can also memorize and repeat personal data from its training rows).

SituationUse thisWhy
You have fewer than a few hundred real labeled examplesReal examples plus light augmentation, reviewedReal data carries the variety that templates miss; augmentation must happen after the split
You need coverage of rare cases you have no real examples ofA small, reviewed synthetic set, capped as a share of trainingFills gaps without letting template text dominate; diversity is measurable
You have a strong model that labels well and the terms allow itDistillation, with provenance on every rowCheapest source of many labels; provenance lets you drop them later if terms or quality change
You have years of production logsLogs, scrubbed, auto-checked, and sampled for human reviewThe most realistic data, and the most likely to teach bad habits and leak personal data
Terms are unclear or the data is personalStop and ask legal or privacy before trainingA trained model cannot easily "forget" a row you should not have used