Part E: From Base Model to Assistant: The Lifecycle
A model you call through an API has usually been through several training stages. Each one changes behavior you will observe, so knowing the stages helps you diagnose what you see.
Real pipelines vary in order and often loop (Qwen3's post-training, for example, runs a reasoning cold start, reasoning reinforcement learning, a fusion step that merges thinking and non-thinking behavior, then general reinforcement learning (Qwen3 Technical Report)). The stages are still the right mental model.
Stage 1: Pretraining
What goes in: enormous amounts of text. Llama 3.1 was pretrained on about 15 trillion tokens; Qwen3 on about 36 trillion tokens in 119 languages and dialects. Data work (deduplication, filtering low-quality pages, balancing code, math, and languages) is a large part of the effort.
What it costs: gpt-oss-120b's training took about 2.1 million H100 GPU-hours according to its model card. Pretraining is by far the most expensive stage, and it is why only a handful of organizations produce frontier base models.
What is learned: everything you saw in TinyLM, at scale: language, formats, facts in proportion to how often they appear, and the skills that help predict text. What is not learned: that it should answer questions, follow instructions, stop at the end of its turn, or refuse anything.
Stage 2: Mid-training and continued pretraining
Mid-training is a second phase of next-token training on a smaller, higher-quality or more targeted mix, often with longer sequences. Qwen3's pretraining ended with a reasoning stage of about 5 trillion tokens emphasizing STEM and code, then a long-context stage on hundreds of billions of tokens at 32,768-token sequence length. Ai2's fully open Olmo 3 documents a 100-billion-token mid-training mix ("Dolmino") and a long-context mix ("Longmino") (Ai2, Olmo 3).
Continued pretraining is the same idea applied by someone else later: take a released base model and keep training it on domain text (legal, medical, a company's own documents) so it learns that vocabulary. Module 9 covers when that is worth it; usually retrieval is cheaper.
What you observe: better math and code, a longer usable context window, and fluency in the domain that was added. Still a base model: still no instruction following.
Stage 3: Supervised fine-tuning (instruction following)
A base model continues documents. It does not know that it is supposed to answer, obey, or stop. TinyLM shows all three problems.
"""Module 1: how a base model behaves, and what instruction tuning would change."""
import torch
from supportdesk.tinylm import SamplingParams, generate, load, next_token_distribution
torch.set_num_threads(1) # one thread is plenty for a 1M-parameter model
model, tokenizer = load()
# 1. A base model continues documents. It does not know its turn is over.
prompt = "Customer (Maya): Hi, I forgot my password and the reset link expired.\nAgent (Dara):"
full = generate(model, tokenizer, prompt, SamplingParams(max_new_tokens=60, temperature=0))
print("== Continuation, no stop sequence ==")
print(prompt + full.text)
print(f"[stop_reason={full.stop_reason}, {len(full.token_ids)} tokens]\n")
# 2. The pre-chat-era trick: frame the prompt as a transcript and cut at the next speaker.
cut = generate(model, tokenizer, prompt, SamplingParams(max_new_tokens=60, temperature=0, stop=["\nCustomer"]))
print("== Same prompt, stop=['\\nCustomer'] ==")
print(repr(cut.text), f"[stop_reason={cut.stop_reason}, {len(cut.token_ids)} tokens]\n")
# 3. Instructions are just more text to continue. The base model never learned to obey them.
print("== Instructions to a base model ==")
for instruction in [
"Classify this ticket as billing or bug.\nTicket: I was charged twice this month.\nCategory:",
"Translate to Spanish: Reset links expire after 30 minutes.\nSpanish:",
]:
top = next_token_distribution(model, tokenizer, instruction, top=3)
out = generate(model, tokenizer, instruction, SamplingParams(max_new_tokens=12, temperature=0))
print(repr(instruction.splitlines()[0]))
print(" top next tokens:", ", ".join(f"{t!r} {p:.2f}" for t, p in top))
print(" continuation: ", repr(out.text))
# 4. What an instruction-tuning (SFT) example looks like once it is flattened into text.
# The chat roles become plain markers; training computes loss only on the assistant part.
example = {
"messages": [
{"role": "system", "content": "You are Brightlane's support assistant. Answer from the help center."},
{"role": "user", "content": "Classify this ticket as billing or bug: I was charged twice this month."},
{"role": "assistant", "content": "billing"},
]
}
text, mask = "", []
for m in example["messages"]:
header, body = f"<|{m['role']}|>", f"{m['content']}<|end|>\n"
text += header + body
mask += [0] * len(tokenizer.encode(header).ids) # role marker: context
mask += [int(m["role"] == "assistant")] * len(tokenizer.encode(body).ids) # answer: trained
print("\n== One SFT example as the model sees it ==")
print(text, end="")
print(f"tokens: {len(mask)}, trained on (assistant answer): {sum(mask)}, context only: {len(mask) - sum(mask)}")Code explained
- In simple words: four experiments that show why a base model is not an assistant, and what the training example that would fix it looks like.
- What happens:
- Part 1 generates 60 tokens with no stop condition. A base model has no idea its turn is over, so it keeps writing the document.
- Part 2 uses the trick people used before chat models existed: frame the prompt as a transcript and cut generation at the next speaker label with
stop=["\nCustomer"]. - Part 3 gives two instructions (classify, translate) in the format an assistant would understand, and shows the top next tokens.
- Part 4 flattens one supervised fine-tuning (SFT) example into the text the model would train on. Chat messages become plain text with role markers. The loss mask says which tokens count toward training: only the assistant's answer is trained, the system and user text are context.
loss_onintinylm.pyaccepts exactly such a mask. - Comes out: real output.
== Continuation, no stop sequence ==
Customer (Maya): Hi, I forgot my password and the reset link expired.
Agent (Dara): Reset links expire after 30 minutes and can be used once. Please request a new link from the sign-in page.
Customer (Ana): Thank you.
Customer (Ana): Hello, how do I export my invoice data?
Agent (Ana): Export a board to CSV or JSON
[stop_reason=length, 60 tokens]
== Same prompt, stop=['\nCustomer'] ==
' Reset links expire after 30 minutes and can be used once. Please request a new link from the sign-in page.' [stop_reason=stop, 26 tokens]
== Instructions to a base model ==
'Classify this ticket as billing or bug.'
top next tokens: ' Gantt' 0.22, ' use' 0.12, ' Team' 0.10
continuation: ' Gantt view dependencies and a workspace owner can I forgot my password'
'Translate to Spanish: Reset links expire after 30 minutes.'
top next tokens: ' Gantt' 0.24, ' use' 0.19, ' Settings' 0.08
continuation: ' Gantt view dependencies and a workspace owner can cancel or after upgrading'
== One SFT example as the model sees it ==
<|system|>You are Brightlane's support assistant. Answer from the help center.<|end|>
<|user|>Classify this ticket as billing or bug: I was charged twice this month.<|end|>
<|assistant|>billing<|end|>
tokens: 93, trained on (assistant answer): 9, context only: 84What the base model does:
- It does not stop. After a correct reply it writes the customer's thank-you (with a different customer name, because "Maya" never appeared in training) and starts a new conversation. The stop sequence fixes this mechanically, which is fine for a transcript-shaped prompt.
- It does not follow instructions. "Classify this ticket as billing or bug ... Category:" gets
' Gantt'at 22 percent, because the only "X: Y" line in the corpus that resembles this pattern is "Currently planned for Q4 2026: Gantt view dependencies". To a base model an instruction is just more text to continue. - It cannot do a task it never saw. Translation fails for the same reason; a large base model would likely translate (it has seen bilingual text), but would still not reliably follow your output format.
Supervised fine-tuning fixes these by training on thousands to millions of examples of the shape above: a conversation followed by the ideal response. Real chat formats use dedicated special tokens for the role markers (TinyLM's tokenizer has none, so its markers split into several ordinary tokens, which is why the trained answer, "billing<|end|>" plus a newline, costs 9 tokens). Here is the shape of one Brightlane SFT example in the common JSON Lines chat format:
{"messages": [
{"role": "system", "content": "You are Brightlane's support assistant. Answer from the help center."},
{"role": "user", "content": "Classify this ticket as billing or bug: I was charged twice this month."},
{"role": "assistant", "content": "billing"}
]}Code explained
- In simple words: one training example is one conversation that ends with the answer you want the model to learn to give.
- What happens: the trainer renders
messageswith the model's chat template (the exact text format with role markers and special tokens) and computes loss only on the assistant turn. Thousands of such examples, covering many tasks and formats, teach a general habit: read the instruction, answer in the requested format, stop. - Comes out: nothing runs here. Module 9 builds a real SFT dataset from Brightlane tickets and fine-tunes TinyLM on it, with before-and-after numbers.
What you observe after SFT: the model answers instead of continuing, follows format instructions, stops at the end of its turn, and adopts the tone of the examples. It also inherits the examples' blind spots: if the SFT data never says "I don't know", the model rarely will.
Stage 4: Preference optimization and alignment
SFT teaches the model to imitate good answers. Preference optimization teaches it to prefer better answers over worse ones. People (or a strong model) compare two responses to the same prompt and pick the better one. Then either:
- RLHF (reinforcement learning from human feedback): train a reward model to predict those preferences, then use reinforcement learning to push the language model toward high-reward responses. This is how InstructGPT was trained; its labelers preferred outputs of the 1.3B-parameter InstructGPT over the 175B-parameter GPT-3 (Ouyang et al., 2022).
- DPO (direct preference optimization) and related methods: skip the separate reward model and train directly on (prompt, chosen, rejected) triples with a simple loss (Rafailov et al., 2023). Cheaper and common in open models.
Alignment is the broader goal these methods serve: behavior that is helpful, honest, and harmless according to the developer's policy. It includes refusing some requests, following a system prompt's rules, and avoiding unsafe content.
What you observe: a consistent assistant personality, refusals, hedging, and a tendency to agree with the user (sycophancy), which is a known side effect of optimizing for human approval. One measurable side effect matters for this module: OpenAI's GPT-4 technical report showed the pretrained model's confidence was well calibrated on a multiple-choice benchmark and that post-training reduced calibration (OpenAI, 2023). Tuned models can sound surer than they are. Module 11 covers alignment's limits in practice.
Stage 5: Reasoning training with verifiable rewards
For math, code, and logic puzzles, correctness can be checked automatically: run the unit tests, compare the final number. Reinforcement learning with verifiable rewards (RLVR) exploits this. The model writes a long chain of intermediate reasoning, then an answer; a program checks the answer; the model is rewarded when it is right. Over many rounds it learns to spend more tokens thinking, to check its work, and to backtrack. Qwen3's reasoning RL stage used 3,995 query-verifier pairs with the GRPO algorithm; gpt-oss was post-trained with chain-of-thought reinforcement learning similar to OpenAI's o3 (Qwen3 Technical Report, gpt-oss model card).
What you observe: a reasoning model that emits hidden or visible "thinking" tokens before answering, is much stronger on multi-step math and code, and is slower and more expensive per answer because those thinking tokens are generated and billed. Many expose a dial: reasoning_effort (low, medium, high) for gpt-oss, a thinking level for Gemini, /think and /no_think switches for Qwen3. The reward only checks what the verifier checks, so models can learn shortcuts that satisfy the checker without solving the problem (reward hacking, Module 9).
Which stage explains what you see
| Stage | Changes | Behavior you will observe later in the course |
|---|---|---|
| Pretraining | Knowledge, language, skills, memorized text | Hallucination on rare facts, knowledge cutoff, memorized phrases |
| Mid-training / continued pretraining | Domain fluency, math and code, long context | Better long-context recall; vocabulary of a domain |
| Supervised fine-tuning | Instruction following, format, stopping, tone | Obeys system prompts, returns JSON when asked, answers instead of continuing |
| Preference optimization | Preferred style, refusals, safety behavior | Hedging, refusals, sycophancy, confident tone, weaker calibration |
| Reasoning training | Long thinking before answering | Better math, code, and multi-step tasks; more output tokens, cost, and latency |
Base, instruction-tuned, reasoning-tuned, aligned
These labels describe where a model stopped in the pipeline, and they overlap: almost every instruction-tuned model you can call has also been preference-tuned (aligned), and reasoning models are instruction-tuned models with extra RL. Model names often signal it: -Base for base checkpoints, -Instruct or -it for instruction-tuned ones, and "thinking" or an effort setting for reasoning.
| Situation | Use this | Why |
|---|---|---|
| Continue text in a fixed style, score how likely text is, or start your own fine-tune | Base model | No assistant habits to fight; probabilities reflect the data |
| Triage, extraction, drafting, chat: most application work | Instruction-tuned (aligned) model | Follows instructions and formats, stops cleanly, safer defaults |
| Multi-step math, code, planning, hard analysis where accuracy beats speed | Reasoning model, at the lowest effort that passes your eval | Verifiable-reward training helps exactly here; effort costs tokens and time |
| High-volume simple classification with tight latency | Small instruction-tuned model, reasoning off | Thinking tokens add cost and delay without helping easy tasks |
Open weights is not open source
Open weights means you can download the trained parameters and run them yourself. It says nothing on its own about what you may do with them, or whether you can see how they were made. Read the licence.
- gpt-oss and Qwen3 are released under Apache 2.0, a permissive open-source software licence: commercial use, modification, and redistribution are allowed with attribution.
- Llama 3.1 uses the Llama 3.1 Community License. It allows commercial use but adds conditions: products with more than 700 million monthly active users (on the release date) must request a separate licence from Meta; distributors must display "Built with Llama"; and a model trained using Llama materials must have a name beginning with "Llama" when distributed (Llama 3.1 license). That is open weights, not open source in the usual sense.
- Open source AI, as defined by the Open Source Initiative's Open Source AI Definition 1.0 (28 October 2024), requires more than weights: sufficiently detailed information about the training data for a skilled person to build a substantially equivalent system, the complete code used to train and run it, and the parameters, all under terms that allow use, study, modification, and sharing for any purpose (OSI). Most "open" models do not publish their training data or full training code.
- Fully open models exist. Ai2's Olmo 3 (November 2025, 7B and 32B, Apache 2.0) released its data (the roughly 9.3-trillion-token Dolma 3 corpus and its training mixes), training code, and intermediate checkpoints from each stage (Ai2). That is what lets researchers study a model's lifecycle end to end.
| Situation | Use this | Why |
|---|---|---|
| You must run inside your own network and may ship commercially | Apache 2.0 or MIT open-weight model | Fewest legal conditions; your data never leaves |
| You need to audit or reproduce training data and code (research, strict compliance) | A fully open model (weights, data, code) | Only these let you inspect what the model learned from |
| A custom licence looks attractive for quality | Read the licence with your legal team first | Usage caps, naming, attribution, and acceptable-use clauses apply to you |
| You plan to fine-tune on outputs of another provider's model | Check both licences and the provider's terms | Some terms restrict using outputs to train competing models (Module 9) |