CourseLarge Language Models · Module 3: Inference Behaviour and Decoding Control · part 15 of 80
Part 15 · Module 3: Inference Behaviour and Decoding Control

Part D: Reasoning-model behaviour

22 min read·22 Sept 2026

Thinking tokens, budgets, and effort levels

A reasoning model is trained (Module 1 covered reasoning training with verifiable rewards) to generate a stretch of working-out text before its answer. Those thinking tokens are decoded exactly like any other tokens: one sequential step each, billed at the output rate, and they occupy context. Some providers return them, some return a summary, some hide them entirely but still bill them.

You control how much thinking happens with an effort level or a budget:

Provider and modelParameterValuesNotes
Groq, openai/gpt-oss-120b and openai/gpt-oss-20breasoning_effortlow, medium, highTrace returned in message.reasoning unless include_reasoning: false
Groq, Qwen reasoning modelsreasoning_effort, reasoning_formateffort includes none to disable; format parsed, raw, hiddeninclude_reasoning and reasoning_format are mutually exclusive
Gemini 3.x (native API)thinking_levelminimal (some models), low, medium, highGoogle notes minimal "does not guarantee that thinking is off"; the thinking page lists gemini-3.5-flash as thinking on by default
Gemini 3.x via OpenAI-compatible endpointreasoning_effortlow, medium, highMapped to thinking_level
Gemini 2.5thinking_budget (tokens)a token count; reasoning_effort maps low, medium, high to 1,024, 8,192, 24,576Still accepted for backward compatibility; do not send both a budget and a level

llm.chat passes reasoning_effort straight through and reads usage.completion_tokens_details.reasoning_tokens into Usage.reasoning_tokens when the provider reports it. Reasoning tokens are already included in output_tokens (the provider's completion_tokens), so cost_usd bills them correctly without double counting. For qwen3:8b on Ollama, thinking is controlled by Ollama's own think option; check your Ollama version's documentation for how its OpenAI-compatible endpoint maps reasoning_effort.

Latency and cost implications

Reasoning tokens are output tokens, and output tokens are both the expensive and the slow half (Part A, Module 2). The math is simple enough to do before you spend anything:

examples/m03_reasoning_cost.py:

python
"""Cost and latency math for reasoning effort levels, using pricing.py.

The reasoning-token counts below are ASSUMPTIONS for a short triage task, not
measurements. Run examples/m03_reasoning_measure.py with a key to get yours,
then put your medians into REASONING_TOKENS.
"""
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, cost_usd

INPUT_TOKENS = 450           # system prompt + one ticket (Module 2 showed how to count this)
VISIBLE_OUTPUT = 60          # the triage answer itself
REASONING_TOKENS = {"none (standard model)": 0, "low": 150, "medium": 600, "high": 2500}   # assumptions
DECODE_TPS = {"openai/gpt-oss-120b": 500}   # Groq's model page lists about 500 tokens/s (checked Sep 2026)
TICKETS_PER_MONTH = 30_000

model = "openai/gpt-oss-120b"
print(f"{model}: input {PRICES[model].input} / output {PRICES[model].output} USD per 1M tokens\n")
print(f"{'effort':>22} | {'output tokens':>13} | {'USD per call':>12} | {'USD per month':>13} | {'decode seconds':>14}")
for effort, reasoning in REASONING_TOKENS.items():
    # Reasoning tokens are billed as output tokens; providers include them in completion_tokens.
    usage = Usage(input_tokens=INPUT_TOKENS, output_tokens=VISIBLE_OUTPUT + reasoning, reasoning_tokens=reasoning)
    per_call = cost_usd(usage, model)
    seconds = usage.output_tokens / DECODE_TPS[model]
    print(f"{effort:>22} | {usage.output_tokens:>13} | {per_call:>12.6f} | {per_call * TICKETS_PER_MONTH:>13.2f} |"
          f" {seconds:>14.2f}")

print("\nSame token counts on gemini-3.5-flash (output price", PRICES["gemini-3.5-flash"].output, "USD per 1M):")
for effort, reasoning in REASONING_TOKENS.items():
    usage = Usage(input_tokens=INPUT_TOKENS, output_tokens=VISIBLE_OUTPUT + reasoning, reasoning_tokens=reasoning)
    print(f"{effort:>22} | USD per month {cost_usd(usage, 'gemini-3.5-flash') * TICKETS_PER_MONTH:>9.2f}")

Code explained

  • In simple words: multiply out what a triage call costs and how long it spends decoding at each effort level, per call and per month.
  • What happens: for a 450-token prompt and a 60-token visible answer, each effort level adds an assumed number of reasoning tokens. The script builds a Usage, prices it with cost_usd from pricing.py, scales it to 30,000 tickets a month, and divides output tokens by Groq's published speed for gpt-oss-120b (about 500 tokens per second on its model page, checked September 2026) to estimate decode time. The reasoning-token counts are assumptions, not measurements, and are labelled as such in the file.
  • Comes out:
text
openai/gpt-oss-120b: input 0.15 / output 0.6 USD per 1M tokens

                effort | output tokens | USD per call | USD per month | decode seconds
 none (standard model) |            60 |     0.000103 |          3.10 |           0.12
                   low |           210 |     0.000193 |          5.80 |           0.42
                medium |           660 |     0.000463 |         13.90 |           1.32
                  high |          2560 |     0.001604 |         48.11 |           5.12

Same token counts on gemini-3.5-flash (output price 9.0 USD per 1M):
 none (standard model) | USD per month     36.45
                   low | USD per month     76.95
                medium | USD per month    198.45
                  high | USD per month    711.45

At these assumptions, going from low to high effort multiplies output tokens by about 12, cost per call by about 8, and decode time from about 0.4 to about 5 seconds before the first visible word appears. On a model with a higher output price the same token counts cost 15 times more (the Gemini lines). Prices in pricing.py are labelled as needing verification; Groq's model page currently lists a lower cached-input price for gpt-oss-120b (0.075 USD per million) than pricing.py does (0.15), which does not affect these uncached numbers.

Replace the assumptions with measurements from your own tickets:

examples/m03_reasoning_measure.py:

python
"""Measure reasoning tokens, latency, cost, and accuracy per effort level on dev tickets (needs a key)."""
import os
import statistics
import sys

from supportdesk.data import CATEGORIES, load_tickets
from supportdesk.llm import PROVIDERS, chat, resolve
from supportdesk.pricing import PRICES, cost_usd

provider, model = resolve()
key_env = PROVIDERS[provider]["key_env"]
if key_env and not os.environ.get(key_env):
    sys.exit(f"Set {key_env} to run this example against {provider}.")

SYSTEM = ("Classify the Brightlane support ticket into exactly one category: "
          + ", ".join(CATEGORIES) + ". Reply with the category name only.")
tickets = load_tickets("dev")[:12]          # small on purpose: 12 tickets x 3 levels = 36 calls

for effort in ("low", "medium", "high"):
    rows = []
    for t in tickets:
        r = chat([{"role": "system", "content": SYSTEM}, {"role": "user", "content": t.text}],
                 reasoning_effort=effort, temperature=None, max_tokens=4000)
        answer = r.text.strip().lower()
        rows.append((answer == t.gold["category"], r.usage.reasoning_tokens, r.latency_ms,
                     cost_usd(r.usage, model) if model in PRICES else float("nan"), r.finish_reason))
    correct = sum(ok for ok, *_ in rows)
    print(f"{effort:>6}: correct {correct}/{len(rows)} | reasoning tokens median "
          f"{statistics.median(x[1] for x in rows):.0f} (max {max(x[1] for x in rows)}) | "
          f"latency median {statistics.median(x[2] for x in rows):.0f} ms | "
          f"cost per call {statistics.mean(x[3] for x in rows):.6f} USD | "
          f"finish reasons {sorted(set(x[4] for x in rows))}")

Code explained

  • In simple words: classify 12 real dev tickets at each effort level and record accuracy, reasoning tokens, latency, and cost.
  • What happens: a short system prompt asks for the category name only. Each call passes reasoning_effort and temperature=None (provider default). max_tokens=4000 leaves room for the thinking, because on most providers the cap covers reasoning plus answer: set it too low and you get finish_reason="length" and an empty answer. The script prints medians and the set of finish reasons so truncation is visible.
  • Comes out: without a key this build exits with the missing-key message. Illustrative sample run (not captured in this build; produced for teaching). Your output will differ.

    textCopy

    text
       low: correct 10/12 | reasoning tokens median 96 (max 310) | latency median 610 ms | cost per call 0.000170 USD | finish reasons ['stop']
    medium: correct 10/12 | reasoning tokens median 380 (max 1150) | latency median 1320 ms | cost per call 0.000340 USD | finish reasons ['stop']
      high: correct 11/12 | reasoning tokens median 1450 (max 3900) | latency median 3900 ms | cost per call 0.000990 USD | finish reasons ['length', 'stop']
    

    When you run it, read it the way this module reads every table. With 12 tickets, 10/12 vs 11/12 is one ticket: noise. Look at the cost and latency ratios, which are large and stable, and at the finish reasons, where length means the thinking consumed the whole budget. Then run on all 48 dev tickets before deciding.

When extended reasoning helps and when it wastes money

SituationUse thisWhy
Ticket triage into six categoriesLowest effort, or a non-reasoning modelPattern recognition, not multi-step logic; thinking adds cost and latency, rarely accuracy
Refund eligibility across plan, date, and policy rulesMedium effort, verify against the rules in codeMulti-step conditions are where reasoning measurably helps
Agent planning a sequence of tool calls (Module 8)Medium to high effortPlanning errors are expensive; reasoning reduces them
Streaming customer-facing chatLow effortTTFT includes all thinking; users feel every hidden token
Math, code, or anything with a checkable answerHigher effort plus an automatic checkReasoning training targeted exactly these; verify rather than trust
You cannot tellMeasure on your eval set at two effort levelsThe measurement script above; decide on cost per correct answer, not cost per call

A useful single number is cost per correct answer: total cost divided by the number of correct outputs. If high effort costs 6 times more and fixes 1 extra ticket in 48, its cost per correct answer is far worse. If it turns an agent's 60 percent task success into 90 percent and failures cost a human 10 minutes each, it pays for itself immediately.

Reading, and not over-trusting, reasoning traces

A visible trace is useful for debugging: it can show that the model misread the ticket, applied the wrong policy, or never considered the refund window. Here is how to read one from Groq:

examples/m03_read_trace.py:

python
"""Read a gpt-oss reasoning trace on Groq (needs GROQ_API_KEY). Treat the trace as a clue, not a proof."""
import os
import sys

from supportdesk.llm import make_client

if not os.environ.get("GROQ_API_KEY"):
    sys.exit("Set GROQ_API_KEY to run this example.")

ticket = ("Subject: Charged after cancelling\n\nI cancelled our Team plan last week but was charged "
          "again today. Refund please, and make sure it does not happen again.")
response = make_client("groq").chat.completions.create(
    model="openai/gpt-oss-120b",
    messages=[{"role": "user", "content": "Category (billing or cancellation) and one-line reason:\n" + ticket}],
    reasoning_effort="medium",
    extra_body={"include_reasoning": True},          # Groq returns the trace in message.reasoning
)
message = response.choices[0].message
print("REASONING TRACE:\n", getattr(message, "reasoning", None))
print("\nANSWER:\n", message.content)
details = response.usage.completion_tokens_details
print("\nreasoning tokens billed:", getattr(details, "reasoning_tokens", None) if details else None)

Code explained

  • In simple words: ask gpt-oss to categorize an ambiguous ticket and print its reasoning next to its answer.
  • What happens: llm.chat returns text only, so this example uses make_client("groq") directly. include_reasoning: True goes through extra_body and Groq returns the trace in message.reasoning; the OpenAI Python SDK keeps unknown fields, so getattr reads it. The script also prints the billed reasoning tokens.
  • Comes out: without a key this build exits with the missing-key message. Illustrative sample run (not captured in this build; produced for teaching). Your output will differ.
text
REASONING TRACE:
 The customer cancelled the Team plan but was charged again. Main ask is a refund of a charge: billing. Could be cancellation, but the cancellation already happened; the issue is the charge. Answer billing.

ANSWER:
 billing: the customer wants a refund for a charge after cancelling.

reasoning tokens billed: 71

That trace reads well. The research says to hold it loosely:

  • Turpin et al., "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting" (NeurIPS 2023), biased prompts (for example, always making the answer "(A)" in the few-shot examples). Models followed the bias, and their explanations rationalized the biased answer without mentioning the bias. Accuracy dropped by as much as 36 percent across 13 BIG-Bench Hard tasks.
  • Anthropic's "Reasoning Models Don't Always Say What They Think" (April 2025) gave reasoning models hints to the answer and checked whether the chain of thought admitted using them. Claude 3.7 Sonnet mentioned the hint 25 percent of the time and DeepSeek R1 39 percent. When trained into reward hacks, models verbalized them less than 2 percent of the time in most settings, and unfaithful traces were on average longer than faithful ones, so length is not a sign of honesty.

Practical rules for the Brightlane team:

  1. Use traces to form hypotheses when debugging ("it thinks the refund window is 30 days"), then test the hypothesis with a changed input, not by reading more traces.
  2. Never use a trace as the justification shown to a customer or an auditor. If you need a justification, have the model cite the knowledge-base article and check that the citation exists (Module 7).
  3. Never grade correctness from the trace. Grade the answer (Module 10).
  4. Remember that some providers return summaries rather than the raw trace, and that hidden reasoning is billed whether or not you see it.

Module Lab

The lab combines the module into one script the Brightlane assistant can build on: decoding presets per task, a translation from presets to llm.chat arguments, a streaming loop that never leaks stop sequences, constrained triage with a pressure check that routes weak answers to a human, and a latency and cost report.

examples/m03_lab.py

python
"""Module 3 lab: decoding presets, safe streaming, constrained triage, and a latency and cost report."""
import statistics
import time

import torch

from examples.m03_constrained import prompt_for, score_options
from examples.m03_setup import MODEL, TOK
from supportdesk.data import CATEGORIES, load_tickets
from supportdesk.llm import Usage
from supportdesk.pricing import cost_usd
from supportdesk.tinylm import SamplingParams, apply_sampling

# 1. Decoding presets by task. TinyLM uses all fields; hosted APIs get the subset they support.
PRESETS = {
    "triage": SamplingParams(temperature=0, max_new_tokens=8),
    "reply_draft": SamplingParams(temperature=0.7, top_p=0.9, max_new_tokens=60, stop=["\nCustomer"], seed=0),
    "brainstorm": SamplingParams(temperature=1.0, min_p=0.1, max_new_tokens=40, stop=["\n"], seed=0),
}


def to_chat_kwargs(params: SamplingParams) -> dict:
    """Translate a preset into llm.chat keyword arguments. top_k, min_p, and logit_bias are left out:
    the Chat Completions API has no top_k or min_p, and Groq rejects logit_bias."""
    kwargs = {"temperature": params.temperature, "max_tokens": params.max_new_tokens,
              "top_p": params.top_p, "stop": params.stop[:4] or None, "seed": params.seed}
    extra = {k: v for k, v in (("frequency_penalty", params.frequency_penalty),
                               ("presence_penalty", params.presence_penalty)) if v}
    if extra:
        kwargs["extra"] = extra
    return {k: v for k, v in kwargs.items() if v is not None}


# 2. Streaming that never shows the stop sequence: hold back any tail that could be the start of one.
@torch.no_grad()
def safe_stream(prompt: str, params: SamplingParams):
    started = time.perf_counter()
    generator = torch.Generator().manual_seed(params.seed) if params.seed is not None else None
    ids = TOK.encode(prompt).ids[-(MODEL.cfg.context - params.max_new_tokens):]
    caches = [dict() for _ in MODEL.blocks]
    logits = MODEL(torch.tensor([ids]), caches)[0, -1]
    generated, sent = [], 0
    for _ in range(params.max_new_tokens):
        probs = apply_sampling(logits, params, generated)
        next_id = int(torch.multinomial(probs, 1, generator=generator)) if params.temperature > 0 else int(probs.argmax())
        generated.append(next_id)
        text = TOK.decode(generated)
        hits = [text.index(s) for s in params.stop if s in text]
        if hits:                                             # a stop sequence completed: flush up to it, end
            if min(hits) > sent:
                yield text[sent:min(hits)], (time.perf_counter() - started) * 1000
            return
        hold = max((k for s in params.stop for k in range(1, len(s)) if text.endswith(s[:k])), default=0)
        if len(text) - hold > sent:
            yield text[sent:len(text) - hold], (time.perf_counter() - started) * 1000
            sent = len(text) - hold
        if len(ids) + len(generated) >= MODEL.cfg.context:
            break
        logits = MODEL(torch.tensor([[next_id]]), caches, start_pos=len(ids) + len(generated) - 1)[0, -1]
    if sent < len(text):
        yield text[sent:], (time.perf_counter() - started) * 1000


# 3. Constrained triage with a pressure check: how much did the model want any category at all?
@torch.no_grad()
def triage(ticket, min_mass: float = 0.05) -> dict:
    prompt = prompt_for(ticket)
    options = [" " + c for c in CATEGORIES]
    scores = score_options(prompt, options)
    ids = TOK.encode(prompt).ids[-100:]
    probs = torch.softmax(MODEL(torch.tensor([ids]))[0, -1], dim=-1)
    mass = probs[sorted({TOK.encode(o).ids[0] for o in options})].sum().item()
    best = max(scores, key=scores.get).strip()
    return {"category": best, "mass": mass, "route": "auto" if mass >= min_mass else "human review"}


if __name__ == "__main__":
    print("hosted kwargs for each preset:")
    for name, preset in PRESETS.items():
        print(f"  {name:>12}: {to_chat_kwargs(preset)}")

    tickets = load_tickets("dev")[:6]
    print(f"\n{'ticket':>7} | {'category':>14} | {'gold':>14} | {'mass':>7} | {'route':>12} | {'TTFT ms':>7} | "
          f"{'total ms':>8} | draft")
    ttfts, totals, routes = [], [], []
    for t in tickets:
        tri = triage(t)
        draft_prompt = f"Customer (Ana): Hi, {t.subject.lower()}.\nAgent (Lena):"
        events = list(safe_stream(draft_prompt, PRESETS["reply_draft"]))
        draft = "".join(piece for piece, _ in events)
        ttfts.append(events[0][1])
        totals.append(events[-1][1])
        routes.append(tri["route"])
        print(f"{t.id:>7} | {tri['category']:>14} | {t.gold['category']:>14} | {tri['mass']:>7.4f} | {tri['route']:>12} |"
              f" {events[0][1]:>7.1f} | {events[-1][1]:>8.1f} | {draft.strip()[:40]!r}")

    print(f"\nTTFT median {statistics.median(ttfts):.1f} ms, total median {statistics.median(totals):.1f} ms "
          f"(TinyLM on a shared CPU; rerun for your machine)")
    print(f"routed to human review: {routes.count('human review')}/{len(routes)}")
    usage = Usage(input_tokens=450, output_tokens=60 + 600, reasoning_tokens=600)     # medium-effort assumption
    print(f"hosted estimate for the same triage on gpt-oss-120b at medium effort (assumed 600 reasoning tokens): "
          f"{cost_usd(usage, 'openai/gpt-oss-120b'):.6f} USD per ticket")

Code explained

  • In simple words: one script that decodes the Brightlane way: the right settings per task, streaming that behaves, and triage that knows when it is guessing.
  • What happens:
    • PRESETS holds three SamplingParams: greedy triage, reply drafts at temperature 0.7 with top-p 0.9 and a stop at the next customer turn, and brainstorming at temperature 1.0 with min-p 0.1.
    • to_chat_kwargs turns a preset into llm.chat arguments, dropping what the Chat Completions API lacks (top-k, min-p) or Groq rejects (logit bias), trimming stops to the 4 Groq allows, and sending penalties through extra.
    • safe_stream is the streaming loop from Part A with the fix: after each token it computes how many trailing characters could be the beginning of a stop sequence and holds them back. When a stop sequence completes, it flushes only the text before it.
    • triage scores the six categories (Part C) and measures the probability mass on their first tokens; below 5 percent it routes the ticket to human review instead of trusting the forced label.
    • The main block prints the hosted arguments, then for six dev tickets the category, gold label, mass, route, TTFT, total time, and the start of a streamed draft, and ends with a reasoning-cost estimate.
  • Comes out: PYTHONPATH=. python examples/m03_lab.py:

    text
    hosted kwargs for each preset:
            triage: {'temperature': 0, 'max_tokens': 8}
       reply_draft: {'temperature': 0.7, 'max_tokens': 60, 'top_p': 0.9, 'stop': ['\nCustomer'], 'seed': 0}
        brainstorm: {'temperature': 1.0, 'max_tokens': 40, 'stop': ['\n'], 'seed': 0}
    
     ticket |       category |           gold |    mass |        route | TTFT ms | total ms | draft
     T-1001 |        billing |        billing |  0.0007 | human review |     3.0 |     39.1 | 'Sorry about the duplicate charge. Duplic'
     T-1002 |        billing |        billing |  0.0004 | human review |     2.1 |      9.2 | 'Great, thanks.'
     T-1004 |        billing |   cancellation |  0.0005 | human review |     2.0 |      8.5 | 'Great, thanks.'
     T-1005 |        billing |   cancellation |  0.0005 | human review |     2.0 |      8.4 | 'Great, thanks.'
     T-1007 |        billing |        billing |  0.0006 | human review |     2.3 |     61.7 | 'Great, I was charged twice for the Team '
     T-1008 |        billing | account_access |  0.0005 | human review |     2.0 |      8.7 | 'Great, thanks.'
    
    TTFT median 2.0 ms, total median 8.9 ms (TinyLM on a shared CPU; rerun for your machine)
    routed to human review: 6/6
    hosted estimate for the same triage on gpt-oss-120b at medium effort (assumed 600 reasoning tokens): 0.000463 USD per ticket
    

    All six tickets route to human review, which is the correct behavior: TinyLM's mass on any category is under 0.1 percent, so its labels (all billing) are not judgments, and the lab does not pretend otherwise. The streamed drafts contain no Customer text, confirming the hold-back fix (a test checks that the streamed text equals generate's trimmed text). TTFT is about 2 ms and totals vary with draft length, exactly as Part A predicts. When you plug a hosted model in, recalibrate the 5 percent threshold on your own dev set: real models put most of their mass on valid labels, and the useful threshold will be much higher.

Run the tests for the whole module:

bash
cd supportdesk
PYTHONPATH=. python -m pytest -q tests/test_m03_decoding.py

Code explained

  • In simple words: check every mechanism in this module automatically, with TinyLM and no API key.
  • What happens: 14 tests cover the one-hot greedy distribution, filter survivor counts (including min-p staying narrower than top-p at high temperature), seeded repeatability, cache and no-cache token equality, stop trimming, logit-bias bans, the JSON constraint always parsing (five seeds, on a nonsense prompt), safe_stream never showing stop text, the hosted-argument translation, reasoning-token cost math, and batch logit differences staying tiny.
  • Comes out:
text
..............                                                           [100%]
14 passed in 2.28s

Project Milestone

After this module the supportdesk repository contains:

FileWhat it adds
examples/m03_setup.pyShared TinyLM loading and measurement helpers (entropy, survivors, timing spread, token probabilities)
examples/m03_one_step.py, m03_prefill_decode.py, m03_kv_cache.py, m03_stream_sim.pyThe decoding loop, prefill vs decode timings, the KV cache, and token-level streaming
examples/m03_stream_hosted.py, m03_hosted_repeat.pyHosted streaming stats and a repeatability check (need a key)
examples/m03_temperature.py, m03_filters.py, m03_diversity.py, m03_penalties.py, m03_stop_max.py, m03_seeds.pySampling experiments with real TinyLM numbers
examples/m03_constrained.py, m03_json_constrained.py, m03_logit_bias.pyConstrained decoding, a JSON constraint, and logit bias
examples/m03_reasoning_cost.py, m03_reasoning_measure.py, m03_read_trace.pyReasoning cost math, an effort-level harness, and trace reading
examples/m03_lab.pyPresets, to_chat_kwargs, safe_stream, and triage with a pressure check
tests/test_m03_decoding.py14 tests, all runnable offline

Decisions the team now has on record: triage decodes greedily (or with provider structured outputs from Module 6) with a small max_tokens; reply drafts stream, use moderate temperature with a relative filter, and stop at the next turn; finish_reason == "length" is logged and alerted on; nothing depends on hosted outputs repeating exactly; reasoning effort starts low for triage and is raised only when a measurement shows cost per correct answer improves. Maya's team gets drafts that appear within a second, and triage answers that are flagged when the model is guessing.

Interview Questions

1. A user says the assistant "feels slow". Where does the time go, and what do you change first? Split latency into TTFT and inter-token time. TTFT covers network, queueing, prefill, and, on reasoning models, all hidden thinking. Inter-token time times output length is usually the bulk. So first stream the reply, which cuts perceived wait to TTFT. Then shorten outputs (tighter instructions, lower max_tokens), lower reasoning effort if it is a reasoning model, and cache long stable prompt prefixes. In our TinyLM measurement one extra output token cost about 37 times as much time as one extra prompt token, and on GPU-served models decode is similarly the bottleneck.

2. Why does latency scale with output length more than input length? Prefill processes all prompt tokens in one parallel pass, so hardware amortizes the work. Decode must produce tokens one at a time, and each token needs a full model pass (on GPUs, reading all weights from memory). Our run: prefill grew from 1.1 to 3.2 ms as the prompt went from 8 to 96 tokens, while each output token cost about 0.9 ms. Very long prompts do raise TTFT, but for typical support traffic the reply length dominates.

3. What does the KV cache do, and does it change the output? It stores the attention keys and values for tokens already processed so each decode step only computes them for the new token. It does not change the math, only avoids recomputing it: our greedy runs produced identical tokens with and without it, while per-token decode time without the cache grew with length (2.0x to 2.8x slower here). Logits can differ in the last bits (about 2e-6 in our test) because operations happen in a different order.

4. Explain temperature to a product manager. What does it not do? Temperature controls how much the model's less likely options get picked. Low temperature concentrates on the top choice; high temperature spreads probability toward unlikely tokens. It does not reorder the options, add knowledge, or make a confident model creative: on a prompt where TinyLM put 99.9 percent on one token, temperatures up to 1.0 changed nothing, while on an uncertain prompt T=1.5 raised the chance of picking outside the top ten from 0.5 to 10.7 percent.

5. Top-k, top-p, or min-p? Top-k keeps a fixed count, which is wrong for both very certain steps (keeps junk) and very open ones (cuts valid options: it kept 5 of 16 equally valid names). Top-p adapts to the distribution, but when temperature flattens the tail first, it admits hundreds of junk tokens (286 at T=1.5 with p=0.95 in our test). Min-p cuts relative to the top token, so it stayed at 16 names and 9 openings at both temperatures. For creative variety, min-p with a higher temperature; for moderate variety, top-p around 0.9 at moderate temperature; for extraction, neither: decode greedily.

6. You set temperature to 0 and a seed, yet the API returns different answers. Why? Serving systems batch your request with others, and GPU kernels are generally not batch invariant: the numeric result for your request depends on batch size, which depends on server load. Tiny differences (we measured about 2e-6 in TinyLM logits between batch sizes) flip near-tied tokens, and one flip changes the rest of the reply. Thinking Machines found 80 unique completions out of 1,000 at temperature 0 on a Qwen3 235B model, and fixed it with batch-invariant kernels. MoE routing and provider-side updates add more variation. Providers call seeds best effort; design so nothing depends on exact repeats.

7. When would you use frequency or presence penalties, and what can go wrong? To discourage loops and repeated phrases, mostly with small models or greedy decoding. They act on tokens, not meaning, so they also penalize necessary repeats like product names and punctuation. In our test, mild values did nothing, presence 1.0 made greedy find a different loop, and 3.0 removed the loop but spliced unrelated facts together. Some providers (Groq, per its docs) do not support them at all. Prefer fixing the cause: a better model, moderate temperature, or a stop sequence.

8. How does schema-constrained decoding work, and what does it not guarantee? At each step the engine computes which tokens keep the output consistent with the grammar or schema and masks the rest before sampling. It guarantees the output parses. It does not guarantee correctness: TinyLM produced valid categories 48/48 and valid JSON 48/48, yet its accuracy (11/48) was indistinguishable from always guessing the majority class (12/48). Check the probability mass the model gave the allowed tokens; when it is tiny, the constraint is doing all the work and the answer deserves review.

9. Your constrained classifier returns the same label for every input. How do you debug it? Suspect the decoding before the model. Log the model's unconstrained top tokens and the allowed set per step. In TinyLM, greedy decoding with a mask picked the global top token, found it disallowed, fell back to a uniform distribution over allowed tokens, and took argmax, which returns the lowest token id, so every answer became how_to. The fix is to mask the logits first and then take the best allowed token, or to score each option's log-probability directly.

10. When is extended reasoning worth paying for? When the task has multiple dependent steps (policy rules with dates and plans, tool planning, code, math) and a measurement on your eval set shows a better cost per correct answer. It is usually waste for classification and short extraction, and it hurts streaming chat because TTFT includes all thinking. Reasoning tokens bill at the output rate: at our assumed counts, gpt-oss-120b triage went from about 5.80 USD per month at low effort to about 48 USD at high for 30,000 tickets, and from 0.4 to 5 seconds of decode.

11. Can you trust a model's visible reasoning as an explanation of its answer? Not as proof. Turpin et al. showed explanations rationalizing answers driven by a hidden bias, with accuracy drops up to 36 percent; Anthropic found reasoning models mentioned a hint they used only 25 percent (Claude 3.7 Sonnet) and 39 percent (DeepSeek R1) of the time, and unfaithful traces were longer. Use traces to generate debugging hypotheses, then test them with changed inputs; grade answers, not traces, and never show a trace as a customer-facing justification.

12. A reply ends mid-sentence. What do you check? finish_reason (or stop_reason in TinyLM). length means max_tokens was hit; on reasoning models the thinking may have consumed the budget, leaving little or no visible answer. stop means a stop sequence fired, possibly one that appears legitimately inside the answer. Log the reason on every call, alert when the length rate rises, and never parse or send a truncated reply as if complete.

Other Tools and Providers

NeedUsed in this moduleAlternatives
Local model for mechanism demosTinyLM (supportdesk/tinylm.py)Hugging Face Transformers with a small open model (GPT-2, Qwen3 0.6B); llama.cpp; nanoGPT
Hosted inferenceGroq, Gemini, Ollama through llm.chatOpenAI, Anthropic, Together AI, Fireworks, Cerebras, Mistral, AWS Bedrock, Azure AI Foundry, Vertex AI
Self-hosted serving with batching and KV cache managementNot usedvLLM, SGLang, Hugging Face Text Generation Inference, llama.cpp server, NVIDIA TensorRT-LLM
Constrained decodingA token trie in generate(allowed=...)XGrammar, llguidance, Outlines, llama.cpp GBNF, lm-format-enforcer; provider structured outputs (Groq strict mode, Gemini, Ollama, OpenAI, Anthropic)
Deterministic inferenceSeeds on TinyLMBatch-invariant kernels (Thinking Machines' batch_invariant_ops), otherwise, caching results
Latency measurementtime.perf_counter, StreamStatsProvider dashboards, OpenTelemetry tracing (Module 13), Artificial Analysis for published TTFT and throughput comparisons
Reasoning controlreasoning_effort (Groq, Gemini compatibility layer)Gemini thinking_level and thinking_budget, Anthropic extended thinking budgets, Ollama think, OpenAI reasoning.effort

Coming Up in Module 4

You now control how a model turns probabilities into text. Module 4, Prompt Engineering Fundamentals, controls what those probabilities are in the first place: the anatomy of a prompt, few-shot examples and their ordering, formatting that signals structure, and the iteration discipline of building a test set before tuning. You will reuse this module's habits directly: decode triage greedily while you compare prompts, so that a change in accuracy comes from the prompt and not from sampling noise, and always report n and the noise band before calling a prompt better.