Part B: Sampling parameters
Inside apply_sampling
Every sampling setting in this part is a transformation of one probability distribution. Here is the function that applies them, from supportdesk/tinylm.py:
Excerpt from supportdesk/tinylm.py:
def apply_sampling(logits: torch.Tensor, params: SamplingParams, generated: list[int]) -> torch.Tensor:
"""Turn raw logits into a probability distribution using the sampling parameters."""
logits = logits.clone()
for token_id, bias in params.logit_bias.items():
logits[token_id] += bias
if generated and (params.frequency_penalty or params.presence_penalty):
counts = torch.bincount(torch.tensor(generated), minlength=logits.shape[0]).float()
logits -= params.frequency_penalty * counts + params.presence_penalty * (counts > 0).float()
if params.temperature <= 0:
probs = torch.zeros_like(logits)
probs[logits.argmax()] = 1.0
return probs
probs = torch.softmax(logits / params.temperature, dim=-1)
if params.top_k is not None:
kth = probs.topk(params.top_k).values[-1]
probs = torch.where(probs >= kth, probs, torch.zeros_like(probs))
if params.min_p is not None:
probs = torch.where(probs >= params.min_p * probs.max(), probs, torch.zeros_like(probs))
if params.top_p is not None:
sorted_probs, order = probs.sort(descending=True)
keep = sorted_probs.cumsum(0) - sorted_probs < params.top_p
mask = torch.zeros_like(probs, dtype=torch.bool)
mask[order[keep]] = True
probs = torch.where(mask, probs, torch.zeros_like(probs))
return probs / probs.sum()
Code explained
- In simple words: start from the model's scores, nudge some tokens up or down, sharpen or flatten, then cut away unlikely tokens, and hand back probabilities that sum to 1.
- What happens: in order: (1) logit bias adds a fixed number to chosen tokens' scores; (2) frequency and presence penalties subtract from tokens already generated; (3) temperature 0 or below returns a one-hot distribution on the top token (greedy); otherwise scores are divided by the temperature before the softmax; (4) top-k keeps the k most likely tokens; (5) min-p keeps tokens whose probability is at least
min_ptimes the top token's; (6) top-p keeps the smallest set of top tokens whose probability adds up totop_p; (7) the survivors are renormalized. - Comes out: a 2,048-long probability vector. One detail if you combine filters: top-p here measures its cumulative sum against the probabilities left after top-k and min-p, without renormalizing first, so combining filters keeps slightly more tokens than top-p alone would suggest. Real inference engines differ in filter order too (vLLM, llama.cpp, and Hugging Face each document their own), which is one reason the same settings behave a little differently across providers.
Temperature: what it changes and what it does not
Temperature divides every score before the softmax. Below 1 it sharpens the distribution toward the top token; above 1 it flattens it toward uniform; 0 is treated as "always take the top token". It changes how much probability the tail gets. It does not change the ranking of tokens, it does not add knowledge, and it does nothing to a prompt where the model is already certain.
examples/m03_temperature.py:
"""What temperature does to the whole next-token distribution."""
from examples.m03_setup import CERTAIN, UNCERTAIN, distribution, entropy_bits, top_tokens
from supportdesk.tinylm import SamplingParams
for prompt in (UNCERTAIN, CERTAIN):
print(f"\nprompt: {prompt!r}")
print(f"{'T':>4} | {'entropy bits':>12} | {'tail mass':>9} | top 4 tokens")
for t in (0.2, 0.7, 1.0, 1.5):
probs = distribution(prompt, SamplingParams(temperature=t))
tail = 1 - probs.topk(10).values.sum().item() # chance of picking something outside the top 10
print(f"{t:>4} | {entropy_bits(probs):>12.2f} | {tail:>9.4f} | {top_tokens(probs, 4)}")
print(f" 0 | {0.0:>12.2f} | {0.0:>9.4f} | greedy: always the top token")
Code explained
- In simple words: look at the same next-token distribution through four temperature settings, for an uncertain prompt and a certain one.
- What happens: for each temperature,
distributionappliesapply_sampling, then the script prints the entropy, the tail mass (the total chance of picking something outside the top 10 tokens), and the top four tokens. - Comes out:
prompt: 'Customer (Ana): Hi,'
T | entropy bits | tail mass | top 4 tokens
0.2 | 0.50 | 0.0000 | [(' how', 0.924), (' I', 0.054), (' can', 0.009), (' where', 0.004)]
0.7 | 2.77 | 0.0001 | [(' how', 0.367), (' I', 0.163), (' can', 0.099), (' where', 0.078)]
1.0 | 3.06 | 0.0047 | [(' how', 0.27), (' I', 0.153), (' can', 0.108), (' where', 0.091)]
1.5 | 4.34 | 0.1073 | [(' how', 0.184), (' I', 0.126), (' can', 0.1), (' where', 0.089)]
0 | 0.00 | 0.0000 | greedy: always the top token
prompt: 'Customer (Ana): Hi, my'
T | entropy bits | tail mass | top 4 tokens
0.2 | 0.00 | 0.0000 | [(' account', 1.0), (' password', 0.0), (' annual', 0.0), (',', 0.0)]
0.7 | 0.00 | 0.0000 | [(' account', 1.0), (' password', 0.0), (' annual', 0.0), (',', 0.0)]
1.0 | 0.02 | 0.0006 | [(' account', 0.999), (' password', 0.0), (' annual', 0.0), (',', 0.0)]
1.5 | 1.05 | 0.0577 | [(' account', 0.928), (' password', 0.005), (' annual', 0.002), (',', 0.001)]
0 | 0.00 | 0.0000 | greedy: always the top token
On the uncertain prompt, T=0.2 gives how 92 percent: nearly greedy. T=0.7 and 1.0 keep the nine plausible openings but reshape their shares. At T=1.5 the ranking is unchanged, yet the tail mass jumps from 0.5 percent to 10.7 percent: one pick in ten now comes from outside the top ten, which on this model means fragments that never follow a greeting. On the certain prompt, temperatures up to 1.0 change nothing visible, and at 1.5 the top token still keeps 93 percent while 5.8 percent leaks into junk. So temperature is not a creativity dial. It is a "how much do I trust the tail" dial, and its effect depends entirely on how peaked the model already is. That is why the same temperature=0.7 can feel deterministic on an extraction prompt and wild on an open-ended one.
Top-k, top-p, and min-p: cutting the tail
Truncation filters remove unlikely tokens before sampling:
- Top-k keeps a fixed number of tokens, whatever their probabilities.
- Top-p (nucleus sampling, from Holtzman et al., "The Curious Case of Neural Text Degeneration", ICLR 2020) keeps the smallest set of top tokens that together reach probability p.
- Min-p (Nguyen et al., "Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs", ICLR 2025) keeps tokens with at least
min_ptimes the top token's probability, so the cut scales with the model's confidence.
The clearest way to compare them is to count what survives on prompts of different certainty.
examples/m03_filters.py:
"""How many tokens survive top-k, top-p, and min-p on three prompts of different certainty."""
from examples.m03_setup import CERTAIN, UNCERTAIN, distribution, entropy_bits, survivors
from supportdesk.tinylm import SamplingParams
PROMPTS = {"certain": CERTAIN, "uncertain": UNCERTAIN, "flat (agent name)": "Agent ("}
SETTINGS = {
"no filter": {},
"top_k=5": {"top_k": 5},
"top_p=0.5": {"top_p": 0.5},
"top_p=0.9": {"top_p": 0.9},
"top_p=0.95": {"top_p": 0.95},
"min_p=0.05": {"min_p": 0.05},
"min_p=0.5": {"min_p": 0.5},
}
for temperature in (1.0, 1.5):
print(f"\ntemperature {temperature}: tokens that can still be sampled (entropy in bits)")
print(f"{'setting':>11} | " + " | ".join(f"{name:>18}" for name in PROMPTS))
for label, extra in SETTINGS.items():
cells = []
for prompt in PROMPTS.values():
probs = distribution(prompt, SamplingParams(temperature=temperature, **extra))
cells.append(f"{survivors(probs):>6} ({entropy_bits(probs):4.2f} b)")
print(f"{label:>11} | " + " | ".join(f"{c:>18}" for c in cells))
Code explained
- In simple words: apply each filter to three distributions (certain, uncertain, and flat across sixteen names) and count how many tokens could still be chosen.
- What happens:
survivorscounts nonzero probabilities afterapply_sampling; the entropy in brackets shows how spread the survivors are. The whole grid runs at temperature 1.0 and again at 1.5. - Comes out:
temperature 1.0: tokens that can still be sampled (entropy in bits)
setting | certain | uncertain | flat (agent name)
no filter | 2048 (0.02 b) | 2048 (3.06 b) | 2048 (4.05 b)
top_k=5 | 5 (0.01 b) | 5 (2.17 b) | 5 (2.32 b)
top_p=0.5 | 1 (0.00 b) | 3 (1.48 b) | 8 (3.00 b)
top_p=0.9 | 1 (0.00 b) | 8 (2.82 b) | 15 (3.90 b)
top_p=0.95 | 1 (0.00 b) | 9 (2.99 b) | 16 (3.99 b)
min_p=0.05 | 1 (0.00 b) | 9 (2.99 b) | 16 (3.99 b)
min_p=0.5 | 1 (0.00 b) | 2 (0.94 b) | 16 (3.99 b)
temperature 1.5: tokens that can still be sampled (entropy in bits)
setting | certain | uncertain | flat (agent name)
no filter | 2048 (1.05 b) | 2048 (4.34 b) | 2048 (5.00 b)
top_k=5 | 5 (0.10 b) | 5 (2.26 b) | 5 (2.32 b)
top_p=0.5 | 1 (0.00 b) | 4 (1.94 b) | 9 (3.17 b)
top_p=0.9 | 1 (0.00 b) | 17 (3.18 b) | 16 (4.00 b)
top_p=0.95 | 25 (0.26 b) | 253 (3.71 b) | 286 (4.39 b)
min_p=0.05 | 1 (0.00 b) | 9 (3.10 b) | 16 (4.00 b)
min_p=0.5 | 1 (0.00 b) | 3 (1.54 b) | 16 (4.00 b)
Three lessons, each visible in one row:
- Top-k ignores context.
top_k=5keeps 5 tokens on the certain prompt (four of them junk) and also only 5 of the 16 valid agent names on the flat prompt, cutting real options. - Top-p adapts, but temperature fools it. At T=1.0,
top_p=0.95finds exactly the 16 names and exactly the 9 openings. At T=1.5 the flattened tail means 95 percent of the mass now needs 286 tokens on the name prompt and 253 on the uncertain prompt: the filter lets in hundreds of junk tokens because temperature was applied first. - Min-p stays put.
min_p=0.05keeps 1, 9, and 16 tokens at both temperatures, because its cut is relative to the top token, which temperature scales together with everything else.min_p=0.5is aggressive: only the tokens at least half as likely as the favorite survive (2 or 3 openings), while all 16 names survive because they are nearly tied.
- This is the argument the min-p paper makes, reproduced on a 1-million-parameter model: min-p lets you raise temperature for variety without letting the long tail in.
Measuring diversity: 20 samples per setting
Survivor counts describe one step. What matters to a product is the whole output. We sample 20 continuations per setting on the uncertain prompt and measure three things: how many distinct outputs we got, how many are coherent, and how many tail tokens were picked. For coherence we use a check that only works because TinyLM memorized its corpus: a line counts as "seen in corpus" if some customer in the training text wrote exactly that line after a greeting.
examples/m03_diversity.py:
"""Diversity vs coherence: 20 samples per setting on the uncertain prompt."""
from pathlib import Path
from examples.m03_setup import MODEL, TOK, UNCERTAIN, token_probs
from supportdesk.tinylm import SamplingParams, generate
CORPUS = Path("data/corpus.txt").read_text(encoding="utf-8")
N = 20
SETTINGS = {
"greedy (T=0)": {"temperature": 0},
"T=0.2": {"temperature": 0.2},
"T=0.7": {"temperature": 0.7},
"T=1.0": {"temperature": 1.0},
"T=1.0, top_p=0.9": {"temperature": 1.0, "top_p": 0.9},
"T=1.5": {"temperature": 1.5},
"T=1.5, min_p=0.1": {"temperature": 1.5, "min_p": 0.1},
"T=2.0": {"temperature": 2.0},
}
print(f"prompt {UNCERTAIN!r}; {N} samples per setting, seeds 0 to {N - 1}, up to 20 tokens, stop at newline\n")
print(f"{'setting':>17} | {'distinct':>8} | {'seen in corpus':>14} | {'tail picks':>10} | an output not seen in corpus")
for label, extra in SETTINGS.items():
outs = [generate(MODEL, TOK, UNCERTAIN, SamplingParams(max_new_tokens=20, stop=["\n"], seed=s, **extra))
for s in range(N)]
texts = [g.text for g in outs]
# "Seen in corpus": some customer in the training text wrote exactly this line after a greeting.
seen = [f",{t}\n" in CORPUS for t in texts]
# "Tail picks": sampled tokens the raw model (T=1, no filters) gave less than a 1% chance.
tail = sum(p < 0.01 for g in outs for p in token_probs(UNCERTAIN, g.token_ids))
odd = next((t for t, ok in zip(texts, seen) if not ok), "(none)")
print(f"{label:>17} | {len(set(texts)):>5}/{N} | {sum(seen):>11}/{N} | {tail:>10} | {odd[:44]!r}")
Code explained
- In simple words: roll the dice 20 times per setting and count how many different answers you get, how many make sense, and how often the model picked a token it considered unlikely.
- What happens: each setting generates 20 samples with seeds 0 to 19 (so the experiment is repeatable), up to 20 tokens, stopping at a newline. "Distinct" counts unique strings. "Seen in corpus" is the crude coherence check. "Tail picks" re-scores every sampled token under the raw model (temperature 1, no filters) and counts tokens that had less than a 1 percent chance.
- Comes out:
prompt 'Customer (Ana): Hi,'; 20 samples per setting, seeds 0 to 19, up to 20 tokens, stop at newline
setting | distinct | seen in corpus | tail picks | an output not seen in corpus
greedy (T=0) | 1/20 | 20/20 | 0 | '(none)'
T=0.2 | 8/20 | 20/20 | 0 | '(none)'
T=0.7 | 15/20 | 20/20 | 1 | '(none)'
T=1.0 | 13/20 | 20/20 | 1 | '(none)'
T=1.0, top_p=0.9 | 13/20 | 20/20 | 0 | '(none)'
T=1.5 | 18/20 | 10/20 | 80 | ' our automations stopped and there is severa'
T=1.5, min_p=0.1 | 12/20 | 20/20 | 0 | '(none)'
T=2.0 | 20/20 | 0/20 | 272 | 'gon nSAMLDara): Great runs nativeest and att'
Greedy gives 1 distinct output out of 20, as it must. T=0.2 already gives 8: low temperature is not deterministic when the top choices are close. T=0.7 and T=1.0 give 13 to 15 distinct outputs, all coherent. T=1.5 gives 18 distinct outputs, but only 10 of 20 are coherent and the samples picked 80 tail tokens. T=2.0 is pure noise. The interesting row is T=1.5, min_p=0.1: 12 distinct outputs, 20 of 20 coherent, zero tail picks. With n=20 per row, treat differences of two or three distinct outputs (13 vs 15) as noise; the jump from 20/20 to 10/20 coherent is not.
The general rule this supports: to get variety, raise temperature and add a relative filter (min-p, or top-p at moderate temperature); never raise temperature alone.
Frequency and presence penalties
Small models and greedy decoding tend to loop. Two penalties push against repetition by lowering the scores of tokens that already appeared in the generated text:
- Frequency penalty subtracts
penalty x count: the more a token has appeared, the bigger the push. - Presence penalty subtracts a flat
penaltyonce a token has appeared at all.
Export a board is a prompt where greedy TinyLM falls into a loop, so we measure there.
examples/m03_penalties.py:
"""Frequency and presence penalties on a prompt where greedy TinyLM falls into a loop."""
from collections import Counter
from examples.m03_setup import MODEL, TOK
from supportdesk.tinylm import SamplingParams, generate
PROMPT = "Export a board"
def repeated_4grams(ids: list[int]) -> int:
"""How many 4-token windows are exact repeats of an earlier window: a simple loop detector."""
grams = Counter(tuple(ids[i:i + 4]) for i in range(len(ids) - 3))
return sum(n - 1 for n in grams.values())
print(f"prompt {PROMPT!r}, greedy, 80 new tokens\n")
print(f"{'frequency':>9} {'presence':>8} | {'distinct tokens':>15} | {'repeated 4-grams':>16} | tail of the output")
for freq, pres in [(0.0, 0.0), (0.5, 0.0), (1.0, 0.0), (0.0, 1.0), (0.0, 3.0), (3.0, 0.0)]:
g = generate(MODEL, TOK, PROMPT, SamplingParams(max_new_tokens=80, temperature=0,
frequency_penalty=freq, presence_penalty=pres))
print(f"{freq:>9} {pres:>8} | {len(set(g.token_ids)):>15} | {repeated_4grams(g.token_ids):>16} | {g.text[-60:]!r}")
Code explained
- In simple words: generate 80 tokens greedily with different penalty strengths and count repeated 4-token chunks, a simple loop detector.
- What happens:
repeated_4gramscounts 4-token windows that already appeared earlier in the output. Each row changes only the penalties; everything else is greedy and identical. - Comes out:
prompt 'Export a board', greedy, 80 new tokens
frequency presence | distinct tokens | repeated 4-grams | tail of the output
0.0 0.0 | 46 | 11 | 're, how do I cancel my Enterprise subscription?\nAgent (Ana):'
0.5 0.0 | 45 | 11 | ' helps.\n\nCustomer (Lena): Hey there, how do I cancel my Team'
1.0 0.0 | 54 | 9 | 'adata after rotating certificates.\nTwo-factor authentication'
0.0 1.0 | 42 | 19 | 'ello, how do I export my invoice data?\nAgent (Ana): Export a'
0.0 3.0 | 61 | 0 | 'rting a bug, include the browser or app version, include the'
3.0 0.0 | 66 | 0 | 'iew dependencies and labels an incident that breached the 99'
Without a penalty the full output contains are available to workspace and attachments) are available to workspace and attachments) are available under Settings > Billing > Invoices are available...: a loop, then a drift into another article. With frequency penalty 3.0 the loop is gone (0 repeated 4-grams) and the text continues are available to workspace owners from Settings > Data privacy@brightlane.example for live updates every 30 minutes...: no repetition, but it now splices unrelated facts. Moderate settings (0.5 or 1.0) barely change anything, and presence 1.0 actually repeated more (19): once the penalized tokens drop out, greedy finds another loop.
Two cautions. Penalties act on tokens, not ideas, so they also penalize the spaces, punctuation, and product names a good reply must repeat; strong values damage exactly the text you care about. And provider support varies: Groq's API reference currently says frequency_penalty and presence_penalty are "not yet supported by any of our models" (checked 21 September 2026). If your drafts loop, first check whether you are decoding greedily from a small model; a better model or moderate temperature usually fixes it more cleanly.
Stop sequences and max tokens
Stop sequences end generation as soon as a given string appears. Max tokens caps the number of output tokens. Both decide where a reply ends, and the reason it ended is information you should always log.
examples/m03_stop_max.py:
"""Stop sequences and max tokens: where a reply ends, and why it ended."""
from examples.m03_setup import MODEL, TOK
from supportdesk.tinylm import SamplingParams, generate
PROMPT = ("Customer (Ana): Hi, I was charged twice for the Team plan this month. "
"Can you refund the duplicate?\nAgent (Lena):")
cases = {
"no stop, 60 tokens": SamplingParams(max_new_tokens=60, temperature=0),
"stop at next turn": SamplingParams(max_new_tokens=60, temperature=0, stop=["\nCustomer"]),
"stop, but only 12 tokens": SamplingParams(max_new_tokens=12, temperature=0, stop=["\nCustomer"]),
}
for label, params in cases.items():
g = generate(MODEL, TOK, PROMPT, params)
print(f"--- {label}: stop_reason={g.stop_reason}, tokens generated={len(g.token_ids)}, "
f"characters returned={len(g.text)}")
print(g.text)
print("--- raw decoded tokens end with:", repr(TOK.decode(g.token_ids)[-20:]))
Code explained
- In simple words: let TinyLM draft an agent reply three ways: with no stop, stopping at the next customer turn, and with the stop but a tiny token cap.
- What happens: each case prints
stop_reason, how many tokens were generated, and how many characters came back, then the tail of the raw decoded tokens so you can see what was trimmed. - Comes out:
--- no stop, 60 tokens: stop_reason=length, tokens generated=60, characters returned=273
Sorry about the duplicate charge. Duplicate charges are always refunded in full within 5 to 10 business days to the original payment method.
Customer (Omar): Great, thanks.
Customer (Eli): Good morning, how do I cancel my Team subscription?
Agent (Eli): You can cancel at
--- raw decoded tokens end with: '): You can cancel at'
--- stop at next turn: stop_reason=stop, tokens generated=27, characters returned=141
Sorry about the duplicate charge. Duplicate charges are always refunded in full within 5 to 10 business days to the original payment method.
--- raw decoded tokens end with: 'ent method.\nCustomer'
--- stop, but only 12 tokens: stop_reason=length, tokens generated=12, characters returned=75
Sorry about the duplicate charge. Duplicate charges are always refunded in
--- raw decoded tokens end with: 'e always refunded in'
Without a stop sequence the model writes the agent reply and then keeps going: it invents the customer's thanks and a whole new conversation, and runs out of tokens mid-sentence (stop_reason=length). With stop=["\nCustomer"] it ends after 27 tokens with a clean reply; the raw tokens end in \nCustomer, which was generated (and on a hosted API, billed) but trimmed from the text. With a 12-token cap the reply is cut mid-sentence and stop_reason is length.
In production, finish_reason == "length" on a hosted call (the same idea as stop_reason here; llm.chat returns it in ChatResult.finish_reason) means the answer is truncated. For JSON outputs it means the JSON is probably broken. Log it, alert on its rate, and never treat a length-truncated reply as complete. Groq accepts up to 4 stop sequences.
Seeds, determinism, and why identical inputs still vary
A seed initializes the random number generator, so the same seed replays the same dice rolls. On TinyLM that makes sampling perfectly repeatable. Hosted models are a different story, and the reasons are worth understanding because they surprise everyone who sets temperature=0 and expects identical outputs.
examples/m03_seeds.py:
"""Seeds, determinism, and the tiny numeric differences that make hosted models vary."""
import torch
from examples.m03_setup import MODEL, TOK, UNCERTAIN
from supportdesk.tinylm import SamplingParams, generate
# 1. A seed fixes the random draws, so the same seed gives the same sample.
def sample(seed):
p = SamplingParams(max_new_tokens=12, temperature=1.0, seed=seed, stop=["\n"])
return generate(MODEL, TOK, UNCERTAIN, p).text
print("seed 7, twice: ", [sample(7), sample(7)])
print("seeds 1, 2, 3: ", [sample(s) for s in (1, 2, 3)])
# 2. Floating point addition is not associative: grouping changes the last bits.
print("\n(0.1 + 0.2) + 0.3 =", (0.1 + 0.2) + 0.3, " 0.1 + (0.2 + 0.3) =", 0.1 + (0.2 + 0.3))
values = torch.randn(100_000, generator=torch.Generator().manual_seed(0), dtype=torch.float32)
in_order = values.sum()
shuffled = values[torch.randperm(len(values), generator=torch.Generator().manual_seed(1))].sum()
chunked = torch.stack([c.sum() for c in values.split(4096)]).sum()
print(f"float32 sum, 3 orders: {in_order.item():.6f} {shuffled.item():.6f} {chunked.item():.6f}")
# 3. The same prompt computed alone or inside a batch of other requests.
corpus = open("data/corpus.txt", encoding="utf-8").read()
ids = TOK.encode(UNCERTAIN).ids
n = len(ids)
others = [TOK.encode(corpus[i * 500:]).ids[:n] for i in range(1, 16)] # 15 other same-length "requests"
with torch.no_grad():
alone = MODEL(torch.tensor([ids]))[0, -1]
for batch_size in (2, 4, 16):
batch = torch.tensor([ids] + others[: batch_size - 1])
in_batch = MODEL(batch)[0, -1]
diff = (alone - in_batch).abs().max().item()
print(f"batch of {batch_size:>2}: max logit difference vs alone = {diff:.2e}, "
f"same top token: {alone.argmax().item() == in_batch.argmax().item()}")
# 4. KV cache vs full recompute: a different order of operations for the same math.
caches = [dict() for _ in MODEL.blocks]
MODEL(torch.tensor([ids[:-1]]), caches)
cached_last = MODEL(torch.tensor([ids[-1:]]), caches, start_pos=n - 1)[0, -1]
print(f"cached vs recomputed: max logit difference = {(alone - cached_last).abs().max().item():.2e}")
Code explained
- In simple words: show that seeds make TinyLM repeatable, then show the tiny arithmetic differences that break repeatability on real serving systems.
- What happens: (1) Two samples with seed 7 and three with seeds 1, 2, 3. (2) Floating point addition in a different grouping gives a different last digit; a sum of 100,000 float32 numbers in three orders gives three slightly different totals. (3) The same prompt is run through TinyLM alone and inside batches of 2, 4, and 16 other same-length requests, and the logits are compared. (4) The same logits are computed with and without the KV cache.
- Comes out:
seed 7, twice: [' can I get a refund on my annual Free plan?', ' can I get a refund on my annual Free plan?']
seeds 1, 2, 3: [' does the Business plan include SSO?', ' slack notifications stopped posting.', ' our automations stopped and there is a banner about a limit.']
(0.1 + 0.2) + 0.3 = 0.6000000000000001 0.1 + (0.2 + 0.3) = 0.6
float32 sum, 3 orders: -244.559769 -244.559677 -244.559769
batch of 2: max logit difference vs alone = 0.00e+00, same top token: True
batch of 4: max logit difference vs alone = 1.91e-06, same top token: True
batch of 16: max logit difference vs alone = 1.91e-06, same top token: True
cached vs recomputed: max logit difference = 1.91e-06
The seed replays exactly. The arithmetic differs in the sixth significant digit depending on order. And the headline: the same prompt's logits change by about 2e-6 depending on how many other requests share the batch, and again depending on whether the cache was used. The top token did not flip here, but a real model decodes hundreds of tokens, and one step where two tokens are nearly tied is enough to send the rest of the reply down a different path.
That is exactly what happens on hosted APIs. Thinking Machines Lab's "Defeating Nondeterminism in LLM Inference" (Horace He and colleagues, September 2025) ran Qwen/Qwen3-235B-A22B-Instruct-2507 1,000 times at temperature 0 and got 80 unique completions; they first diverged at token 103, where 992 said "Queens, New York" and 8 said "New York City". Their analysis: the common explanation ("concurrency plus floating point") is incomplete; individual GPU kernels are run-to-run deterministic, but they are not batch invariant: the result for your request depends on how many other requests the server batched with it, and server load changes that from moment to moment. After they wrote batch-invariant kernels for matrix multiplication, attention, and normalization, all 1,000 completions were identical, at some cost in speed. Our TinyLM batch test is the same effect in miniature.
Mixture-of-experts models add another source. In an MoE model (gpt-oss-120b routes each token to 4 of 128 experts), a tiny numeric difference can change which experts a token is routed to, and some serving systems cap how many tokens each expert handles per batch, so your token's routing can depend on other users' tokens. Provider-side changes (new hardware, new kernels, a quietly updated model) add slower drift. This is why providers describe seeds as best effort: Groq's API reference says that with a seed "our system will make a best effort to sample deterministically", and Groq converts temperature=0 to 1e-8.
To see this on your provider, run:
examples/m03_hosted_repeat.py:
"""Send the identical request N times to a hosted model and count distinct replies."""
import os
import sys
from collections import Counter
from supportdesk.llm import PROVIDERS, chat, resolve
N = 10
provider, model = resolve()
key_env = PROVIDERS[provider]["key_env"]
if key_env and not os.environ.get(key_env):
sys.exit(f"Set {key_env} (or LLM_PROVIDER=ollama) to run this example against {provider}.")
messages = [{"role": "user", "content": "In two sentences, explain to a Brightlane customer why a CSV export "
"might not include comments."}]
for label, kwargs in {"temperature=0, seed=42": {"temperature": 0.0, "seed": 42},
"temperature=1.0, seed=42": {"temperature": 1.0, "seed": 42},
"temperature=1.0, no seed": {"temperature": 1.0}}.items():
replies = [chat(messages, max_tokens=300, **kwargs).text.strip() for _ in range(N)]
counts = Counter(replies)
print(f"{label:>26}: {len(counts)} distinct out of {N}; most common appears {counts.most_common(1)[0][1]} times")
Code explained
- In simple words: send the identical request ten times per setting and count how many different answers come back.
- What happens: three settings: temperature 0 with a seed, temperature 1 with a seed, temperature 1 without.
Countercounts identical replies. - Comes out: without a key this build prints the missing-key message. Illustrative sample run (not captured in this build; produced for teaching). Your output will differ.
temperature=0, seed=42: 3 distinct out of 10; most common appears 7 times
temperature=1.0, seed=42: 6 distinct out of 10; most common appears 3 times
temperature=1.0, no seed: 10 distinct out of 10; most common appears 1 times
Whatever your numbers, the design conclusion is the same: never build a system that depends on the model repeating itself exactly. Cache the answer you got if you need it again, compare outputs by meaning rather than string equality in tests, and run evaluations with several samples (Module 10).
Choosing parameters by task
Here is how the pieces combine for the Brightlane assistant. These are starting points to test against your eval set, not laws.
| Situation | Use this | Why |
|---|---|---|
| Ticket triage, extraction, classification | Temperature 0 (or the provider's lowest), low max_tokens, schema constraint (Part C) | You want the single most likely answer; variety is a bug |
| Code or SQL generation | Temperature 0 to 0.3, stop sequences at the end of the code block | Small deviations break syntax; tests catch the rest |
| Customer reply drafts | Temperature 0.5 to 0.8 with top-p 0.9, or temperature 1.0 with min-p 0.05 to 0.1 | Natural wording without tail tokens; the agent reviews anyway |
| Brainstorming (macro ideas, FAQ titles) | Temperature 1.0 or higher with min-p 0.1, several samples | Variety is the goal; min-p keeps it coherent (our 20-sample run) |
| Self-consistency voting (Module 5) | Temperature about 0.7 to 1.0, 5 or more samples | Voting needs genuinely different reasoning paths |
| Gemini 3 models | Leave temperature at its default of 1.0 (temperature=None in llm.chat) | Google "strongly recommend[s] keeping the temperature parameter at its default value of 1.0" for Gemini 3, citing looping and degraded performance otherwise |
| Reasoning models on Groq | Temperature in the provider's recommended range; control depth with reasoning_effort instead | Groq's reasoning docs say reasoning models perform best around 0.5 to 0.7 |
| Tests that must be repeatable | Do not rely on seeds; use ScriptedLLM for plumbing, and multiple samples for quality | Hosted determinism is best effort (batch invariance, MoE routing) |
Note that llm.chat defaults to temperature=0.0. That is a sensible default for this course's extraction tasks on Groq, and the wrong one for Gemini 3: pass temperature=None there.
Provider support differs more than most people expect. As checked on 21 September 2026:
| Parameter | TinyLM | Groq (Chat Completions) | Gemini (OpenAI-compatible endpoint) | Ollama |
|---|---|---|---|---|
| temperature, top_p, max_tokens, stop | Yes | Yes (stop: up to 4) | Yes | Yes |
| top_k, min_p | Yes | Not in the Chat Completions API | top_k in the native API | Yes, as model options |
| frequency / presence penalty | Yes | Documented as "not yet supported by any of our models" | In the native API | Yes |
| seed | Yes, exact | Best effort | In the native API; verify on the compatibility layer | Yes |
| logit_bias | Yes | Returns an error (400) | Not documented for the compatibility layer | Not in the OpenAI-compatible endpoint |
The Gemini compatibility layer "silently ignore[s]" parameters it does not support, so a setting you send may have no effect without any error. Always verify a parameter changes behavior before you rely on it.