Part B: Self-hosting
Part B: Self-hosting
Self-hosting means you run the model yourself: you pick a GPU, load open weights into its memory, and run a serving engine (software that turns a model file into an HTTP endpoint that many users can call at once). No open-weight model can be downloaded in this course's sandbox, so every mechanism below is measured for real on TinyLM, and every number about real models is computed from their published configuration files, not guessed.
A note on timings before we start. This module was written on a 2-core CPU that other jobs were sharing. Every timing script sets torch.set_num_threads(1) and prints the machine load next to its numbers; the "threads" diagnosis later in this part shows why. Timings on your machine will differ; the shapes of the curves are what transfer.
Hardware sizing and memory arithmetic
A GPU has a fixed amount of fast memory (VRAM: 24 GB on an NVIDIA L4, 80 GB on an A100 80GB or H100). Serving a model needs three things to fit in it:
- Weights: parameters times bytes per parameter. At bf16 (16-bit floats, the usual serving precision) that is 2 bytes each, so an 8-billion-parameter model needs about 16 GB before it has answered anything.
- KV cache: during generation, every layer stores a key vector and a value vector for every token of every active conversation, so the model does not recompute them (Module 3). Per token that is
2 (K and V) x layers x KV heads x head size x bytes. Many modern models use grouped-query attention (GQA): several query heads share one key/value head, which shrinks this cache. Llama 3.1 8B has 32 query heads but only 8 KV heads. - Overhead: activations, the CUDA runtime, and the engine's own bookkeeping. We assume 2 GiB; measure yours.
Whatever is left after weights and overhead is KV cache, and KV cache is what decides how many users you can serve at once and how long their conversations can be. That is the real sizing question.
The script below checks the formulas on TinyLM, where we can count the real tensors, then applies them to three real open models using the values in their published config.json files on Hugging Face (meta-llama/Llama-3.1-8B, Qwen/Qwen3-8B, and openai/gpt-oss-120b, read 21 September 2026).
examples/m13_memory.py
"""Memory arithmetic for self-hosting: weights plus KV cache, from a model's config.
TinyLM is computed exactly and checked against the real tensors. Real open
models are computed from their published config.json files (values copied
below, with the source named) so you can size a GPU before renting one.
"""
from __future__ import annotations
from dataclasses import dataclass
import torch
from supportdesk.tinylm import load
GIB = 1024 ** 3
BYTES = {"fp32": 4.0, "bf16": 2.0, "int8": 1.0, "int4": 0.5, "mxfp4": 4.25 / 8} # MXFP4: 4 bits + 1 shared 8-bit scale per 32 values
@dataclass(frozen=True)
class Arch:
"""The config.json fields that decide memory. Values come from each model's published config."""
name: str
hidden: int
layers: int
heads: int
kv_heads: int
head_dim: int
intermediate: int
vocab: int
tied: bool
experts: int = 1 # mixture-of-experts: experts per layer (1 = dense)
active_experts: int = 1 # experts each token actually uses
full_attention_layers: int | None = None # others use a sliding window
window: int | None = None
# Sources: config.json on Hugging Face for meta-llama/Llama-3.1-8B, Qwen/Qwen3-8B, and
# openai/gpt-oss-120b (read 21 Sep 2026). gpt-oss-120b alternates sliding-window (128 tokens)
# and full-attention layers, and routes each token to 4 of its 128 experts.
LLAMA_31_8B = Arch("Llama 3.1 8B", 4096, 32, 32, 8, 128, 14336, 128256, False)
QWEN3_8B = Arch("Qwen3-8B", 4096, 36, 32, 8, 128, 12288, 151936, False)
GPT_OSS_120B = Arch("gpt-oss-120b", 2880, 36, 64, 8, 64, 2880, 201088, False, experts=128,
active_experts=4, full_attention_layers=18, window=128)
def params(a: Arch) -> dict[str, int]:
"""Parameter count by part, for a Llama-style block (gated MLP, grouped-query attention).
Small terms (norm weights, attention biases, router, Qwen3's q/k norms) are
left out; they change the total by far less than 0.1 percent.
"""
embed = a.vocab * a.hidden * (1 if a.tied else 2)
attn = a.hidden * a.heads * a.head_dim * 2 + a.hidden * a.kv_heads * a.head_dim * 2
mlp = 3 * a.hidden * a.intermediate * a.experts
active = embed // (1 if a.tied else 2) + (attn + 3 * a.hidden * a.intermediate * a.active_experts) * a.layers
return {"embeddings": embed, "attention": attn * a.layers, "mlp": mlp * a.layers,
"total": embed + (attn + mlp) * a.layers, "active": active}
def kv_bytes_per_token(a: Arch, dtype: str = "bf16") -> float:
"""K and V, for every layer, for every KV head: 2 * layers * kv_heads * head_dim * bytes."""
layers = a.full_attention_layers if a.full_attention_layers is not None else a.layers
return 2 * layers * a.kv_heads * a.head_dim * BYTES[dtype]
def kv_bytes_for_sequence(a: Arch, tokens: int, dtype: str = "bf16") -> float:
full = kv_bytes_per_token(a, dtype) * tokens
if a.window is not None: # sliding-window layers keep at most `window` tokens
sliding_layers = a.layers - (a.full_attention_layers or 0)
full += 2 * sliding_layers * a.kv_heads * a.head_dim * BYTES[dtype] * min(tokens, a.window)
return full
def tinylm_check() -> None:
model, tokenizer = load()
cfg = model.cfg
d, v, c, n = cfg.d_model, cfg.vocab_size, cfg.context, cfg.n_layers
per_layer = (2 * d) + (d * 3 * d + 3 * d) + (d * d + d) + (2 * d) + (d * 4 * d + 4 * d) + (4 * d * d + d)
by_hand = v * d + c * d + n * per_layer + 2 * d # head is tied to tok_emb, so it adds nothing
weights_fp32 = sum(p.numel() * p.element_size() for p in model.parameters())
print(f"TinyLM parameters: by hand {by_hand:,} model.num_parameters() {model.num_parameters():,}")
print(f"TinyLM weights in fp32: {weights_fp32:,} bytes = {weights_fp32 / 1024**2:.2f} MiB")
used = tokenizer.get_vocab_size()
print(f"embedding rows: {v:,} in the config, {used:,} used by the tokenizer, "
f"{(v - used) * d * 4:,} bytes of rows no token can reach")
kv_formula = 2 * n * cfg.n_heads * (d // cfg.n_heads) * 4 # TinyLM has no grouped-query attention
ids = tokenizer.encode("Customer (Ben): I was charged twice for the Team plan.\nAgent (Dara):").ids
caches = [dict() for _ in model.blocks]
with torch.no_grad():
model(torch.tensor([ids]), caches)
measured = sum(cache["k"].nbytes + cache["v"].nbytes for cache in caches)
print(f"TinyLM KV cache: formula {kv_formula:,} bytes/token; measured {measured:,} bytes "
f"for {len(ids)} tokens = {measured // len(ids):,} bytes/token")
print(f"TinyLM full context ({c} tokens): {kv_formula * c / 1024:.0f} KiB per sequence")
def real_models() -> None:
print(f"\n{'model':<14}{'params (B)':>11}{'bf16 GiB':>10}{'int4 GiB':>10}{'KV KiB/tok':>12}{'KV GiB @8k':>12}{'KV GiB @128k':>14}")
for a in (LLAMA_31_8B, QWEN3_8B):
p = params(a)["total"]
print(f"{a.name:<14}{p / 1e9:>11.2f}{p * 2 / GIB:>10.1f}{p * 0.5 / GIB:>10.1f}"
f"{kv_bytes_per_token(a) / 1024:>12.0f}{kv_bytes_for_sequence(a, 8192) / GIB:>12.2f}"
f"{kv_bytes_for_sequence(a, 131072) / GIB:>14.2f}")
a = GPT_OSS_120B
p = params(a)
experts, rest = p["mlp"], p["total"] - p["mlp"]
size = experts * BYTES["mxfp4"] + rest * BYTES["bf16"]
print(f"{a.name:<14}{p['total'] / 1e9:>11.2f}{p['total'] * 2 / GIB:>10.1f}{'':>10}"
f"{kv_bytes_per_token(a) / 1024:>12.0f}{kv_bytes_for_sequence(a, 8192) / GIB:>12.2f}"
f"{kv_bytes_for_sequence(a, 131072) / GIB:>14.2f}")
print(f"gpt-oss-120b as shipped (experts MXFP4, rest bf16): {size / GIB:.1f} GiB of weights")
print(f"gpt-oss-120b parameters used per token (4 of 128 experts): {p['active'] / 1e9:.2f} B "
f"(counting the {a.vocab * a.hidden / 1e9:.2f} B output head, not the input lookup table)")
def fits(a: Arch, gpu_gib: float, weight_dtype: str, context: int, overhead_gib: float = 2.0) -> int:
"""How many sequences of `context` tokens fit next to the weights (overhead: activations, runtime)."""
weights = params(a)["total"] * BYTES[weight_dtype] / GIB
free = gpu_gib - weights - overhead_gib
return max(int(free // (kv_bytes_for_sequence(a, context) / GIB)), 0)
if __name__ == "__main__":
tinylm_check()
real_models()
print("\nConcurrent sequences that fit (2 GiB runtime overhead assumed):")
for gpu, gib in (("L4 24 GB", 22.5), ("A100 80 GB", 80.0)):
for dtype in ("bf16", "int4"):
print(f" Qwen3-8B on {gpu:<10} weights {dtype}: "
f"{fits(QWEN3_8B, gib, dtype, 4096):>4} at 4k tokens, {fits(QWEN3_8B, gib, dtype, 32768):>3} at 32k")
Code explained
- In simple words: a memory calculator that proves itself on a model it can open, then sizes models it cannot download.
- What happens:
BYTES: bytes per parameter for common formats. MXFP4 is the 4-bit format gpt-oss ships its expert weights in: blocks of 32 four-bit values share one 8-bit scale, so it costs 4.25 bits per value.Arch: the handful of config fields that decide memory: hidden size, layers, attention heads, KV heads, head size, MLP width, vocabulary, whether input and output embeddings are shared (tied), and for mixture-of-experts models how many experts exist and how many each token uses. gpt-oss-120b also alternates sliding-window layers (which only keep the last 128 tokens) with full-attention layers.params(a): counts parameters for a Llama-style block (attention projections plus a gated MLP with three matrices).activecounts what one token actually touches in a mixture-of-experts (MoE) model, where a small router picks a few experts per token.kv_bytes_per_tokenandkv_bytes_for_sequence: the KV formula; sliding-window layers are capped at the window.tinylm_check(): counts TinyLM's parameters by hand, then runs a real prompt and adds up the bytes in the actual cache tensors. It also compares the embedding table (2,048 rows in the config) with the tokenizer (1,503 real tokens).fits(...): how many sequences of a given length fit beside the weights.
- Comes out:text
TinyLM parameters: by hand 1,071,872 model.num_parameters() 1,071,872 TinyLM weights in fp32: 4,287,488 bytes = 4.09 MiB embedding rows: 2,048 in the config, 1,503 used by the tokenizer, 279,040 bytes of rows no token can reach TinyLM KV cache: formula 4,096 bytes/token; measured 73,728 bytes for 18 tokens = 4,096 bytes/token TinyLM full context (128 tokens): 512 KiB per sequence model params (B) bf16 GiB int4 GiB KV KiB/tok KV GiB @8k KV GiB @128k Llama 3.1 8B 8.03 15.0 3.7 128 1.00 16.00 Qwen3-8B 8.19 15.3 3.8 144 1.12 18.00 gpt-oss-120b 116.78 217.5 36 0.29 4.50 gpt-oss-120b as shipped (experts MXFP4, rest bf16): 60.7 GiB of weights gpt-oss-120b parameters used per token (4 of 128 experts): 5.12 B (counting the 0.58 B output head, not the input lookup table) Concurrent sequences that fit (2 GiB runtime overhead assumed): Qwen3-8B on L4 24 GB weights bf16: 9 at 4k tokens, 1 at 32k Qwen3-8B on L4 24 GB weights int4: 29 at 4k tokens, 3 at 32k Qwen3-8B on A100 80 GB weights bf16: 111 at 4k tokens, 13 at 32k Qwen3-8B on A100 80 GB weights int4: 131 at 4k tokens, 16 at 32kRead it from the top. The by-hand count matches
num_parameters()exactly, and the measured cache is exactly 4,096 bytes per token, so the formula is right. TinyLM also carries 545 embedding rows that no token id can reach (the tokenizer learned only 1,503 tokens): 279 KB of dead weight. Real models pad vocabularies too, usually on purpose, to a multiple that suits GPU kernels; it is harmless but it does count toward memory.The real-model rows match their published sizes (Llama 3.1 8B is 8.03 B parameters, Qwen3-8B is 8.19 B), and two lessons stand out. First, the KV cache is not small: one 128k-token conversation on Llama 3.1 8B needs 16 GiB of cache, more than the weights. (Qwen3-8B's native context is 32,768 tokens and reaches 131,072 only with YaRN scaling, so its 128k column is a what-if.) Second, sliding windows and MoE change everything: gpt-oss-120b has 117 B parameters, but its cache is 36 KiB per token and each token uses only 5.1 B parameters, which matches the 5.1 B "active parameters" OpenAI publishes. That is why a 117 B model can generate quickly on one 80 GB GPU: memory is set by total parameters, speed mostly by active ones.
The last block is the sizing answer. On a 24 GB L4, Qwen3-8B in bf16 serves 9 conversations of 4k tokens at once, but only one of 32k. Quantizing the weights to 4 bits frees memory for 29. Nothing here depends on the GPU's speed; that comes next.
| Situation | Use this | Why |
|---|---|---|
| Model under about 10 B parameters, short contexts, modest traffic | One 24 GB GPU (L4 class), 4-bit or 8-bit weights | Weights fit with room left for KV cache |
| 8 B model with long contexts or many concurrent users | An 80 GB GPU | KV cache, not weights, becomes the limit |
| Large MoE model such as gpt-oss-120b | An 80 GB GPU with the shipped 4-bit experts | Total parameters decide memory; active parameters decide speed |
| 70 B dense model or larger | Several GPUs with tensor parallelism, or a hosted API | Weights alone exceed one GPU |
Serving engines and what they optimize
You do not serve a model with a for loop over generate(). Serving engines exist because naive generation wastes the GPU in two ways: memory for the KV cache gets reserved for the longest possible reply and mostly sits empty, and one user's request occupies the whole GPU while it waits for the next token.
The main ideas, which every engine below implements in some form:
- Continuous batching: the engine runs one decode step for every active request together, and new requests join the batch between steps instead of waiting for the whole batch to finish.
- Paged KV cache: cache memory is handed out in small fixed-size blocks, like pages in an operating system, so it is used as needed rather than reserved up front. This is the PagedAttention idea from the vLLM paper (Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention", SOSP 2023).
- Prefix caching: requests that start with the same tokens (a system prompt, a shared help-center excerpt) reuse the stored KV cache for that prefix instead of recomputing it. You will measure this on TinyLM in Part C.
- Quantization and speculative decoding: smaller weights, and a small draft model that proposes tokens the big model checks in one pass.
| Engine | What it optimizes | Where it fits |
|---|---|---|
| vLLM | Throughput on GPUs: paged KV cache, continuous batching, automatic prefix caching, OpenAI-compatible server | The default choice for serving open models on NVIDIA or AMD GPUs |
| SGLang | Prefix reuse (its RadixAttention keeps a tree of cached prefixes), a low-overhead scheduler, prefill-decode disaggregation, speculative decoding, FP4/FP8/INT4 quantization | Workloads with heavy shared prefixes: agents, multi-turn chat, structured generation |
| TensorRT-LLM | Peak speed on NVIDIA hardware by compiling optimized kernels, in-flight batching, FP8 and FP4 | Large, stable deployments on NVIDIA GPUs where a build step is acceptable |
| llama.cpp | Running quantized models (GGUF format) on CPUs, Apple Silicon, and consumer GPUs | Laptops, edge devices, single-user or low-traffic servers |
| Ollama | A friendly local wrapper: model library, one-command downloads, OpenAI-compatible endpoint | Development and the course's local provider (llm.py talks to it on port 11434) |
One engine has left the table: Hugging Face's text-generation-inference (TGI) README now says it "is now in maintenance mode" and points users to vLLM, SGLang, and local engines such as llama.cpp and MLX (TGI on GitHub). If a tutorial you are following uses TGI, it still works, but new projects should pick one of the others. The SGLang feature list above is from its README (sgl-project/sglang).
| Situation | Use this | Why |
|---|---|---|
| First production deployment of an open model on GPUs | vLLM | Widest model support, OpenAI-compatible API, good defaults |
| Most requests share long prefixes (agents, RAG with fixed context) | SGLang | Prefix reuse is its central design |
| Fixed model, very high volume, NVIDIA only, team can own a build pipeline | TensorRT-LLM | Highest speed per GPU at the cost of flexibility |
| No GPU, or a single user | llama.cpp or Ollama | Quantized CPU and Apple Silicon inference |
Quantization and the quality it costs
Quantization stores weights with fewer bits: int8 (1 byte) or 4-bit formats instead of bf16 (2 bytes) or fp32 (4 bytes). It halves or quarters memory, which (as the table above showed) turns into more concurrent users. The question is always what it costs in quality, and the only honest answer is a measurement on your own task.
PyTorch's dynamic int8 quantization is a real, one-line version: it stores every nn.Linear weight as int8 with a scale, and converts activations on the fly. The script applies it to TinyLM and measures size, speed, perplexity on the held-out 5 percent of the corpus (the same 8,244 tokens pretraining held out), how often the top prediction changes, and whether greedy replies change.
examples/m13_quantize.py
"""Quantize TinyLM's linear layers to int8 and measure what it saves and what it costs.
Measures, on this machine: size on disk, forward-pass latency, validation
perplexity, agreement of the top next-token prediction, and greedy generations.
"""
from __future__ import annotations
import copy
import io
import math
import os
import statistics
import time
import warnings
import torch
from supportdesk.data import DATA_DIR
from supportdesk.tinylm import SamplingParams, generate, load, loss_on
warnings.filterwarnings("ignore") # torch.ao.quantization prints deprecation notices
torch.manual_seed(0)
torch.set_num_threads(1) # one thread: steadier timings on a shared 2-core machine
def validation_ids(tokenizer) -> torch.Tensor:
"""The same held-out tail (last 5 percent of tokens) that pretraining used for validation."""
ids = torch.tensor(tokenizer.encode((DATA_DIR / "corpus.txt").read_text(encoding="utf-8")).ids)
return ids[int(len(ids) * 0.95):]
@torch.no_grad()
def evaluate(model, ids: torch.Tensor, context: int = 128) -> tuple[float, list[int]]:
"""Perplexity over non-overlapping windows, plus the argmax prediction at every position."""
total, count, predictions = 0.0, 0, []
for start in range(0, len(ids) - 1, context):
chunk = ids[start: start + context + 1]
if len(chunk) < 2:
break
x, y = chunk[:-1].unsqueeze(0), chunk[1:].unsqueeze(0)
total += loss_on(model, x, y).item() * y.numel()
count += y.numel()
predictions += model(x)[0].argmax(-1).tolist()
return math.exp(total / count), predictions
def size_on_disk(model) -> int:
buffer = io.BytesIO()
torch.save(model.state_dict(), buffer)
return buffer.getbuffer().nbytes
@torch.no_grad()
def forward_ms(models: dict, batch: torch.Tensor, repeats: int = 30) -> dict[str, float]:
"""Median forward time per model. Runs alternate between models so background load hits both."""
times: dict[str, list[float]] = {name: [] for name in models}
for i in range(repeats + 3):
for name, model in models.items():
started = time.perf_counter()
model(batch)
if i >= 3: # the first rounds are warm-up
times[name].append((time.perf_counter() - started) * 1000)
return {name: statistics.median(t) for name, t in times.items()}
def main() -> None:
print(f"machine load (1-minute average) {os.getloadavg()[0]:.2f} on {os.cpu_count()} cores; "
f"torch threads {torch.get_num_threads()}")
fp32, tokenizer = load()
int8 = torch.ao.quantization.quantize_dynamic(copy.deepcopy(fp32), {torch.nn.Linear}, dtype=torch.qint8)
ids = validation_ids(tokenizer)
batch = ids[:128].unsqueeze(0)
models = {"fp32": fp32, "int8": int8}
speed = forward_ms(models, batch)
results = {}
for name, model in models.items():
ppl, preds = evaluate(model, ids)
results[name] = preds
print(f"{name}: size {size_on_disk(model) / 1024:5.0f} KiB forward (1 x 128 tokens) "
f"{speed[name]:6.2f} ms validation perplexity {ppl:.4f}")
same = sum(a == b for a, b in zip(results["fp32"], results["int8"]))
print(f"top-1 next-token agreement: {same} of {len(results['fp32'])} positions "
f"({same / len(results['fp32']):.2%})")
prompts = ["Customer (Ben): How much does the Team plan cost?\nAgent (Dara):",
"Customer (Mo): I forgot my password and the reset link expired.\nAgent (Lena):",
"Customer (Ana): Can I export my tasks to CSV?\nAgent (Kofi):"]
params = SamplingParams(max_new_tokens=30, temperature=0.0, stop=["\n"])
for prompt in prompts:
a = generate(fp32, tokenizer, prompt, params).text
b = generate(int8, tokenizer, prompt, params).text
print(f"greedy identical: {a == b!s:<5} | fp32: {a.strip()[:70]}")
if a != b:
print(f" | int8: {b.strip()[:70]}")
if __name__ == "__main__":
main()
Code explained
- In simple words: make a half-size copy of TinyLM and check, token by token, whether it still predicts the same things.
- What happens:
validation_ids: the held-out tail of the corpus, tokenized.evaluate: perplexity (the exponential of the average next-token loss; lower means the text looks more predictable to the model, Module 3) over non-overlapping 128-token windows, plus the model's top prediction at every position.size_on_disk: the bytestorch.savewrites for the state dict.forward_ms: median forward time, alternating between the two models so background load hits both equally.main:torch.ao.quantization.quantize_dynamicconverts the linear layers of a deep copy. Then it compares everything and runs three greedy generations with each model. The script silences aDeprecationWarning:torch.ao.quantizationannounces its own removal (the message says 2.10) and points to the separatetorchaopackage'squantize_API. It still works in torch 2.14.0, which is all a one-line demonstration needs; for new production code, usetorchaoor your serving engine's quantization.
- Comes out:text
machine load (1-minute average) 0.22 on 2 cores; torch threads 1 fp32: size 4203 KiB forward (1 x 128 tokens) 6.11 ms validation perplexity 1.6531 int8: size 2167 KiB forward (1 x 128 tokens) 4.62 ms validation perplexity 1.6535 top-1 next-token agreement: 8180 of 8243 positions (99.24%) greedy identical: True | fp32: Team costs 12 USD per user per month, or 10 USD billed annually. Busin greedy identical: True | fp32: Reset links expire after 30 minutes and can be used once. Please reque greedy identical: True | fp32: Sorry about the duplicate charge. Duplicate charges are always refundeThe file halves (4,203 KiB to 2,167 KiB; the embedding table is not a linear layer, so it stays fp32). Perplexity moves from 1.6531 to 1.6535, a 0.02 percent change. The top prediction changes at 63 of 8,243 positions (0.76 percent), and all three greedy replies are identical. The forward pass got about a quarter faster here (6.1 to 4.6 ms), but at this tiny size and on a shared machine that number moves from run to run (an earlier run gave 4.0 to 3.3 ms); do not generalize it.
What does transfer: quantization error is small on average and concentrated in a few positions, so a perplexity or average-score check can look perfect while specific outputs change. Check both an aggregate metric and a set of exact outputs that matter to you (here, the three replies). For real models, published results are the guide: 8-bit weights are usually close to lossless, 4-bit costs more and varies by method and model, and small models suffer more than large ones. Run your Module 10 eval suite on the quantized model before you ship it.
| Situation | Use this | Why |
|---|---|---|
| Memory is tight and quality must not move | 8-bit weights | Near-lossless in most published results; measure anyway |
| You need more concurrent users or a smaller GPU | 4-bit weights (AWQ, GPTQ, GGUF Q4 variants, or the model's shipped format) | Quarter-size weights; quality cost varies, so run the eval suite |
| The model ships in a quantized format (gpt-oss MXFP4) | Use it as shipped | The model was trained or calibrated for it |
| CPU or laptop inference | llama.cpp GGUF quantizations | Built for exactly that |
Throughput, batching, and utilization
Decoding one token for one user reads every weight from memory to do a small amount of math. Decoding one token for 32 users reads the same weights once and does 32 times the math. That is why batching raises throughput (tokens per second for the whole server) far more than it raises latency (seconds per request for each user), up to the point where the math, not memory, becomes the limit.
The script runs TinyLM on batches of 1 to 32 requests, each a 64-token prompt followed by 32 greedy decode steps with the KV cache, exactly the way a serving engine steps a batch.
examples/m13_batching.py
"""Throughput vs batch size on TinyLM: decode many requests together, measure the tradeoff.
Each request is a 64-token prompt from the held-out corpus tail followed by
32 greedy decode steps with the KV cache. A batch of B requests runs those
steps together, the way a serving engine batches concurrent users.
"""
from __future__ import annotations
import os
import statistics
import time
import torch
from supportdesk.data import DATA_DIR
from supportdesk.tinylm import load
torch.manual_seed(0)
torch.set_num_threads(1) # one thread: steadier timings on a shared 2-core machine
PROMPT, DECODE = 64, 32
@torch.no_grad()
def run_batch(model, prompts: torch.Tensor) -> tuple[float, float]:
"""Return (prefill seconds, decode seconds) for one batch of prompts."""
caches = [dict() for _ in model.blocks]
started = time.perf_counter()
logits = model(prompts, caches)[:, -1]
prefill = time.perf_counter() - started
started = time.perf_counter()
for step in range(DECODE):
next_ids = logits.argmax(-1, keepdim=True)
logits = model(next_ids, caches, start_pos=PROMPT + step)[:, -1]
return prefill, time.perf_counter() - started
def main() -> None:
print(f"machine load (1-minute average) {os.getloadavg()[0]:.2f} on {os.cpu_count()} cores; "
f"torch threads {torch.get_num_threads()}")
model, tokenizer = load()
ids = torch.tensor(tokenizer.encode((DATA_DIR / "corpus.txt").read_text(encoding="utf-8")).ids)
tail = ids[int(len(ids) * 0.95):]
pool = torch.stack([tail[i * PROMPT:(i + 1) * PROMPT] for i in range(32)])
run_batch(model, pool[:4]) # warm-up
print(f"{'batch':>5} {'decode tok/s':>13} {'ms per step':>12} {'request latency ms':>19} {'prefill ms':>11}")
rows = []
for batch in (1, 2, 4, 8, 16, 32):
trials = [run_batch(model, pool[:batch]) for _ in range(5)]
prefill = statistics.median(t[0] for t in trials)
decode = statistics.median(t[1] for t in trials)
tps = batch * DECODE / decode
latency = (prefill + decode) * 1000
rows.append((batch, tps, latency))
print(f"{batch:>5} {tps:>13.0f} {decode / DECODE * 1000:>12.2f} {latency:>19.0f} {prefill * 1000:>11.1f}")
base_tps, base_latency = rows[0][1], rows[0][2]
print("\nrelative to batch 1:")
for batch, tps, latency in rows:
print(f" batch {batch:>2}: throughput x{tps / base_tps:5.1f}, each request waits x{latency / base_latency:4.1f}")
if __name__ == "__main__":
main()
Code explained
- In simple words: serve 1, 2, 4, up to 32 customers at once and time how the whole kitchen and each customer fare.
- What happens:
run_batchprefills all prompts in one forward pass, then runs 32 decode steps where each step feeds one new token per request through the model, reusing the caches.mainbuilds 32 prompts from the held-out corpus, warms up, and for each batch size takes the median of 5 runs. Throughput isbatch x 32 tokens / decode time; request latency is prefill plus decode. - Comes out:text
machine load (1-minute average) 0.28 on 2 cores; torch threads 1 batch decode tok/s ms per step request latency ms prefill ms 1 1039 0.96 34 2.7 2 1999 1.00 36 4.3 4 3042 1.31 49 7.2 8 4565 1.75 69 13.0 16 6098 2.62 110 25.8 32 7087 4.52 194 49.2 relative to batch 1: batch 1: throughput x 1.0, each request waits x 1.0 batch 2: throughput x 1.9, each request waits x 1.1 batch 4: throughput x 2.9, each request waits x 1.5 batch 8: throughput x 4.4, each request waits x 2.1 batch 16: throughput x 5.9, each request waits x 3.3 batch 32: throughput x 6.8, each request waits x 5.8Going from 1 to 32 concurrent requests multiplies throughput by about 7 while each request waits nearly 6 times longer. On a GPU the curve is much flatter at the start (small batches are nearly free) and throughput keeps rising for longer, but the tradeoff has the same shape: batch size is a dial between cost per token and latency per user, and a serving engine lets you cap it.
Utilization is the share of time the hardware does useful work. A GPU you rent by the hour costs the same whether it serves 32 requests at once or sits idle at 3 a.m. That fact, not the speed of the GPU, decides most self-hosting economics, as the next section shows.
A real diagnosis from writing this module. The first runs of these timing scripts were wildly inconsistent: sometimes a decode step took under 1 ms, sometimes over 100 ms. The trace pointed to threads: by default PyTorch uses one thread per core, and on a 2-core machine that other processes were also using, the two threads kept waiting for each other. Here is the check, run once on a quiet machine and once with two busy background processes:
"""Diagnose slow decoding on a shared machine: time TinyLM decode steps with 1 and with 2 threads."""
import os
import statistics
import time
import torch
from supportdesk.tinylm import load
model, _ = load()
x = torch.randint(0, 1503, (1, 64), generator=torch.Generator().manual_seed(0))
print(f"machine load (1-minute average) {os.getloadavg()[0]:.2f} on {os.cpu_count()} cores")
results: dict[int, list[float]] = {1: [], 2: []}
with torch.no_grad():
for _ in range(5): # alternate settings so background load hits both
for threads in (2, 1):
torch.set_num_threads(threads)
model(x) # warm-up after switching
started = time.perf_counter()
for _ in range(100):
model(x[:, :1])
results[threads].append((time.perf_counter() - started) / 100 * 1000)
for threads, runs in results.items():
print(f"{threads} thread(s): median {statistics.median(runs):.2f} ms per decode step, "
f"worst {max(runs):.2f} ms (5 rounds of 100 steps)")
Code explained
- In simple words: time the same decode step with one thread and with two, taking turns so background load hits both.
- What happens: the script alternates 5 rounds of 100 single-token forward passes per setting and reports the median and the worst round, next to the machine's 1-minute load average.
- Comes out: first on the machine as it was, then with two CPU-burning processes started in the background (
python3 -c "while True: pass" &, twice, stopped afterwards by PID):textmachine load (1-minute average) 0.39 on 2 cores 1 thread(s): median 0.70 ms per decode step, worst 0.84 ms (5 rounds of 100 steps) 2 thread(s): median 0.67 ms per decode step, worst 0.69 ms (5 rounds of 100 steps)textmachine load (1-minute average) 2.40 on 2 cores 1 thread(s): median 2.33 ms per decode step, worst 2.57 ms (5 rounds of 100 steps) 2 thread(s): median 134.99 ms per decode step, worst 162.71 ms (5 rounds of 100 steps)With spare cores, the two settings are close. With every core busy, two threads took about 135 ms per step against 2.3 ms for one thread, roughly 58 times slower, because each step waits for a thread that the operating system has parked. The same thing happens on real inference servers when too many workers share too few cores. The fix is to set thread counts explicitly (
torch.set_num_threads,OMP_NUM_THREADS, or the engine's own setting) and always report load next to a timing.
The real total cost vs API pricing
Now the question finance will ask: should Brightlane rent GPUs instead of paying per token? The script measures real input sizes (triage prompt plus draft prompt with three retrieved help-center articles, for all 72 tickets), prices them with supportdesk.pricing, and compares with renting GPUs.
GPU prices are on-demand prices per GPU-hour read on 21 September 2026: RunPod lists the L4 at 0.49 USD, the A100 SXM 80 GB at 1.59 USD, and the H100 SXM at 3.49 USD (runpod.io/pricing); Lambda lists the H100 SXM at 3.99 USD and the A100 SXM 80 GB at 2.79 USD (lambda.ai/pricing). Prices change monthly; verify them.
One correction to the shared price table first. supportdesk/pricing.py bills cached input tokens for Groq's gpt-oss models at the full input price, but Groq's prompt-caching page says "There is a 50% discount for cached input tokens", and that "the prompt caching discount does not stack with the batch discount" (Groq docs: prompt caching, read 21 September 2026). The script corrects this locally instead of editing the shared file; the difference is shown next to the original.
examples/m13_tco.py
"""Total cost of ownership: hosted API vs a rented GPU, for Brightlane's ticket volume.
Token counts per ticket are measured on the real dataset. API prices come from
supportdesk.pricing. GPU prices and every operating assumption are labeled
below; replace them with your own quotes and, above all, with throughput you
measured on your own serving engine.
"""
from __future__ import annotations
import statistics
from dataclasses import dataclass
from supportdesk.data import get_article, load_tickets
from supportdesk.kb_search import KBSearch
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, Price, cost_usd
from supportdesk.tokens import count_messages
HOURS_PER_MONTH = 730
# GPU on-demand prices, USD per GPU-hour, read on 21 Sep 2026 (verify before use):
# RunPod secure cloud (runpod.io/pricing): L4 0.49, A100 SXM 80 GB 1.59, H100 SXM 3.49
# Lambda (lambda.ai/pricing): H100 SXM 3.99, A100 SXM 80 GB 2.79
GPU_PRICES = {"L4 (RunPod)": 0.49, "A100 80GB (RunPod)": 1.59, "H100 (RunPod)": 3.49, "H100 (Lambda)": 3.99}
# Assumptions (ours, for teaching; change them):
TICKETS_PER_MONTH = 30_000
OUTPUT_TOKENS_PER_TICKET = 350 # triage JSON (about 60) + draft reply (about 290)
OPS_HOURS_PER_MONTH = 20 # patching, upgrades, on-call for the serving stack
OPS_RATE_USD = 100 # loaded cost of an engineer hour
REPLICAS = 2 # one GPU is a single point of failure
# supportdesk.pricing bills Groq's cached input at the full input price. Groq's prompt-caching
# page (console.groq.com/docs/prompt-caching, read 21 Sep 2026) lists a 50 percent discount on
# cached input tokens for gpt-oss models, which does not stack with the 50 percent batch discount.
# We correct it here instead of editing the shared price table.
GROQ_CACHED = {m: Price(PRICES[m].input, PRICES[m].output, PRICES[m].input * 0.5)
for m in ("openai/gpt-oss-20b", "openai/gpt-oss-120b")}
def groq_cost(usage: Usage, model: str, batch: bool = False) -> float:
"""Cost with Groq's cached-input discount; in batch mode the cache discount does not stack."""
if batch:
return cost_usd(Usage(usage.input_tokens, usage.output_tokens), model, batch=True)
p = GROQ_CACHED[model]
uncached = usage.input_tokens - usage.cached_tokens
return (uncached * p.input + usage.cached_tokens * p.cached_input + usage.output_tokens * p.output) / 1e6
def measured_input_tokens() -> list[int]:
"""Input tokens per ticket for triage + draft, built from the real tickets and KB articles."""
search = KBSearch()
system_triage = "Classify the Brightlane support ticket. Reply with JSON: category, priority, language, summary, needs_human."
system_draft = "Draft a reply for a human agent to review. Use only the articles below and cite their ids."
totals = []
for t in load_tickets():
triage = [{"role": "system", "content": system_triage}, {"role": "user", "content": t.text}]
articles = "\n\n".join(f"[{h.article_id}]\n{get_article(h.article_id).body}" for h in search.search(t.text, k=3))
draft = [{"role": "system", "content": system_draft + "\n\n" + articles}, {"role": "user", "content": t.text}]
totals.append(count_messages(triage) + count_messages(draft))
return totals
@dataclass
class SelfHost:
gpu: str
tokens_per_second: float # sustained output tokens/s for the whole GPU at your latency target
utilization: float # share of the month the GPU does useful work
def monthly_usd(self) -> float:
return GPU_PRICES[self.gpu] * HOURS_PER_MONTH * REPLICAS + OPS_HOURS_PER_MONTH * OPS_RATE_USD
def capacity_tickets(self) -> float:
return self.tokens_per_second * 3600 * HOURS_PER_MONTH * self.utilization / OUTPUT_TOKENS_PER_TICKET * REPLICAS
def main() -> None:
inputs = measured_input_tokens()
mean_in = statistics.mean(inputs)
print(f"measured input tokens per ticket (o200k estimate, n={len(inputs)}): "
f"mean {mean_in:.0f}, median {statistics.median(inputs):.0f}, max {max(inputs)}")
usage = Usage(input_tokens=round(mean_in), output_tokens=OUTPUT_TOKENS_PER_TICKET)
print(f"\nHosted API at {TICKETS_PER_MONTH:,} tickets per month:")
for model in ("openai/gpt-oss-20b", "openai/gpt-oss-120b", "gemini-3.5-flash-lite", "gemini-3.5-flash"):
per_ticket = cost_usd(usage, model)
print(f" {model:<24} {per_ticket * 1000:6.3f} USD per 1k tickets {per_ticket * TICKETS_PER_MONTH:9.2f} USD per month")
cached = Usage(round(mean_in), OUTPUT_TOKENS_PER_TICKET, cached_tokens=round(mean_in * 0.9))
print("\nIf 90% of input tokens hit the prompt cache (USD per 1k tickets):")
for model in ("openai/gpt-oss-20b", "openai/gpt-oss-120b"):
print(f" {model:<24} pricing.py {cost_usd(cached, model) * 1000:.3f} with Groq's 50% cached rate "
f"{groq_cost(cached, model) * 1000:.3f} batch (no stacking) {groq_cost(cached, model, batch=True) * 1000:.3f}")
print(f"\nSelf-hosted, {REPLICAS} replicas, plus {OPS_HOURS_PER_MONTH} ops hours at {OPS_RATE_USD} USD:")
for gpu in GPU_PRICES:
fixed = SelfHost(gpu, 1, 1).monthly_usd()
print(f" {gpu:<20} {fixed:8.2f} USD per month, whether you send 0 tickets or a million")
print("\nBreak-even: the volume where self-hosting the same open model costs what the API costs")
for gpu, model in (("L4 (RunPod)", "openai/gpt-oss-20b"), ("H100 (RunPod)", "openai/gpt-oss-120b")):
api = cost_usd(usage, model)
volume = SelfHost(gpu, 1, 1).monthly_usd() / api
tps = volume * OUTPUT_TOKENS_PER_TICKET / (HOURS_PER_MONTH * 3600) / REPLICAS
print(f" {model:<20} on {gpu:<14} {volume:>12,.0f} tickets per month, "
f"which needs {tps:,.0f} output tok/s per GPU around the clock")
print("\nWhat one self-hosted ticket costs, by sustained throughput and utilization (L4 pair):")
print(f" {'tok/s per GPU':>13} " + " ".join(f"{'util ' + format(u, '.0%'):>12}" for u in (0.1, 0.3, 0.6)))
for tps in (100, 300, 1000):
cells = []
for u in (0.1, 0.3, 0.6):
host = SelfHost("L4 (RunPod)", tps, u)
cells.append(f"{host.monthly_usd() / host.capacity_tickets() * 1000:>9.3f}/1k")
print(f" {tps:>13} " + " ".join(f"{c:>12}" for c in cells))
need = TICKETS_PER_MONTH * OUTPUT_TOKENS_PER_TICKET / (HOURS_PER_MONTH * 3600)
print(f"\nBrightlane's average load is only {need:.1f} output tokens/s: the GPUs would sit mostly idle.")
if __name__ == "__main__":
main()
Code explained
- In simple words: a spreadsheet in code that puts the per-token bill next to the rent-a-GPU bill for the same monthly work.
- What happens:
- The constants at the top are labeled assumptions: 30,000 tickets a month, 350 output tokens per ticket, 20 hours a month of engineering time for the serving stack at 100 USD an hour, and 2 replicas because one GPU is a single point of failure.
GROQ_CACHEDandgroq_cost: Groq's real cached-input discount, with batch mode getting the batch discount and no cache discount on top.measured_input_tokens(): builds both prompts for every real ticket (withKBSearchpicking three articles) and counts them withcount_messages. These are o200k estimates; each provider's own tokenizer will differ a little (Module 2).SelfHost: the monthly bill (GPU hours times replicas plus ops time) and the monthly capacity (sustained output tokens per second times utilization).main: API cost per model, the cached variants, the fixed self-hosting bill, the break-even volume, and a grid of cost per 1,000 tickets across throughput and utilization.
- Comes out:
| Situation | Use this | Why |
|---|---|---|
| Low or spiky volume (Brightlane today) | Hosted API | You pay only for tokens; idle GPUs cost the same as busy ones |
| Very high, steady volume on a task an open model handles | Self-hosting, after measuring sustained throughput | Break-even arrives only when the GPUs stay busy |
| Privacy rules forbid sending tickets out, at any volume | Self-hosting, priced honestly | Cost stops being the deciding factor; budget for it |
| Work that can wait hours | Batch APIs | Half price at most providers, no infrastructure |