CourseLarge Language Models · Module 13: Deployment, Operations, and Economics · part 67 of 80
Part 67 · Module 13: Deployment, Operations, and Economics

Part B: Self-hosting

30 min read·22 Sept 2026

Part B: Self-hosting

Self-hosting means you run the model yourself: you pick a GPU, load open weights into its memory, and run a serving engine (software that turns a model file into an HTTP endpoint that many users can call at once). No open-weight model can be downloaded in this course's sandbox, so every mechanism below is measured for real on TinyLM, and every number about real models is computed from their published configuration files, not guessed.

A note on timings before we start. This module was written on a 2-core CPU that other jobs were sharing. Every timing script sets torch.set_num_threads(1) and prints the machine load next to its numbers; the "threads" diagnosis later in this part shows why. Timings on your machine will differ; the shapes of the curves are what transfer.

Hardware sizing and memory arithmetic

A GPU has a fixed amount of fast memory (VRAM: 24 GB on an NVIDIA L4, 80 GB on an A100 80GB or H100). Serving a model needs three things to fit in it:

  1. Weights: parameters times bytes per parameter. At bf16 (16-bit floats, the usual serving precision) that is 2 bytes each, so an 8-billion-parameter model needs about 16 GB before it has answered anything.
  2. KV cache: during generation, every layer stores a key vector and a value vector for every token of every active conversation, so the model does not recompute them (Module 3). Per token that is 2 (K and V) x layers x KV heads x head size x bytes. Many modern models use grouped-query attention (GQA): several query heads share one key/value head, which shrinks this cache. Llama 3.1 8B has 32 query heads but only 8 KV heads.
  3. Overhead: activations, the CUDA runtime, and the engine's own bookkeeping. We assume 2 GiB; measure yours.

Whatever is left after weights and overhead is KV cache, and KV cache is what decides how many users you can serve at once and how long their conversations can be. That is the real sizing question.

The script below checks the formulas on TinyLM, where we can count the real tensors, then applies them to three real open models using the values in their published config.json files on Hugging Face (meta-llama/Llama-3.1-8B, Qwen/Qwen3-8B, and openai/gpt-oss-120b, read 21 September 2026).

examples/m13_memory.py

python
"""Memory arithmetic for self-hosting: weights plus KV cache, from a model's config.

TinyLM is computed exactly and checked against the real tensors. Real open
models are computed from their published config.json files (values copied
below, with the source named) so you can size a GPU before renting one.
"""
from __future__ import annotations

from dataclasses import dataclass

import torch

from supportdesk.tinylm import load

GIB = 1024 ** 3
BYTES = {"fp32": 4.0, "bf16": 2.0, "int8": 1.0, "int4": 0.5, "mxfp4": 4.25 / 8}  # MXFP4: 4 bits + 1 shared 8-bit scale per 32 values


@dataclass(frozen=True)
class Arch:
    """The config.json fields that decide memory. Values come from each model's published config."""

    name: str
    hidden: int
    layers: int
    heads: int
    kv_heads: int
    head_dim: int
    intermediate: int
    vocab: int
    tied: bool
    experts: int = 1            # mixture-of-experts: experts per layer (1 = dense)
    active_experts: int = 1     # experts each token actually uses
    full_attention_layers: int | None = None   # others use a sliding window
    window: int | None = None


# Sources: config.json on Hugging Face for meta-llama/Llama-3.1-8B, Qwen/Qwen3-8B, and
# openai/gpt-oss-120b (read 21 Sep 2026). gpt-oss-120b alternates sliding-window (128 tokens)
# and full-attention layers, and routes each token to 4 of its 128 experts.
LLAMA_31_8B = Arch("Llama 3.1 8B", 4096, 32, 32, 8, 128, 14336, 128256, False)
QWEN3_8B = Arch("Qwen3-8B", 4096, 36, 32, 8, 128, 12288, 151936, False)
GPT_OSS_120B = Arch("gpt-oss-120b", 2880, 36, 64, 8, 64, 2880, 201088, False, experts=128,
                    active_experts=4, full_attention_layers=18, window=128)


def params(a: Arch) -> dict[str, int]:
    """Parameter count by part, for a Llama-style block (gated MLP, grouped-query attention).

    Small terms (norm weights, attention biases, router, Qwen3's q/k norms) are
    left out; they change the total by far less than 0.1 percent.
    """
    embed = a.vocab * a.hidden * (1 if a.tied else 2)
    attn = a.hidden * a.heads * a.head_dim * 2 + a.hidden * a.kv_heads * a.head_dim * 2
    mlp = 3 * a.hidden * a.intermediate * a.experts
    active = embed // (1 if a.tied else 2) + (attn + 3 * a.hidden * a.intermediate * a.active_experts) * a.layers
    return {"embeddings": embed, "attention": attn * a.layers, "mlp": mlp * a.layers,
            "total": embed + (attn + mlp) * a.layers, "active": active}


def kv_bytes_per_token(a: Arch, dtype: str = "bf16") -> float:
    """K and V, for every layer, for every KV head: 2 * layers * kv_heads * head_dim * bytes."""
    layers = a.full_attention_layers if a.full_attention_layers is not None else a.layers
    return 2 * layers * a.kv_heads * a.head_dim * BYTES[dtype]


def kv_bytes_for_sequence(a: Arch, tokens: int, dtype: str = "bf16") -> float:
    full = kv_bytes_per_token(a, dtype) * tokens
    if a.window is not None:  # sliding-window layers keep at most `window` tokens
        sliding_layers = a.layers - (a.full_attention_layers or 0)
        full += 2 * sliding_layers * a.kv_heads * a.head_dim * BYTES[dtype] * min(tokens, a.window)
    return full


def tinylm_check() -> None:
    model, tokenizer = load()
    cfg = model.cfg
    d, v, c, n = cfg.d_model, cfg.vocab_size, cfg.context, cfg.n_layers
    per_layer = (2 * d) + (d * 3 * d + 3 * d) + (d * d + d) + (2 * d) + (d * 4 * d + 4 * d) + (4 * d * d + d)
    by_hand = v * d + c * d + n * per_layer + 2 * d        # head is tied to tok_emb, so it adds nothing
    weights_fp32 = sum(p.numel() * p.element_size() for p in model.parameters())
    print(f"TinyLM parameters: by hand {by_hand:,}  model.num_parameters() {model.num_parameters():,}")
    print(f"TinyLM weights in fp32: {weights_fp32:,} bytes = {weights_fp32 / 1024**2:.2f} MiB")
    used = tokenizer.get_vocab_size()
    print(f"embedding rows: {v:,} in the config, {used:,} used by the tokenizer, "
          f"{(v - used) * d * 4:,} bytes of rows no token can reach")

    kv_formula = 2 * n * cfg.n_heads * (d // cfg.n_heads) * 4   # TinyLM has no grouped-query attention
    ids = tokenizer.encode("Customer (Ben): I was charged twice for the Team plan.\nAgent (Dara):").ids
    caches = [dict() for _ in model.blocks]
    with torch.no_grad():
        model(torch.tensor([ids]), caches)
    measured = sum(cache["k"].nbytes + cache["v"].nbytes for cache in caches)
    print(f"TinyLM KV cache: formula {kv_formula:,} bytes/token; measured {measured:,} bytes "
          f"for {len(ids)} tokens = {measured // len(ids):,} bytes/token")
    print(f"TinyLM full context ({c} tokens): {kv_formula * c / 1024:.0f} KiB per sequence")


def real_models() -> None:
    print(f"\n{'model':<14}{'params (B)':>11}{'bf16 GiB':>10}{'int4 GiB':>10}{'KV KiB/tok':>12}{'KV GiB @8k':>12}{'KV GiB @128k':>14}")
    for a in (LLAMA_31_8B, QWEN3_8B):
        p = params(a)["total"]
        print(f"{a.name:<14}{p / 1e9:>11.2f}{p * 2 / GIB:>10.1f}{p * 0.5 / GIB:>10.1f}"
              f"{kv_bytes_per_token(a) / 1024:>12.0f}{kv_bytes_for_sequence(a, 8192) / GIB:>12.2f}"
              f"{kv_bytes_for_sequence(a, 131072) / GIB:>14.2f}")
    a = GPT_OSS_120B
    p = params(a)
    experts, rest = p["mlp"], p["total"] - p["mlp"]
    size = experts * BYTES["mxfp4"] + rest * BYTES["bf16"]
    print(f"{a.name:<14}{p['total'] / 1e9:>11.2f}{p['total'] * 2 / GIB:>10.1f}{'':>10}"
          f"{kv_bytes_per_token(a) / 1024:>12.0f}{kv_bytes_for_sequence(a, 8192) / GIB:>12.2f}"
          f"{kv_bytes_for_sequence(a, 131072) / GIB:>14.2f}")
    print(f"gpt-oss-120b as shipped (experts MXFP4, rest bf16): {size / GIB:.1f} GiB of weights")
    print(f"gpt-oss-120b parameters used per token (4 of 128 experts): {p['active'] / 1e9:.2f} B "
          f"(counting the {a.vocab * a.hidden / 1e9:.2f} B output head, not the input lookup table)")


def fits(a: Arch, gpu_gib: float, weight_dtype: str, context: int, overhead_gib: float = 2.0) -> int:
    """How many sequences of `context` tokens fit next to the weights (overhead: activations, runtime)."""
    weights = params(a)["total"] * BYTES[weight_dtype] / GIB
    free = gpu_gib - weights - overhead_gib
    return max(int(free // (kv_bytes_for_sequence(a, context) / GIB)), 0)


if __name__ == "__main__":
    tinylm_check()
    real_models()
    print("\nConcurrent sequences that fit (2 GiB runtime overhead assumed):")
    for gpu, gib in (("L4 24 GB", 22.5), ("A100 80 GB", 80.0)):
        for dtype in ("bf16", "int4"):
            print(f"  Qwen3-8B on {gpu:<10} weights {dtype}: "
                  f"{fits(QWEN3_8B, gib, dtype, 4096):>4} at 4k tokens, {fits(QWEN3_8B, gib, dtype, 32768):>3} at 32k")

Code explained

  • In simple words: a memory calculator that proves itself on a model it can open, then sizes models it cannot download.
  • What happens:
    • BYTES: bytes per parameter for common formats. MXFP4 is the 4-bit format gpt-oss ships its expert weights in: blocks of 32 four-bit values share one 8-bit scale, so it costs 4.25 bits per value.
    • Arch: the handful of config fields that decide memory: hidden size, layers, attention heads, KV heads, head size, MLP width, vocabulary, whether input and output embeddings are shared (tied), and for mixture-of-experts models how many experts exist and how many each token uses. gpt-oss-120b also alternates sliding-window layers (which only keep the last 128 tokens) with full-attention layers.
    • params(a): counts parameters for a Llama-style block (attention projections plus a gated MLP with three matrices). active counts what one token actually touches in a mixture-of-experts (MoE) model, where a small router picks a few experts per token.
    • kv_bytes_per_token and kv_bytes_for_sequence: the KV formula; sliding-window layers are capped at the window.
    • tinylm_check(): counts TinyLM's parameters by hand, then runs a real prompt and adds up the bytes in the actual cache tensors. It also compares the embedding table (2,048 rows in the config) with the tokenizer (1,503 real tokens).
    • fits(...): how many sequences of a given length fit beside the weights.
  • Comes out:

    text
    TinyLM parameters: by hand 1,071,872  model.num_parameters() 1,071,872
    TinyLM weights in fp32: 4,287,488 bytes = 4.09 MiB
    embedding rows: 2,048 in the config, 1,503 used by the tokenizer, 279,040 bytes of rows no token can reach
    TinyLM KV cache: formula 4,096 bytes/token; measured 73,728 bytes for 18 tokens = 4,096 bytes/token
    TinyLM full context (128 tokens): 512 KiB per sequence
    
    model          params (B)  bf16 GiB  int4 GiB  KV KiB/tok  KV GiB @8k  KV GiB @128k
    Llama 3.1 8B         8.03      15.0       3.7         128        1.00         16.00
    Qwen3-8B             8.19      15.3       3.8         144        1.12         18.00
    gpt-oss-120b       116.78     217.5                    36        0.29          4.50
    gpt-oss-120b as shipped (experts MXFP4, rest bf16): 60.7 GiB of weights
    gpt-oss-120b parameters used per token (4 of 128 experts): 5.12 B (counting the 0.58 B output head, not the input lookup table)
    
    Concurrent sequences that fit (2 GiB runtime overhead assumed):
      Qwen3-8B on L4 24 GB   weights bf16:    9 at 4k tokens,   1 at 32k
      Qwen3-8B on L4 24 GB   weights int4:   29 at 4k tokens,   3 at 32k
      Qwen3-8B on A100 80 GB weights bf16:  111 at 4k tokens,  13 at 32k
      Qwen3-8B on A100 80 GB weights int4:  131 at 4k tokens,  16 at 32k
    

    Read it from the top. The by-hand count matches num_parameters() exactly, and the measured cache is exactly 4,096 bytes per token, so the formula is right. TinyLM also carries 545 embedding rows that no token id can reach (the tokenizer learned only 1,503 tokens): 279 KB of dead weight. Real models pad vocabularies too, usually on purpose, to a multiple that suits GPU kernels; it is harmless but it does count toward memory.

    The real-model rows match their published sizes (Llama 3.1 8B is 8.03 B parameters, Qwen3-8B is 8.19 B), and two lessons stand out. First, the KV cache is not small: one 128k-token conversation on Llama 3.1 8B needs 16 GiB of cache, more than the weights. (Qwen3-8B's native context is 32,768 tokens and reaches 131,072 only with YaRN scaling, so its 128k column is a what-if.) Second, sliding windows and MoE change everything: gpt-oss-120b has 117 B parameters, but its cache is 36 KiB per token and each token uses only 5.1 B parameters, which matches the 5.1 B "active parameters" OpenAI publishes. That is why a 117 B model can generate quickly on one 80 GB GPU: memory is set by total parameters, speed mostly by active ones.

    The last block is the sizing answer. On a 24 GB L4, Qwen3-8B in bf16 serves 9 conversations of 4k tokens at once, but only one of 32k. Quantizing the weights to 4 bits frees memory for 29. Nothing here depends on the GPU's speed; that comes next.

SituationUse thisWhy
Model under about 10 B parameters, short contexts, modest trafficOne 24 GB GPU (L4 class), 4-bit or 8-bit weightsWeights fit with room left for KV cache
8 B model with long contexts or many concurrent usersAn 80 GB GPUKV cache, not weights, becomes the limit
Large MoE model such as gpt-oss-120bAn 80 GB GPU with the shipped 4-bit expertsTotal parameters decide memory; active parameters decide speed
70 B dense model or largerSeveral GPUs with tensor parallelism, or a hosted APIWeights alone exceed one GPU

Serving engines and what they optimize

You do not serve a model with a for loop over generate(). Serving engines exist because naive generation wastes the GPU in two ways: memory for the KV cache gets reserved for the longest possible reply and mostly sits empty, and one user's request occupies the whole GPU while it waits for the next token.

The main ideas, which every engine below implements in some form:

  • Continuous batching: the engine runs one decode step for every active request together, and new requests join the batch between steps instead of waiting for the whole batch to finish.
  • Paged KV cache: cache memory is handed out in small fixed-size blocks, like pages in an operating system, so it is used as needed rather than reserved up front. This is the PagedAttention idea from the vLLM paper (Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention", SOSP 2023).
  • Prefix caching: requests that start with the same tokens (a system prompt, a shared help-center excerpt) reuse the stored KV cache for that prefix instead of recomputing it. You will measure this on TinyLM in Part C.
  • Quantization and speculative decoding: smaller weights, and a small draft model that proposes tokens the big model checks in one pass.
EngineWhat it optimizesWhere it fits
vLLMThroughput on GPUs: paged KV cache, continuous batching, automatic prefix caching, OpenAI-compatible serverThe default choice for serving open models on NVIDIA or AMD GPUs
SGLangPrefix reuse (its RadixAttention keeps a tree of cached prefixes), a low-overhead scheduler, prefill-decode disaggregation, speculative decoding, FP4/FP8/INT4 quantizationWorkloads with heavy shared prefixes: agents, multi-turn chat, structured generation
TensorRT-LLMPeak speed on NVIDIA hardware by compiling optimized kernels, in-flight batching, FP8 and FP4Large, stable deployments on NVIDIA GPUs where a build step is acceptable
llama.cppRunning quantized models (GGUF format) on CPUs, Apple Silicon, and consumer GPUsLaptops, edge devices, single-user or low-traffic servers
OllamaA friendly local wrapper: model library, one-command downloads, OpenAI-compatible endpointDevelopment and the course's local provider (llm.py talks to it on port 11434)

One engine has left the table: Hugging Face's text-generation-inference (TGI) README now says it "is now in maintenance mode" and points users to vLLM, SGLang, and local engines such as llama.cpp and MLX (TGI on GitHub). If a tutorial you are following uses TGI, it still works, but new projects should pick one of the others. The SGLang feature list above is from its README (sgl-project/sglang).

SituationUse thisWhy
First production deployment of an open model on GPUsvLLMWidest model support, OpenAI-compatible API, good defaults
Most requests share long prefixes (agents, RAG with fixed context)SGLangPrefix reuse is its central design
Fixed model, very high volume, NVIDIA only, team can own a build pipelineTensorRT-LLMHighest speed per GPU at the cost of flexibility
No GPU, or a single userllama.cpp or OllamaQuantized CPU and Apple Silicon inference

Quantization and the quality it costs

Quantization stores weights with fewer bits: int8 (1 byte) or 4-bit formats instead of bf16 (2 bytes) or fp32 (4 bytes). It halves or quarters memory, which (as the table above showed) turns into more concurrent users. The question is always what it costs in quality, and the only honest answer is a measurement on your own task.

PyTorch's dynamic int8 quantization is a real, one-line version: it stores every nn.Linear weight as int8 with a scale, and converts activations on the fly. The script applies it to TinyLM and measures size, speed, perplexity on the held-out 5 percent of the corpus (the same 8,244 tokens pretraining held out), how often the top prediction changes, and whether greedy replies change.

examples/m13_quantize.py

python
"""Quantize TinyLM's linear layers to int8 and measure what it saves and what it costs.

Measures, on this machine: size on disk, forward-pass latency, validation
perplexity, agreement of the top next-token prediction, and greedy generations.
"""
from __future__ import annotations

import copy
import io
import math
import os
import statistics
import time
import warnings

import torch

from supportdesk.data import DATA_DIR
from supportdesk.tinylm import SamplingParams, generate, load, loss_on

warnings.filterwarnings("ignore")          # torch.ao.quantization prints deprecation notices
torch.manual_seed(0)
torch.set_num_threads(1)          # one thread: steadier timings on a shared 2-core machine


def validation_ids(tokenizer) -> torch.Tensor:
    """The same held-out tail (last 5 percent of tokens) that pretraining used for validation."""
    ids = torch.tensor(tokenizer.encode((DATA_DIR / "corpus.txt").read_text(encoding="utf-8")).ids)
    return ids[int(len(ids) * 0.95):]


@torch.no_grad()
def evaluate(model, ids: torch.Tensor, context: int = 128) -> tuple[float, list[int]]:
    """Perplexity over non-overlapping windows, plus the argmax prediction at every position."""
    total, count, predictions = 0.0, 0, []
    for start in range(0, len(ids) - 1, context):
        chunk = ids[start: start + context + 1]
        if len(chunk) < 2:
            break
        x, y = chunk[:-1].unsqueeze(0), chunk[1:].unsqueeze(0)
        total += loss_on(model, x, y).item() * y.numel()
        count += y.numel()
        predictions += model(x)[0].argmax(-1).tolist()
    return math.exp(total / count), predictions


def size_on_disk(model) -> int:
    buffer = io.BytesIO()
    torch.save(model.state_dict(), buffer)
    return buffer.getbuffer().nbytes


@torch.no_grad()
def forward_ms(models: dict, batch: torch.Tensor, repeats: int = 30) -> dict[str, float]:
    """Median forward time per model. Runs alternate between models so background load hits both."""
    times: dict[str, list[float]] = {name: [] for name in models}
    for i in range(repeats + 3):
        for name, model in models.items():
            started = time.perf_counter()
            model(batch)
            if i >= 3:                        # the first rounds are warm-up
                times[name].append((time.perf_counter() - started) * 1000)
    return {name: statistics.median(t) for name, t in times.items()}


def main() -> None:
    print(f"machine load (1-minute average) {os.getloadavg()[0]:.2f} on {os.cpu_count()} cores; "
          f"torch threads {torch.get_num_threads()}")
    fp32, tokenizer = load()
    int8 = torch.ao.quantization.quantize_dynamic(copy.deepcopy(fp32), {torch.nn.Linear}, dtype=torch.qint8)
    ids = validation_ids(tokenizer)
    batch = ids[:128].unsqueeze(0)

    models = {"fp32": fp32, "int8": int8}
    speed = forward_ms(models, batch)
    results = {}
    for name, model in models.items():
        ppl, preds = evaluate(model, ids)
        results[name] = preds
        print(f"{name}: size {size_on_disk(model) / 1024:5.0f} KiB   forward (1 x 128 tokens) "
              f"{speed[name]:6.2f} ms   validation perplexity {ppl:.4f}")
    same = sum(a == b for a, b in zip(results["fp32"], results["int8"]))
    print(f"top-1 next-token agreement: {same} of {len(results['fp32'])} positions "
          f"({same / len(results['fp32']):.2%})")

    prompts = ["Customer (Ben): How much does the Team plan cost?\nAgent (Dara):",
               "Customer (Mo): I forgot my password and the reset link expired.\nAgent (Lena):",
               "Customer (Ana): Can I export my tasks to CSV?\nAgent (Kofi):"]
    params = SamplingParams(max_new_tokens=30, temperature=0.0, stop=["\n"])
    for prompt in prompts:
        a = generate(fp32, tokenizer, prompt, params).text
        b = generate(int8, tokenizer, prompt, params).text
        print(f"greedy identical: {a == b!s:<5} | fp32: {a.strip()[:70]}")
        if a != b:
            print(f"                     | int8: {b.strip()[:70]}")


if __name__ == "__main__":
    main()

Code explained

  • In simple words: make a half-size copy of TinyLM and check, token by token, whether it still predicts the same things.
  • What happens:
    • validation_ids: the held-out tail of the corpus, tokenized.
    • evaluate: perplexity (the exponential of the average next-token loss; lower means the text looks more predictable to the model, Module 3) over non-overlapping 128-token windows, plus the model's top prediction at every position.
    • size_on_disk: the bytes torch.save writes for the state dict.
    • forward_ms: median forward time, alternating between the two models so background load hits both equally.
    • main: torch.ao.quantization.quantize_dynamic converts the linear layers of a deep copy. Then it compares everything and runs three greedy generations with each model. The script silences a DeprecationWarning: torch.ao.quantization announces its own removal (the message says 2.10) and points to the separate torchao package's quantize_ API. It still works in torch 2.14.0, which is all a one-line demonstration needs; for new production code, use torchao or your serving engine's quantization.
  • Comes out:

    text
    machine load (1-minute average) 0.22 on 2 cores; torch threads 1
    fp32: size  4203 KiB   forward (1 x 128 tokens)   6.11 ms   validation perplexity 1.6531
    int8: size  2167 KiB   forward (1 x 128 tokens)   4.62 ms   validation perplexity 1.6535
    top-1 next-token agreement: 8180 of 8243 positions (99.24%)
    greedy identical: True  | fp32: Team costs 12 USD per user per month, or 10 USD billed annually. Busin
    greedy identical: True  | fp32: Reset links expire after 30 minutes and can be used once. Please reque
    greedy identical: True  | fp32: Sorry about the duplicate charge. Duplicate charges are always refunde
    

    The file halves (4,203 KiB to 2,167 KiB; the embedding table is not a linear layer, so it stays fp32). Perplexity moves from 1.6531 to 1.6535, a 0.02 percent change. The top prediction changes at 63 of 8,243 positions (0.76 percent), and all three greedy replies are identical. The forward pass got about a quarter faster here (6.1 to 4.6 ms), but at this tiny size and on a shared machine that number moves from run to run (an earlier run gave 4.0 to 3.3 ms); do not generalize it.

    What does transfer: quantization error is small on average and concentrated in a few positions, so a perplexity or average-score check can look perfect while specific outputs change. Check both an aggregate metric and a set of exact outputs that matter to you (here, the three replies). For real models, published results are the guide: 8-bit weights are usually close to lossless, 4-bit costs more and varies by method and model, and small models suffer more than large ones. Run your Module 10 eval suite on the quantized model before you ship it.

SituationUse thisWhy
Memory is tight and quality must not move8-bit weightsNear-lossless in most published results; measure anyway
You need more concurrent users or a smaller GPU4-bit weights (AWQ, GPTQ, GGUF Q4 variants, or the model's shipped format)Quarter-size weights; quality cost varies, so run the eval suite
The model ships in a quantized format (gpt-oss MXFP4)Use it as shippedThe model was trained or calibrated for it
CPU or laptop inferencellama.cpp GGUF quantizationsBuilt for exactly that

Throughput, batching, and utilization

Decoding one token for one user reads every weight from memory to do a small amount of math. Decoding one token for 32 users reads the same weights once and does 32 times the math. That is why batching raises throughput (tokens per second for the whole server) far more than it raises latency (seconds per request for each user), up to the point where the math, not memory, becomes the limit.

The script runs TinyLM on batches of 1 to 32 requests, each a 64-token prompt followed by 32 greedy decode steps with the KV cache, exactly the way a serving engine steps a batch.

examples/m13_batching.py

python
"""Throughput vs batch size on TinyLM: decode many requests together, measure the tradeoff.

Each request is a 64-token prompt from the held-out corpus tail followed by
32 greedy decode steps with the KV cache. A batch of B requests runs those
steps together, the way a serving engine batches concurrent users.
"""
from __future__ import annotations

import os
import statistics
import time

import torch

from supportdesk.data import DATA_DIR
from supportdesk.tinylm import load

torch.manual_seed(0)
torch.set_num_threads(1)          # one thread: steadier timings on a shared 2-core machine
PROMPT, DECODE = 64, 32


@torch.no_grad()
def run_batch(model, prompts: torch.Tensor) -> tuple[float, float]:
    """Return (prefill seconds, decode seconds) for one batch of prompts."""
    caches = [dict() for _ in model.blocks]
    started = time.perf_counter()
    logits = model(prompts, caches)[:, -1]
    prefill = time.perf_counter() - started
    started = time.perf_counter()
    for step in range(DECODE):
        next_ids = logits.argmax(-1, keepdim=True)
        logits = model(next_ids, caches, start_pos=PROMPT + step)[:, -1]
    return prefill, time.perf_counter() - started


def main() -> None:
    print(f"machine load (1-minute average) {os.getloadavg()[0]:.2f} on {os.cpu_count()} cores; "
          f"torch threads {torch.get_num_threads()}")
    model, tokenizer = load()
    ids = torch.tensor(tokenizer.encode((DATA_DIR / "corpus.txt").read_text(encoding="utf-8")).ids)
    tail = ids[int(len(ids) * 0.95):]
    pool = torch.stack([tail[i * PROMPT:(i + 1) * PROMPT] for i in range(32)])
    run_batch(model, pool[:4])  # warm-up

    print(f"{'batch':>5} {'decode tok/s':>13} {'ms per step':>12} {'request latency ms':>19} {'prefill ms':>11}")
    rows = []
    for batch in (1, 2, 4, 8, 16, 32):
        trials = [run_batch(model, pool[:batch]) for _ in range(5)]
        prefill = statistics.median(t[0] for t in trials)
        decode = statistics.median(t[1] for t in trials)
        tps = batch * DECODE / decode
        latency = (prefill + decode) * 1000
        rows.append((batch, tps, latency))
        print(f"{batch:>5} {tps:>13.0f} {decode / DECODE * 1000:>12.2f} {latency:>19.0f} {prefill * 1000:>11.1f}")
    base_tps, base_latency = rows[0][1], rows[0][2]
    print("\nrelative to batch 1:")
    for batch, tps, latency in rows:
        print(f"  batch {batch:>2}: throughput x{tps / base_tps:5.1f}, each request waits x{latency / base_latency:4.1f}")


if __name__ == "__main__":
    main()

Code explained

  • In simple words: serve 1, 2, 4, up to 32 customers at once and time how the whole kitchen and each customer fare.
  • What happens: run_batch prefills all prompts in one forward pass, then runs 32 decode steps where each step feeds one new token per request through the model, reusing the caches. main builds 32 prompts from the held-out corpus, warms up, and for each batch size takes the median of 5 runs. Throughput is batch x 32 tokens / decode time; request latency is prefill plus decode.
  • Comes out:

    text
    machine load (1-minute average) 0.28 on 2 cores; torch threads 1
    batch  decode tok/s  ms per step  request latency ms  prefill ms
        1          1039         0.96                  34         2.7
        2          1999         1.00                  36         4.3
        4          3042         1.31                  49         7.2
        8          4565         1.75                  69        13.0
       16          6098         2.62                 110        25.8
       32          7087         4.52                 194        49.2
    
    relative to batch 1:
      batch  1: throughput x  1.0, each request waits x 1.0
      batch  2: throughput x  1.9, each request waits x 1.1
      batch  4: throughput x  2.9, each request waits x 1.5
      batch  8: throughput x  4.4, each request waits x 2.1
      batch 16: throughput x  5.9, each request waits x 3.3
      batch 32: throughput x  6.8, each request waits x 5.8
    

    Going from 1 to 32 concurrent requests multiplies throughput by about 7 while each request waits nearly 6 times longer. On a GPU the curve is much flatter at the start (small batches are nearly free) and throughput keeps rising for longer, but the tradeoff has the same shape: batch size is a dial between cost per token and latency per user, and a serving engine lets you cap it.

Utilization is the share of time the hardware does useful work. A GPU you rent by the hour costs the same whether it serves 32 requests at once or sits idle at 3 a.m. That fact, not the speed of the GPU, decides most self-hosting economics, as the next section shows.

A real diagnosis from writing this module. The first runs of these timing scripts were wildly inconsistent: sometimes a decode step took under 1 ms, sometimes over 100 ms. The trace pointed to threads: by default PyTorch uses one thread per core, and on a 2-core machine that other processes were also using, the two threads kept waiting for each other. Here is the check, run once on a quiet machine and once with two busy background processes:

python
"""Diagnose slow decoding on a shared machine: time TinyLM decode steps with 1 and with 2 threads."""
import os
import statistics
import time

import torch

from supportdesk.tinylm import load

model, _ = load()
x = torch.randint(0, 1503, (1, 64), generator=torch.Generator().manual_seed(0))
print(f"machine load (1-minute average) {os.getloadavg()[0]:.2f} on {os.cpu_count()} cores")
results: dict[int, list[float]] = {1: [], 2: []}
with torch.no_grad():
    for _ in range(5):                                   # alternate settings so background load hits both
        for threads in (2, 1):
            torch.set_num_threads(threads)
            model(x)                                     # warm-up after switching
            started = time.perf_counter()
            for _ in range(100):
                model(x[:, :1])
            results[threads].append((time.perf_counter() - started) / 100 * 1000)
for threads, runs in results.items():
    print(f"{threads} thread(s): median {statistics.median(runs):.2f} ms per decode step, "
          f"worst {max(runs):.2f} ms (5 rounds of 100 steps)")

Code explained

  • In simple words: time the same decode step with one thread and with two, taking turns so background load hits both.
  • What happens: the script alternates 5 rounds of 100 single-token forward passes per setting and reports the median and the worst round, next to the machine's 1-minute load average.
  • Comes out: first on the machine as it was, then with two CPU-burning processes started in the background (python3 -c "while True: pass" &, twice, stopped afterwards by PID):

    text
    machine load (1-minute average) 0.39 on 2 cores
    1 thread(s): median 0.70 ms per decode step, worst 0.84 ms (5 rounds of 100 steps)
    2 thread(s): median 0.67 ms per decode step, worst 0.69 ms (5 rounds of 100 steps)
    

    text
    machine load (1-minute average) 2.40 on 2 cores
    1 thread(s): median 2.33 ms per decode step, worst 2.57 ms (5 rounds of 100 steps)
    2 thread(s): median 134.99 ms per decode step, worst 162.71 ms (5 rounds of 100 steps)
    

    With spare cores, the two settings are close. With every core busy, two threads took about 135 ms per step against 2.3 ms for one thread, roughly 58 times slower, because each step waits for a thread that the operating system has parked. The same thing happens on real inference servers when too many workers share too few cores. The fix is to set thread counts explicitly (torch.set_num_threads, OMP_NUM_THREADS, or the engine's own setting) and always report load next to a timing.

The real total cost vs API pricing

Now the question finance will ask: should Brightlane rent GPUs instead of paying per token? The script measures real input sizes (triage prompt plus draft prompt with three retrieved help-center articles, for all 72 tickets), prices them with supportdesk.pricing, and compares with renting GPUs.

GPU prices are on-demand prices per GPU-hour read on 21 September 2026: RunPod lists the L4 at 0.49 USD, the A100 SXM 80 GB at 1.59 USD, and the H100 SXM at 3.49 USD (runpod.io/pricing); Lambda lists the H100 SXM at 3.99 USD and the A100 SXM 80 GB at 2.79 USD (lambda.ai/pricing). Prices change monthly; verify them.

One correction to the shared price table first. supportdesk/pricing.py bills cached input tokens for Groq's gpt-oss models at the full input price, but Groq's prompt-caching page says "There is a 50% discount for cached input tokens", and that "the prompt caching discount does not stack with the batch discount" (Groq docs: prompt caching, read 21 September 2026). The script corrects this locally instead of editing the shared file; the difference is shown next to the original.

examples/m13_tco.py

python
"""Total cost of ownership: hosted API vs a rented GPU, for Brightlane's ticket volume.

Token counts per ticket are measured on the real dataset. API prices come from
supportdesk.pricing. GPU prices and every operating assumption are labeled
below; replace them with your own quotes and, above all, with throughput you
measured on your own serving engine.
"""
from __future__ import annotations

import statistics
from dataclasses import dataclass

from supportdesk.data import get_article, load_tickets
from supportdesk.kb_search import KBSearch
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, Price, cost_usd
from supportdesk.tokens import count_messages

HOURS_PER_MONTH = 730

# GPU on-demand prices, USD per GPU-hour, read on 21 Sep 2026 (verify before use):
#   RunPod secure cloud (runpod.io/pricing): L4 0.49, A100 SXM 80 GB 1.59, H100 SXM 3.49
#   Lambda (lambda.ai/pricing): H100 SXM 3.99, A100 SXM 80 GB 2.79
GPU_PRICES = {"L4 (RunPod)": 0.49, "A100 80GB (RunPod)": 1.59, "H100 (RunPod)": 3.49, "H100 (Lambda)": 3.99}

# Assumptions (ours, for teaching; change them):
TICKETS_PER_MONTH = 30_000
OUTPUT_TOKENS_PER_TICKET = 350        # triage JSON (about 60) + draft reply (about 290)
OPS_HOURS_PER_MONTH = 20              # patching, upgrades, on-call for the serving stack
OPS_RATE_USD = 100                    # loaded cost of an engineer hour
REPLICAS = 2                          # one GPU is a single point of failure

# supportdesk.pricing bills Groq's cached input at the full input price. Groq's prompt-caching
# page (console.groq.com/docs/prompt-caching, read 21 Sep 2026) lists a 50 percent discount on
# cached input tokens for gpt-oss models, which does not stack with the 50 percent batch discount.
# We correct it here instead of editing the shared price table.
GROQ_CACHED = {m: Price(PRICES[m].input, PRICES[m].output, PRICES[m].input * 0.5)
               for m in ("openai/gpt-oss-20b", "openai/gpt-oss-120b")}


def groq_cost(usage: Usage, model: str, batch: bool = False) -> float:
    """Cost with Groq's cached-input discount; in batch mode the cache discount does not stack."""
    if batch:
        return cost_usd(Usage(usage.input_tokens, usage.output_tokens), model, batch=True)
    p = GROQ_CACHED[model]
    uncached = usage.input_tokens - usage.cached_tokens
    return (uncached * p.input + usage.cached_tokens * p.cached_input + usage.output_tokens * p.output) / 1e6


def measured_input_tokens() -> list[int]:
    """Input tokens per ticket for triage + draft, built from the real tickets and KB articles."""
    search = KBSearch()
    system_triage = "Classify the Brightlane support ticket. Reply with JSON: category, priority, language, summary, needs_human."
    system_draft = "Draft a reply for a human agent to review. Use only the articles below and cite their ids."
    totals = []
    for t in load_tickets():
        triage = [{"role": "system", "content": system_triage}, {"role": "user", "content": t.text}]
        articles = "\n\n".join(f"[{h.article_id}]\n{get_article(h.article_id).body}" for h in search.search(t.text, k=3))
        draft = [{"role": "system", "content": system_draft + "\n\n" + articles}, {"role": "user", "content": t.text}]
        totals.append(count_messages(triage) + count_messages(draft))
    return totals


@dataclass
class SelfHost:
    gpu: str
    tokens_per_second: float   # sustained output tokens/s for the whole GPU at your latency target
    utilization: float         # share of the month the GPU does useful work

    def monthly_usd(self) -> float:
        return GPU_PRICES[self.gpu] * HOURS_PER_MONTH * REPLICAS + OPS_HOURS_PER_MONTH * OPS_RATE_USD

    def capacity_tickets(self) -> float:
        return self.tokens_per_second * 3600 * HOURS_PER_MONTH * self.utilization / OUTPUT_TOKENS_PER_TICKET * REPLICAS


def main() -> None:
    inputs = measured_input_tokens()
    mean_in = statistics.mean(inputs)
    print(f"measured input tokens per ticket (o200k estimate, n={len(inputs)}): "
          f"mean {mean_in:.0f}, median {statistics.median(inputs):.0f}, max {max(inputs)}")
    usage = Usage(input_tokens=round(mean_in), output_tokens=OUTPUT_TOKENS_PER_TICKET)

    print(f"\nHosted API at {TICKETS_PER_MONTH:,} tickets per month:")
    for model in ("openai/gpt-oss-20b", "openai/gpt-oss-120b", "gemini-3.5-flash-lite", "gemini-3.5-flash"):
        per_ticket = cost_usd(usage, model)
        print(f"  {model:<24} {per_ticket * 1000:6.3f} USD per 1k tickets   {per_ticket * TICKETS_PER_MONTH:9.2f} USD per month")

    cached = Usage(round(mean_in), OUTPUT_TOKENS_PER_TICKET, cached_tokens=round(mean_in * 0.9))
    print("\nIf 90% of input tokens hit the prompt cache (USD per 1k tickets):")
    for model in ("openai/gpt-oss-20b", "openai/gpt-oss-120b"):
        print(f"  {model:<24} pricing.py {cost_usd(cached, model) * 1000:.3f}   with Groq's 50% cached rate "
              f"{groq_cost(cached, model) * 1000:.3f}   batch (no stacking) {groq_cost(cached, model, batch=True) * 1000:.3f}")

    print(f"\nSelf-hosted, {REPLICAS} replicas, plus {OPS_HOURS_PER_MONTH} ops hours at {OPS_RATE_USD} USD:")
    for gpu in GPU_PRICES:
        fixed = SelfHost(gpu, 1, 1).monthly_usd()
        print(f"  {gpu:<20} {fixed:8.2f} USD per month, whether you send 0 tickets or a million")

    print("\nBreak-even: the volume where self-hosting the same open model costs what the API costs")
    for gpu, model in (("L4 (RunPod)", "openai/gpt-oss-20b"), ("H100 (RunPod)", "openai/gpt-oss-120b")):
        api = cost_usd(usage, model)
        volume = SelfHost(gpu, 1, 1).monthly_usd() / api
        tps = volume * OUTPUT_TOKENS_PER_TICKET / (HOURS_PER_MONTH * 3600) / REPLICAS
        print(f"  {model:<20} on {gpu:<14} {volume:>12,.0f} tickets per month, "
              f"which needs {tps:,.0f} output tok/s per GPU around the clock")

    print("\nWhat one self-hosted ticket costs, by sustained throughput and utilization (L4 pair):")
    print(f"  {'tok/s per GPU':>13} " + " ".join(f"{'util ' + format(u, '.0%'):>12}" for u in (0.1, 0.3, 0.6)))
    for tps in (100, 300, 1000):
        cells = []
        for u in (0.1, 0.3, 0.6):
            host = SelfHost("L4 (RunPod)", tps, u)
            cells.append(f"{host.monthly_usd() / host.capacity_tickets() * 1000:>9.3f}/1k")
        print(f"  {tps:>13} " + " ".join(f"{c:>12}" for c in cells))
    need = TICKETS_PER_MONTH * OUTPUT_TOKENS_PER_TICKET / (HOURS_PER_MONTH * 3600)
    print(f"\nBrightlane's average load is only {need:.1f} output tokens/s: the GPUs would sit mostly idle.")


if __name__ == "__main__":
    main()

Code explained

  • In simple words: a spreadsheet in code that puts the per-token bill next to the rent-a-GPU bill for the same monthly work.
  • What happens:
    • The constants at the top are labeled assumptions: 30,000 tickets a month, 350 output tokens per ticket, 20 hours a month of engineering time for the serving stack at 100 USD an hour, and 2 replicas because one GPU is a single point of failure.
    • GROQ_CACHED and groq_cost: Groq's real cached-input discount, with batch mode getting the batch discount and no cache discount on top.
    • measured_input_tokens(): builds both prompts for every real ticket (with KBSearch picking three articles) and counts them with count_messages. These are o200k estimates; each provider's own tokenizer will differ a little (Module 2).
    • SelfHost: the monthly bill (GPU hours times replicas plus ops time) and the monthly capacity (sustained output tokens per second times utilization).
    • main: API cost per model, the cached variants, the fixed self-hosting bill, the break-even volume, and a grid of cost per 1,000 tickets across throughput and utilization.
  • Comes out:

SituationUse thisWhy
Low or spiky volume (Brightlane today)Hosted APIYou pay only for tokens; idle GPUs cost the same as busy ones
Very high, steady volume on a task an open model handlesSelf-hosting, after measuring sustained throughputBreak-even arrives only when the GPUs stay busy
Privacy rules forbid sending tickets out, at any volumeSelf-hosting, priced honestlyCost stops being the deciding factor; budget for it
Work that can wait hoursBatch APIsHalf price at most providers, no infrastructure