Part A: Tokenization in practice
By the end of this module, you'll have:
- Measured, with five real tokenizers, how many tokens Brightlane's tickets, invoice ids, JSON, code, and non-English messages cost, and why the same request can cost 1.6 to 7.7 times more in Hindi than in English.
- Read
supportdesk/tokens.pyandsupportdesk/pricing.pyline by line, and used them to count a full triage request before sending it. - Built a context budget calculator for the Brightlane assistant, and watched what really happens at and past the limit, both in TinyLM and with a provider's overflow error.
- A needle-in-a-haystack harness for position effects that runs against any model through
llm.chat, and that you have proven can detect a blind spot. - Four context management strategies (drop-oldest, sliding window, importance-ranked, summary compaction) compared on a long support chat with an automatic check of which key facts survive.
- A cost model for the whole triage workload: cost per task under every model in
PRICES, output versus input share, reasoning tokens, batch pricing, and a measured cache-aware prompt layout.
Prerequisites: Module 1 (what a model is, the supportdesk repository, llm.chat, ScriptedLLM, TinyLM). Working Python. No API key is needed for any measured output in this module.
Where we are: Module 1 showed that a model reads and writes tokens, one prediction at a time, inside a fixed context window. This module turns those words into numbers you can budget: how many tokens your text becomes, how many fit, which ones the model actually uses, and what each one costs.
How this module is organized
| Part | What it covers |
|---|---|
| Setup | Working copy, how the examples run |
| Part A: Tokenization in practice | Subwords and BPE, tokenizer families, fertility on non-English text, code, numbers and ids, artifacts behind counting and arithmetic failures, counting before you send (tokens.py) |
| Part B: The context window | What fills it, the budget equation, TinyLM and providers at the limit, position effects, a needle-in-a-haystack harness, effective versus advertised context |
| Part C: Managing context | Drop-oldest, sliding window, importance-ranked truncation, summary compaction, what to preserve and when, long context versus retrieval |
| Part D: Cost mechanics | pricing.py, input and output prices, reasoning tokens, prompt caching and cache-aware layout, batch tiers, cost per task |
| Module Lab | A token-aware, budget-checked, cost-reported triage run with a compaction guard |
Setup
The examples run in order and each one is a complete script. They live in examples/ and import the canonical supportdesk package plus a few small Module 2 helpers (m02_prompts.py, m02_conversation.py, m02_context.py) that sit next to them. Run every command from the repository root.
cp -r supportdesk ~/work/m02 && cd ~/work/m02
source ~/venv/bin/activate # Python 3.11 with the packages from requirements.txt
pip install -r requirements.txt # tiktoken==0.14.0, tokenizers==0.23.2, torch==2.14.0, openai==3.16.2, ...
PYTHONPATH=. python examples/m02_tokenizers.py
PYTHONPATH=. python -m pytest -q tests/test_m02_tokens_context_cost.pyCode explained
- In simple words: make your own copy of the repository, activate the environment, and check that the first example and the module's tests run.
- What happens:
cp -rcopies the canonical repository so nothing you do touches the original.PYTHONPATH=.lets Python find thesupportdeskpackage from the repository root. Scripts inexamples/also find their Module 2 neighbors (m02_prompts.pyand friends) because Python puts a script's own folder on the import path. The tokenizer files ship invendor/tokenizers/, so nothing downloads. - Comes out: the first example prints token pieces (shown in Part A), and pytest ends with
9 passed. If you seeModuleNotFoundError: supportdesk, you forgotPYTHONPATH=.or ran from another folder.
Two files carry the triage request and the long chat that later examples reuse. The triage prompt is a working draft: Module 4 teaches how to design prompts properly; here it is simply a realistic payload to count and price.
# examples/m02_prompts.py
"""Module 2: the triage request used by this module's examples (Module 4 teaches prompt design properly)."""
import json
from supportdesk.data import CATEGORIES, PRIORITIES, Ticket
from supportdesk.schemas import triage_json_schema
TRIAGE_INSTRUCTIONS = f"""You are the triage assistant for Brightlane's support desk.
Brightlane is a project-management SaaS. Plans: Free, Team (12 USD per user per month),
Business (24 USD per user per month), Enterprise (custom pricing).
Read one customer ticket and classify it. Reply with one JSON object and nothing else.
Categories: {", ".join(CATEGORIES)}.
- billing: charges, invoices, prices, payment methods.
- cancellation: cancelling a plan or asking for a refund because they are leaving.
- account_access: login, password, SSO, locked accounts, 2FA.
- bug: something that used to work is broken or erroring.
- how_to: a question about using a feature that exists.
- feature_request: asking for something Brightlane does not do yet.
Priorities: {", ".join(PRIORITIES)}.
- urgent: many users blocked or a security risk right now.
- high: one user blocked, or money charged wrongly.
- normal: needs an answer, nobody is blocked.
- low: a question or an idea.
Write the summary in English even when the ticket is in another language.
Set needs_human to true for refunds, account changes, legal, or security issues.
JSON schema of the reply:
{json.dumps(triage_json_schema(), separators=(",", ":"))}"""
EXAMPLES = [
("Subject: Invoice address\n\nCan you add our VAT number to future invoices?",
'{"category":"billing","priority":"low","language":"en","summary":"Customer wants a VAT number on future invoices.","needs_human":false}'),
("Subject: SSO broken\n\nNobody in our company can log in through Okta since this morning.",
'{"category":"account_access","priority":"urgent","language":"en","summary":"Company-wide SSO login failure through Okta.","needs_human":true}'),
]
def triage_messages(ticket: Ticket) -> list[dict]:
"""Stable parts first (instructions, examples), the ticket last: a cache-friendly layout."""
messages = [{"role": "system", "content": TRIAGE_INSTRUCTIONS}]
for user, assistant in EXAMPLES:
messages += [{"role": "user", "content": user}, {"role": "assistant", "content": assistant}]
messages.append({"role": "user", "content": f"Customer tier: {ticket.customer_tier}\n{ticket.text}"})
return messagesCode explained
- In simple words: this is the request the Brightlane assistant sends to classify one ticket: fixed instructions, two worked examples, then the ticket.
- What happens:
TRIAGE_INSTRUCTIONSholds the plan prices, category and priority definitions, and the JSON schema fromschemas.triage_json_schema()(introduced in Module 6; here we only count it).EXAMPLESare two question and answer pairs, a technique called few-shot prompting (Module 4).triage_messages()builds the chat messages with every stable part first and the ticket last. Part D shows, with measurements, why that order matters for cost. - Comes out: nothing when run alone; later scripts import it.
# examples/m02_conversation.py
"""Module 2: a long multi-turn Brightlane support chat and a checker for the facts that must survive."""
import re
SYSTEM = ("You are Brightlane's support assistant. Answer from the help center, be concise, "
"and hand refunds and account changes to a human agent.")
TURNS = [
("user", "Hi, we were charged twice this month, invoice INV-2026-004512. Please refund the duplicate charge."),
("assistant", "Sorry about that. I can pass this to our billing team. Which plan is the workspace on, and how many seats?"),
("user", "We are on the Business plan with 40 seats, paid monthly."),
("assistant", "Thanks. A billing agent will review it; duplicate charges go back to the original card in 5 to 10 business days."),
("user", "While I have you: Slack notifications stopped arriving in our #ops channel yesterday."),
("assistant", "Please open Settings > Integrations > Slack and check that the channel is still connected. Reconnecting fixes most cases."),
("user", "It says connected, but nothing arrives. We did rename the channel last week."),
("assistant", "Renaming breaks the link. Choose the channel again in the Slack integration settings and send a test message."),
("user", "That worked, thank you. Another question: can I export a board to CSV with the custom fields?"),
("assistant", "Yes. Open the board, choose More > Export > CSV, and tick Include custom fields. Large boards arrive by email."),
("user", "The export link in the email says expired."),
("assistant", "Export links last 24 hours. Run the export again and download it the same day."),
("user", "OK. Our designers also want automations that move cards when a checklist is complete. Possible?"),
("assistant", "Yes: Automations > New rule > When checklist completed > Move card to column. Rules run within a minute."),
("user", "The rule ran twice on one card this morning."),
("assistant", "That can happen when two rules match the same card. Check for an older rule with the same trigger and disable one."),
("user", "Found it, there was an old copy. Does the mobile app support offline mode?"),
("assistant", "The mobile app caches boards you opened recently and syncs your edits when you reconnect."),
("user", "Great. Last thing on mobile: push notifications are delayed by an hour on Android."),
("assistant", "Android battery optimization can delay them. Exclude the Brightlane app from battery optimization in system settings."),
("user", "Done, they arrive now. Can guests see private boards?"),
("assistant", "No. Guests only see boards they are explicitly invited to, and private boards stay hidden from them."),
("user", "Good to know. So, where are we on my original request?"),
]
FACTS = {
"invoice id": lambda text: "INV-2026-004512" in text,
"plan and seats": lambda text: "Business" in text and re.search(r"\b40 seats\b", text) is not None,
# "refund" and "duplicate" close together, so the system prompt's "refunds" alone does not count
"customer's ask": lambda text: re.search(r"refund.{0,40}duplicate|duplicate.{0,40}refund", text, re.I) is not None,
}
def conversation() -> list[dict]:
"""System message plus the 23 turns above, as chat messages."""
return [{"role": "system", "content": SYSTEM}] + [{"role": r, "content": c} for r, c in TURNS]
def facts_kept(messages: list[dict]) -> dict[str, bool]:
"""Which key facts are still present anywhere in the messages we would send."""
text = "\n".join(m["content"] for m in messages)
return {name: check(text) for name, check in FACTS.items()}Code explained
- In simple words: a 23-turn support chat in which the important facts come early and the question that needs them comes last, plus a checker that says whether those facts are still in whatever we send.
- What happens: the customer gives an invoice id, the plan and seat count, and the request (refund the duplicate charge) in the first three turns. Then the chat wanders through Slack, exports, automations, and mobile, and finally asks "where are we on my original request?".
FACTSdefines three checks by exact text or regular expression;facts_kept()joins the message contents and runs each check. None of the later turns repeat those facts, so a strategy only passes if it really kept them. - Comes out: nothing when run alone. The full chat is 24 messages and 558 tokens (measured in Part C).
Part A: Tokenization in practice
What a token is
A token is the unit a model reads and writes: a chunk of text that has an integer id in the model's vocabulary (its fixed list of known chunks). A tokenizer is the program that turns text into token ids and back. Frequent words usually become one token; rare words, numbers, and ids break into several subword pieces. Every price, limit, and speed figure you will see is quoted in tokens, not characters or words.
Let's look at one real Brightlane ticket through five tokenizers: three OpenAI families that ship with tiktoken (p50k_base from the GPT-3 era, cl100k_base from GPT-3.5 and GPT-4, o200k_base from GPT-4o onward), an older Anthropic tokenizer (claude-legacy), and TinyLM's own tokenizer from Module 1.
# examples/m02_tokenizers.py
"""Module 2: see one Brightlane ticket through five tokenizers, then count the whole dataset."""
from tokenizers import Tokenizer
from supportdesk.data import load_tickets
from supportdesk.tinylm import MODEL_DIR
from supportdesk.tokens import ENCODINGS, count_tokens, pieces
tiny = Tokenizer.from_file(str(MODEL_DIR / "tokenizer.json")) # TinyLM's own byte-level BPE
def tiny_pieces(text: str) -> list[str]:
return [tiny.decode([i]) for i in tiny.encode(text).ids]
tickets = load_tickets()
first = tickets[0]
print(f"{first.id}: {first.body}\n")
for name in ENCODINGS:
p = pieces(first.body, name)
print(f"{name:13} {len(p):3} tokens {'|'.join(p)}")
p = tiny_pieces(first.body)
print(f"{'tinylm':13} {len(p):3} tokens {'|'.join(p)}")
print("\nAll 72 tickets (subject + body):")
chars = sum(len(t.text) for t in tickets)
words = sum(len(t.text.split()) for t in tickets)
print(f"{'characters':13} {chars:6} words {words}")
for name in ENCODINGS:
n = sum(count_tokens(t.text, name) for t in tickets)
print(f"{name:13} {n:6} tokens {chars / n:4.2f} chars/token {n / words:4.2f} tokens/word")
n = sum(len(tiny.encode(t.text).ids) for t in tickets)
print(f"{'tinylm':13} {n:6} tokens {chars / n:4.2f} chars/token {n / words:4.2f} tokens/word")Code explained
- In simple words: we hold the same sentence up to five different "rulers" and see where each one draws its tick marks.
- What happens:
pieces()fromsupportdesk/tokens.pyreturns the text of each token. For TinyLM we loadmodels/tinylm-base/tokenizer.jsonwith thetokenizerslibrary and decode each id alone. The second half sums characters, words, and tokens over all 72 tickets (subject plus body, asTicket.textformats them). - Comes out:
T-1001: Hi, my card was charged 288 USD twice on 3 September for the Team plan (invoice INV-2026-004512). Please refund the duplicate.
p50k_base 33 tokens Hi|,| my| card| was| charged| 288| USD| twice| on| 3| September| for| the| Team| plan| (|inv|oice| INV|-|20|26|-|00|45|12|).| Please| refund| the| duplicate|.
cl100k_base 33 tokens Hi|,| my| card| was| charged| |288| USD| twice| on| |3| September| for| the| Team| plan| (|invoice| INV|-|202|6|-|004|512|).| Please| refund| the| duplicate|.
o200k_base 33 tokens Hi|,| my| card| was| charged| |288| USD| twice| on| |3| September| for| the| Team| plan| (|invoice| INV|-|202|6|-|004|512|).| Please| refund| the| duplicate|.
claude-legacy 33 tokens Hi|,| my| card| was| charged| 288| USD| twice| on| 3| September| for| the| Team| plan| (|invoice| IN|V|-|20|26|-|00|45|12|).| Please| refund| the| duplicate|.
tinylm 35 tokens H|i|,| my| card| was| charged| 2|8|8| USD| twice| on| 3| S|ep|tember| for| the| Team| plan| (|in|voice| INV|-|2026|-|004512|).| Please| refund| the| duplicate|.
All 72 tickets (subject + body):
characters 7498 words 1266
p50k_base 2004 tokens 3.74 chars/token 1.58 tokens/word
cl100k_base 1818 tokens 4.12 chars/token 1.44 tokens/word
o200k_base 1725 tokens 4.35 chars/token 1.36 tokens/word
claude-legacy 1914 tokens 3.92 chars/token 1.51 tokens/word
tinylm 3149 tokens 2.38 chars/token 2.49 tokens/wordRead the pieces first. Common words (charged, refund, duplicate) are single tokens in every tokenizer, and most tokens carry their leading space ( card, not card). The invoice id INV-2026-004512 falls apart into 5 to 9 pieces. TinyLM, trained only on Brightlane text, splits Hi into letters and September into three pieces, but has learned 2026 and even 004512 as single tokens because its corpus repeats them. On the whole dataset the newest OpenAI tokenizer needs 14% fewer tokens than the oldest (1,725 versus 2,004), and TinyLM's tiny vocabulary needs 83% more than o200k_base. The rule of thumb "about 4 characters per token for English" holds here (3.7 to 4.4), but only for English prose.
How subwords form: train a tiny BPE
Almost every modern tokenizer is built with byte pair encoding (BPE). Training starts from single bytes (256 symbols, enough to spell any text in any language) and repeatedly merges the most frequent adjacent pair into a new symbol, until the vocabulary reaches a target size. Frequent words end up as one token; rare strings stay split. Watching this happen explains most of the surprises in this part, so let's train a few BPE tokenizers on Brightlane's own corpus.
# examples/m02_train_bpe.py
"""Module 2: train byte-level BPE tokenizers of growing size and watch subword merges form."""
import json
from tokenizers import Tokenizer, decoders, models, pre_tokenizers, trainers
from supportdesk.data import DATA_DIR
corpus = str(DATA_DIR / "corpus.txt")
# "_" in the output marks a space that belongs to the token.
WORDS = [" refund", " duplicate", " Brightlane", " unsubscribed", " INV-2026-004512"]
def train(vocab_size: int) -> Tokenizer:
tok = Tokenizer(models.BPE())
tok.pre_tokenizer = pre_tokenizers.ByteLevel(add_prefix_space=False)
tok.decoder = decoders.ByteLevel()
trainer = trainers.BpeTrainer(vocab_size=vocab_size, show_progress=False,
initial_alphabet=pre_tokenizers.ByteLevel.alphabet())
tok.train([corpus], trainer)
return tok
for size in (256, 280, 400, 1000, 4000):
tok = train(size)
shown = ["|".join(tok.decode([i]).replace(" ", "_") for i in tok.encode(w).ids) for w in WORDS]
print(f"vocab {tok.get_vocab_size():5}: " + " ".join(shown))
big = train(4000)
merges = json.loads(big.to_str())["model"]["merges"]
print(f"\nFirst 12 of {len(merges)} merges (_ marks a leading space):")
print(" " + ", ".join(" + ".join(m).replace("\u0120", "_") for m in merges[:12]))
print(f"Asked for 4000, got {big.get_vocab_size()}: training stops when no pair is left to merge.")Code explained
- In simple words: we build the same kind of tokenizer TinyLM uses, five times with bigger and bigger vocabularies, and watch words fuse from letters into whole tokens.
- What happens:
models.BPE()is the BPE model from thetokenizerslibrary.pre_tokenizers.ByteLevelfirst splits text into words (keeping the leading space with the word) and maps bytes to printable symbols;BpeTrainerlearns merges up tovocab_size. Theinitial_alphabetguarantees all 256 bytes are in the vocabulary, so no text is ever unrepresentable. We print five probe words at each size, then the first merges learned and the final size. This mirrorstrain_tokenizer()inscripts/pretrain_tinylm.py, which built TinyLM's tokenizer. - Comes out (about 4 seconds):
vocab 256: _|r|e|f|u|n|d _|d|u|p|l|i|c|a|t|e _|B|r|i|g|h|t|l|a|n|e _|u|n|s|u|b|s|c|r|i|b|e|d _|I|N|V|-|2|0|2|6|-|0|0|4|5|1|2
vocab 280: _|re|f|u|n|d _|d|u|p|l|i|c|at|e _|B|r|i|g|h|t|l|an|e _|u|n|s|u|b|s|c|r|i|b|e|d _|I|N|V|-|2|0|2|6|-|0|0|4|5|1|2
vocab 400: _refund _d|up|l|ic|at|e _B|ri|g|h|t|lan|e _|un|s|u|b|s|c|ri|b|ed _I|N|V|-|2|0|2|6|-|0|0|4|5|1|2
vocab 1000: _refund _duplicate _Brightlane _un|s|ub|scri|b|ed _I|N|V|-|20|26|-|00|4|5|1|2
vocab 1502: _refund _duplicate _Brightlane _un|s|ub|scri|b|ed _INV|-|2026|-|004512
First 12 of 1246 merges (_ marks a leading space):
e + r, a + n, t + h, i + n, u + s, _ + (, ) + :, t + o, e + n, _ + th, _ + a, to + m
Asked for 4000, got 1502: training stops when no pair is left to merge.At 256 entries every word is spelled byte by byte. The first merges are the corpus's most frequent pairs: e + r, a + n, t + h, and, tellingly, _ + ( and ) + :, because the corpus is full of lines like Customer (Ben):. By 400 entries refund is one token; by 1,000 duplicate and Brightlane are too. unsubscribed never becomes one token: the word is rare in this corpus, so its pieces never win a merge. Finally, we asked for 4,000 entries and got 1,502. BPE stops when every word in the training text is already a single token, and this templated corpus has only about 1,250 distinct pairs worth merging. That is also why TinyLM's tokenizer has 1,503 entries (1,502 plus the <|endoftext|> special token) although its model reserves 2,048 embedding rows (vocab_size in models/tinylm-base/config.json); 545 rows are never used.
The lesson for everything that follows: what a tokenizer merges depends on what it was trained on. A tokenizer trained mostly on English web text learns English words, so other scripts and unusual strings stay expensive.
The canonical token counter: supportdesk/tokens.py
This course counts tokens through one small file. Module 2 introduces it; later modules import it.
<strong>supportdesk/tokens.py</strong>
# supportdesk/tokens.py
"""Count tokens offline with real tokenizers from several model families.
Encodings available without network access:
p50k_base GPT-3 era (about 50k vocabulary)
cl100k_base GPT-3.5 / GPT-4 era (about 100k vocabulary)
o200k_base GPT-4o and later OpenAI models (about 200k vocabulary)
claude-legacy an older Anthropic tokenizer (about 65k vocabulary)
Other providers (Llama, Gemini, Qwen) use their own tokenizers; counts differ
by model, so treat any single count as an estimate for a model it was not built for.
"""
from __future__ import annotations
import os
from functools import lru_cache
from pathlib import Path
VENDOR = Path(__file__).resolve().parents[1] / "vendor" / "tokenizers"
os.environ.setdefault("TIKTOKEN_CACHE_DIR", str(VENDOR))
import tiktoken # noqa: E402 (must import after the cache directory is set)
from tokenizers import Tokenizer # noqa: E402
ENCODINGS = ("p50k_base", "cl100k_base", "o200k_base", "claude-legacy")
DEFAULT_ENCODING = "o200k_base"
@lru_cache(maxsize=None)
def _encoder(name: str):
if name == "claude-legacy":
return Tokenizer.from_file(str(VENDOR / "anthropic_tokenizer.json"))
if name not in ENCODINGS:
raise ValueError(f"Unknown encoding {name!r}. Choose one of {ENCODINGS}.")
return tiktoken.get_encoding(name)
def encode(text: str, encoding: str = DEFAULT_ENCODING) -> list[int]:
enc = _encoder(encoding)
if encoding == "claude-legacy":
return enc.encode(text).ids
return enc.encode(text, disallowed_special=())
def count_tokens(text: str, encoding: str = DEFAULT_ENCODING) -> int:
"""Number of tokens `text` becomes under one tokenizer."""
return len(encode(text, encoding))
def pieces(text: str, encoding: str = DEFAULT_ENCODING) -> list[str]:
"""The text each token stands for, useful for seeing where a tokenizer splits."""
enc = _encoder(encoding)
if encoding == "claude-legacy":
return [enc.decode([i]) for i in enc.encode(text).ids]
return [enc.decode_single_token_bytes(i).decode("utf-8", errors="replace") for i in encode(text, encoding)]
def count_messages(messages: list[dict], encoding: str = DEFAULT_ENCODING, per_message_overhead: int = 4) -> int:
"""Estimate prompt tokens for a chat request.
Each message costs its content tokens plus a few tokens of role and
separator markup; 4 is a common approximation. Providers report the exact
number in the response usage, so use this for budgeting before you send.
"""
total = 0
for m in messages:
content = m.get("content") or ""
if not isinstance(content, str):
content = str(content)
total += count_tokens(content, encoding) + per_message_overhead
return total + 3 # the assistant turn the model is primed to writeCode explained
- In simple words: one place to count tokens offline with several real tokenizers, so every budget and cost estimate in the course uses the same ruler.
- What happens (per part):
- Module docstring and
VENDOR: the tokenizer files live invendor/tokenizers/. SettingTIKTOKEN_CACHE_DIRbeforeimport tiktokenmakes tiktoken load them from disk instead of downloading, which is why the import sits below the environment line. ENCODINGSandDEFAULT_ENCODING: the four names you can pass.o200k_baseis the default because the course's default model,openai/gpt-oss-120b, useso200k_harmony, a superset ofo200k_basewith extra special tokens for its chat format, so content counts match closely._encoder(name): builds each tokenizer once (lru_cacheremembers it) and returns either a tiktokenEncodingor atokenizers.Tokenizerforclaude-legacy. An unknown name raisesValueErrorlisting the valid ones.encode(text, encoding): returns token ids.disallowed_special=()tells tiktoken to treat text like<|endoftext|>inside a ticket as ordinary characters instead of raising an error, which matters when you count untrusted customer text.count_tokens(text, encoding): the length ofencode(). This is the function you will call most.pieces(text, encoding): decodes each id on its own so you can see split points. A single token can be part of a multi-byte character (common in Japanese and Hindi);errors="replace"shows such fragments as the replacement character instead of crashing.count_messages(messages, encoding, per_message_overhead): estimates a chat request: content tokens plus about 4 tokens per message for role markers and separators, plus 3 for the assistant turn the model is primed to write. Non-string content (for example a list of content parts) is converted withstr(), which is only a rough estimate.- Comes out: nothing by itself. The docstring's warning is the important output: counts from one tokenizer are estimates for any model built with another.
Tokenizer differences across model families
Each model family trains its own tokenizer, so the same text has different token counts, and therefore different prices and limits, depending on the model. The table below uses the measured totals from the first example.
| Tokenizer | Vocabulary size | Used by | Tokens for all 72 tickets |
|---|---|---|---|
| p50k_base | 50,281 | GPT-3 era OpenAI models | 2,004 |
| cl100k_base | 100,277 | GPT-3.5, GPT-4 | 1,818 |
| o200k_base | 200,019 | GPT-4o and later; gpt-oss via o200k_harmony | 1,725 |
| claude-legacy | 65,000 | older Anthropic models | 1,914 |
| TinyLM BPE | 1,503 | TinyLM only | 3,149 |
Vocabulary sizes were read from the loaded tokenizers (tiktoken.get_encoding(name).n_vocab, Tokenizer.get_vocab_size()). Two practical consequences:
- Bigger vocabularies usually mean fewer tokens, especially outside English, because more words and scripts earned their own merges. They also mean a bigger embedding table in the model; that trade-off is the model builder's, not yours.
- You cannot count exactly for a model whose tokenizer you do not have. Gemini, Llama, and Qwen use their own tokenizers, and current Claude models do not use
claude-legacy. Providers offer exact counting (Gemini'scountTokens, Anthropic's token counting endpoint) and always report exact usage after a call. Use local counts for budgeting, with a safety margin, and calibrate against reported usage (later in this part).
Fertility: non-English text
Fertility is how many tokens a tokenizer needs per unit of text: per word, per character, or, most usefully for cost, relative to the same meaning in English. High fertility means a message costs more, fills the context faster, and is generated more slowly. The ticket dataset has only 8 non-English tickets, so we also measure a parallel set: the same three support requests written in each language.
# examples/m02_fertility.py
"""Module 2: how many tokens the same support request costs in five languages."""
from tokenizers import Tokenizer
from supportdesk.data import load_tickets
from supportdesk.tinylm import MODEL_DIR
from supportdesk.tokens import ENCODINGS, count_tokens
tiny = Tokenizer.from_file(str(MODEL_DIR / "tokenizer.json"))
NAMES = ENCODINGS + ("tinylm",)
def count(text: str, name: str) -> int:
return len(tiny.encode(text).ids) if name == "tinylm" else count_tokens(text, name)
# The same three requests, written in each language (hand translations for this course).
PARALLEL = {
"en": ["Hi, I was charged twice for the Team plan this month. Please refund the duplicate charge.",
"I forgot my password and the reset link has expired. What can I do?",
"How do I export a board as a CSV file?"],
"es": ["Hola, me cobraron dos veces el plan Team este mes. Por favor, reembolsen el cargo duplicado.",
"Olvidé mi contraseña y el enlace para restablecerla ha caducado. ¿Qué puedo hacer?",
"¿Cómo exporto un tablero como archivo CSV?"],
"de": ["Hallo, mir wurde der Team-Plan diesen Monat zweimal berechnet. Bitte erstatten Sie die doppelte Abbuchung.",
"Ich habe mein Passwort vergessen und der Link zum Zurücksetzen ist abgelaufen. Was kann ich tun?",
"Wie exportiere ich ein Board als CSV-Datei?"],
"ja": ["こんにちは。今月、Teamプランの料金が二重に請求されました。重複した請求分を返金してください。",
"パスワードを忘れてしまい、リセット用のリンクの有効期限が切れています。どうすればよいですか?",
"ボードをCSVファイルとしてエクスポートするにはどうすればよいですか?"],
"hi": ["नमस्ते, इस महीने Team प्लान के लिए मुझसे दो बार शुल्क लिया गया। कृपया दोहरा शुल्क वापस करें।",
"मैं अपना पासवर्ड भूल गया हूँ और रीसेट लिंक की समय सीमा समाप्त हो गई है। मैं क्या करूँ?",
"मैं किसी बोर्ड को CSV फ़ाइल के रूप में कैसे एक्सपोर्ट करूँ?"],
}
print("Parallel set: total tokens for the same 3 requests, and the ratio to English")
print(f"{'lang':5}{'chars':>6}{'bytes':>6}" + "".join(f"{n:>15}" for n in NAMES))
english = {n: sum(count(s, n) for s in PARALLEL["en"]) for n in NAMES}
for lang, texts in PARALLEL.items():
chars = sum(len(s) for s in texts)
nbytes = sum(len(s.encode("utf-8")) for s in texts)
cells = []
for n in NAMES:
total = sum(count(s, n) for s in texts)
cells.append(f"{total:>6} ({total / english[n]:.1f}x)")
print(f"{lang:5}{chars:>6}{nbytes:>6}" + "".join(f"{c:>15}" for c in cells))
print("\nReal tickets: tokens per character by language (o200k_base and tinylm)")
tickets = load_tickets()
for lang in ("en", "es", "de", "ja", "hi"):
group = [t for t in tickets if t.language == lang]
chars = sum(len(t.text) for t in group)
o200 = sum(count(t.text, "o200k_base") for t in group)
tl = sum(count(t.text, "tinylm") for t in group)
words = sum(len(t.text.split()) for t in group)
per_word = "no spaces" if lang == "ja" else f"{o200 / words:.2f}/word"
print(f"{lang}: n={len(group):2} chars={chars:5} o200k={o200:5} ({o200 / chars:.2f}/char, "
f"{per_word}) tinylm={tl:5} ({tl / chars:.2f}/char)")Code explained
- In simple words: we say the same three things in five languages and ask each tokenizer what it charges.
- What happens:
PARALLELholds hand translations of a duplicate-charge complaint, a password reset question, and a CSV export question. For each language and tokenizer we total the tokens and divide by the English total. We also print characters and UTF-8 bytes: Latin letters are 1 byte, accented letters 2, Japanese and Devanagari 3. The second table measures the real tickets by language, per character and per word (Japanese is written without spaces, so "per word" does not apply). - Comes out:
Parallel set: total tokens for the same 3 requests, and the ratio to English
lang chars bytes p50k_base cl100k_base o200k_base claude-legacy tinylm
en 194 194 46 (1.0x) 46 (1.0x) 46 (1.0x) 46 (1.0x) 52 (1.0x)
es 216 222 78 (1.7x) 61 (1.3x) 56 (1.2x) 66 (1.4x) 130 (2.5x)
de 245 246 89 (1.9x) 67 (1.5x) 59 (1.3x) 74 (1.6x) 138 (2.7x)
ja 129 373 163 (3.5x) 113 (2.5x) 84 (1.8x) 115 (2.5x) 369 (7.1x)
hi 237 599 353 (7.7x) 233 (5.1x) 73 (1.6x) 250 (5.4x) 592 (11.4x)
Real tickets: tokens per character by language (o200k_base and tinylm)
en: n=64 chars= 6767 o200k= 1508 (0.22/char, 1.29/word) tinylm= 2442 (0.36/char)
es: n= 2 chars= 250 o200k= 67 (0.27/char, 1.81/word) tinylm= 147 (0.59/char)
de: n= 3 chars= 306 o200k= 71 (0.23/char, 1.65/word) tinylm= 162 (0.53/char)
ja: n= 2 chars= 104 o200k= 61 (0.59/char, no spaces) tinylm= 250 (2.40/char)
hi: n= 1 chars= 71 o200k= 18 (0.25/char, 1.64/word) tinylm= 148 (2.08/char)The same meaning costs 1.2 to 1.9 times more tokens in Spanish and German, 1.8 to 3.5 times more in Japanese, and 1.6 to 7.7 times more in Hindi, depending on the tokenizer. The newest tokenizer (o200k_base) narrows the gap dramatically for Hindi (7.7x down to 1.6x) because its larger vocabulary includes many Devanagari merges; older ones fall back to several byte-level tokens per character. TinyLM, which saw almost no non-English text, needs 11.4 times more tokens for Hindi: it is spelling UTF-8 bytes one or two at a time. The real-ticket table agrees in direction, but with n = 1 to 3 tickets per non-English language it is anecdote, not measurement; the parallel set is the fairer comparison, and it is still only three sentences.
What this means for Brightlane: a Japanese customer's ticket costs about 1.8 times as much to process as the English equivalent on a current tokenizer, fills the context window 1.8 times faster, and a reply in Japanese takes about 1.8 times as many decode steps. If your traffic is multilingual, measure fertility on your own messages with your model's tokenizer before you set budgets.
Fertility: code, numbers, ids, and JSON
Support tickets are not just prose. They contain invoice ids, amounts, URLs, pasted code, and the assistant itself produces JSON.
# examples/m02_hard_strings.py
"""Module 2: code, numbers, ids, and JSON under five tokenizers, plus the pieces behind classic failures."""
import json
from tokenizers import Tokenizer
from supportdesk.schemas import Triage
from supportdesk.tinylm import MODEL_DIR
from supportdesk.tokens import ENCODINGS, count_tokens, encode, pieces
tiny = Tokenizer.from_file(str(MODEL_DIR / "tokenizer.json"))
NAMES = ENCODINGS + ("tinylm",)
def count(text: str, name: str) -> int:
return len(tiny.encode(text).ids) if name == "tinylm" else count_tokens(text, name)
triage = Triage(category="billing", priority="high", language="en",
summary="Customer was charged twice and wants the duplicate refunded.", needs_human=True)
SAMPLES = {
"prose": "Please refund the duplicate charge on our Team plan.",
"invoice ids": "INV-2026-004512, INV-2026-004871, INV-2025-019934",
"amounts": "288.00 USD, 1,234.56 EUR, 0.0075 USD per token",
"long number": "Workspace 81736450921 has 3141592653 events.",
"python code": "if ticket.priority == 'urgent':\n notify(oncall, ticket.id)\n",
"json (pretty)": triage.model_dump_json(indent=2),
"json (compact)": json.dumps(triage.model_dump(), separators=(",", ":")),
"url + email": "https://status.brightlane.example/incidents?id=4411 billing@brightlane.example",
}
print(f"{'sample':15}{'chars':>6}" + "".join(f"{n:>14}" for n in NAMES) + " chars/token (o200k)")
for label, text in SAMPLES.items():
counts = [count(text, n) for n in NAMES]
print(f"{label:15}{len(text):>6}" + "".join(f"{c:>14}" for c in counts) + f" {len(text) / counts[2]:.2f}")
print("\nWhere the pieces fall (o200k_base):")
for text in ["INV-2026-004512", "3141592653", "1,234.56", " notify(oncall, ticket.id)"]:
print(f" {text!r:34} -> {pieces(text)}")
print("\nWhy letter counting, spelling, and arithmetic are hard (o200k_base):")
for text in ["Brightlane", "strawberry", " refund", " unsubscribed", "288 / 2 = 144", "14 * 24 = 336", "12345 + 67890"]:
print(f" {text!r:18} -> {pieces(text)}")
print(" 'r' in 'strawberry':", "strawberry".count("r"), "(the model never sees letters, only the pieces above)")
print("\nSame word, different tokens (o200k_base ids):")
for text in ["refund", " refund", " Refund", " REFUND"]:
print(f" {text!r:10} -> ids {encode(text)} pieces {pieces(text)}")Code explained
- In simple words: we measure the strings that are not ordinary words and look at exactly where they break.
- What happens:
SAMPLESholds realistic Brightlane strings, including a realTriageobject fromsupportdesk/schemas.pyserialized two ways: pretty (indented) and compact (no spaces). The table prints token counts per tokenizer and characters per token foro200k_base(higher is cheaper). Thenpieces()shows the split points for ids, numbers, code, and the classic failure cases, andencode()shows that one word gets different ids depending on spacing and case. - Comes out:
sample chars p50k_base cl100k_base o200k_base claude-legacy tinylm chars/token (o200k)
prose 52 10 10 10 10 11 5.20
invoice ids 49 27 23 23 28 29 2.13
amounts 46 19 21 21 19 27 2.19
long number 44 14 14 14 13 26 3.14
python code 62 21 15 15 19 42 4.13
json (pretty) 169 53 45 46 47 85 3.67
json (compact) 148 33 30 31 33 74 4.77
url + email 78 22 18 18 22 29 4.33
Where the pieces fall (o200k_base):
'INV-2026-004512' -> ['INV', '-', '202', '6', '-', '004', '512']
'3141592653' -> ['314', '159', '265', '3']
'1,234.56' -> ['1', ',', '234', '.', '56']
' notify(oncall, ticket.id)' -> [' ', ' notify', '(on', 'call', ',', ' ticket', '.id', ')']
Why letter counting, spelling, and arithmetic are hard (o200k_base):
'Brightlane' -> ['Bright', 'lane']
'strawberry' -> ['st', 'raw', 'berry']
' refund' -> [' refund']
' unsubscribed' -> [' unsub', 'scribed']
'288 / 2 = 144' -> ['288', ' /', ' ', '2', ' =', ' ', '144']
'14 * 24 = 336' -> ['14', ' *', ' ', '24', ' =', ' ', '336']
'12345 + 67890' -> ['123', '45', ' +', ' ', '678', '90']
'r' in 'strawberry': 3 (the model never sees letters, only the pieces above)
Same word, different tokens (o200k_base ids):
'refund' -> ids [148482] pieces ['refund']
' refund' -> ids [18376] pieces [' refund']
' Refund' -> ids [100598] pieces [' Refund']
' REFUND' -> ids [76596, 21592] pieces [' REF', 'UND']Prose runs at 5.2 characters per token; invoice ids and amounts at about 2.1. Ids are expensive because they are unpredictable: INV-2026-004512 becomes INV, -, 202, 6, -, 004, 512. The cl100k_base and o200k_base tokenizers split digits into groups of at most three, from the left, so 3141592653 becomes 314|159|265|3 and the groups do not line up with thousands. Pretty-printed JSON costs 46 tokens where compact JSON costs 31 for the same data: indentation and spaces are tokens too, a 33% saving on every structured reply if you ask for compact output. Python code is fairly efficient in the newer tokenizers (runs of spaces merge into one token) and poor in p50k_base.
Tokenization artifacts: counting, spelling, and arithmetic
The last part of that output explains several failures from Module 1's list of things models are unreliable at.
- Counting letters. The model receives
st|raw|berry, three ids. It never sees the letters r, r, r. To count them it must have learned, from training text, which letters each token contains. It often has, but not reliably, which is why "how many r's in strawberry" became a famous failure. - Spelling and character edits.
refund,Refund, andrefundare three unrelated ids (18376, 100598, 148482), andREFUNDis two tokens. Tasks like "reverse this invoice id", "is this id's check digit correct", or "fix the typo in this SKU" operate on characters the model cannot see directly. - Arithmetic.
12345 + 67890arrives as123|45| +| |678|90. Column addition needs digits aligned by place value, but the pieces are aligned by the tokenizer's left-to-right three-digit grouping. The model has to reconstruct place values from chunk boundaries that mean nothing numerically.
The engineering response is the same in each case: do not ask the model to do character-level or arithmetic work you can do in code. For Brightlane that means computing refund amounts, seat totals, and prorations in Python, extracting invoice ids with a regular expression, and validating them in code (Module 6), then giving the model the results.
| Situation | Use this | Why |
|---|---|---|
| Count, reverse, or validate characters in an id | Python string operations or a regex | The model sees tokens, not characters |
| Compute a refund or a seat total | Python arithmetic, then pass the result to the model | Digit chunks do not align with place value |
| Ask the model to return numbers or ids | Copy them from input, then check them in code | Copying is reliable, transformation is not |
| Structured replies | Compact JSON | About a third fewer tokens than indented JSON here |
Counting before you send
Before any call, you want to know: will this request fit, and what will it cost? count_messages() gives the estimate. Here it is on a full triage request.
# examples/m02_count_request.py
"""Module 2: count a full triage request before sending it, message by message and across tokenizers."""
from m02_prompts import triage_messages
from supportdesk.data import load_tickets
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import ENCODINGS, count_messages, count_tokens
ticket = load_tickets("test")[0]
messages = triage_messages(ticket)
print(f"Triage request for {ticket.id} ({len(messages)} messages), o200k_base:")
for m in messages:
first_line = m["content"].splitlines()[0][:48]
print(f" {m['role']:9} {count_tokens(m['content']):4} tokens {first_line!r}")
print(f" count_messages total: {count_messages(messages)} (content + 4 per message + 3 priming)")
print("\nSame request under each tokenizer:")
for name in ENCODINGS:
print(f" {name:13} {count_messages(messages, encoding=name):5}")
reply = '{"category":"cancellation","priority":"normal","language":"en","summary":"Customer wants to cancel their plan.","needs_human":true}'
fake = ScriptedLLM(replies=[reply]) # plumbing only: not a model, and its usage is itself an estimate
result = fake(messages, max_tokens=200)
print(f"\nScriptedLLM usage (estimated with count_messages, not reported by a provider): {result.usage}")Code explained
- In simple words: we weigh the request message by message before it leaves the building.
- What happens: we build the triage request for the first test ticket, count each message's content, and call
count_messages()for the total. Then we count the same request with every tokenizer. Finally we pass it throughScriptedLLM, which is not a model: it returns our scripted JSON and fillsusageusing the samecount_messages()estimate, so its numbers prove only that the plumbing carries usage through. - Comes out:
Triage request for T-1003 (6 messages), o200k_base:
system 512 tokens "You are the triage assistant for Brightlane's su"
user 15 tokens 'Subject: Invoice address'
assistant 30 tokens '{"category":"billing","priority":"low","language'
user 20 tokens 'Subject: SSO broken'
assistant 32 tokens '{"category":"account_access","priority":"urgent"'
user 28 tokens 'Customer tier: team'
count_messages total: 664 (content + 4 per message + 3 priming)
Same request under each tokenizer:
p50k_base 723
cl100k_base 652
o200k_base 664
claude-legacy 713
ScriptedLLM usage (estimated with count_messages, not reported by a provider): Usage(input_tokens=664, output_tokens=29, cached_tokens=0, reasoning_tokens=0)The system prompt is 77% of the request (512 of 664 tokens), and it is identical for every ticket; the ticket itself is 28 tokens. Keep that ratio in mind: it drives the caching result in Part D. The four tokenizers disagree by up to 11% (652 to 723) on the same request, which is the size of error to expect when you count with the wrong tokenizer. The 29 output tokens are for the compact JSON label.
A provider's reported usage is the ground truth, and it is usually a little higher than a content-only estimate because the chat template adds tokens you never wrote (role headers, a default system preamble for some models, tool definitions rendered as text). Calibrate once per model with a few real calls:
# examples/m02_calibrate.py
"""Module 2: compare our pre-send estimate with the usage a provider reports.
PYTHONPATH=. python examples/m02_calibrate.py --real (needs a provider key)
"""
import argparse
from m02_prompts import triage_messages
from supportdesk.data import load_tickets
from supportdesk.llm import chat
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages
parser = argparse.ArgumentParser()
parser.add_argument("--real", action="store_true")
args = parser.parse_args()
llm = chat if args.real else ScriptedLLM(responder=lambda messages, kwargs: '{"category":"billing"}')
ratios = []
for ticket in load_tickets("test")[:8]:
messages = triage_messages(ticket)
estimate = count_messages(messages)
result = llm(messages, max_tokens=512)
ratios.append(result.usage.input_tokens / estimate)
print(f"{ticket.id} {ticket.language}: estimated {estimate:4} reported {result.usage.input_tokens:4} "
f"output {result.usage.output_tokens:4} (reasoning {result.usage.reasoning_tokens})")
print(f"reported / estimated: min {min(ratios):.3f} max {max(ratios):.3f} "
f"({'real provider' if args.real else 'stand-in: it reports our own estimate, so 1.000 by construction'})")Code explained
- In simple words: send eight real requests, compare what we predicted with what the provider billed, and learn our correction factor.
- What happens: for eight test tickets, estimate with
count_messages(), call the model, and divideresult.usage.input_tokens(read byllm.pyfrom the provider'susage.prompt_tokens) by the estimate. Without--realit usesScriptedLLM, whose "reported" usage is our own estimate. - Comes out (stand-in, captured):
T-1003 en: estimated 664 reported 664 output 5 (reasoning 0)
T-1006 en: estimated 668 reported 668 output 5 (reasoning 0)
T-1009 en: estimated 662 reported 662 output 5 (reasoning 0)
T-1012 en: estimated 660 reported 660 output 5 (reasoning 0)
T-1015 en: estimated 663 reported 663 output 5 (reasoning 0)
T-1018 en: estimated 662 reported 662 output 5 (reasoning 0)
T-1021 en: estimated 669 reported 669 output 5 (reasoning 0)
T-1024 en: estimated 662 reported 662 output 5 (reasoning 0)
reported / estimated: min 1.000 max 1.000 (stand-in: it reports our own estimate, so 1.000 by construction)With --real, you get the provider's numbers. Illustrative sample run (not captured in this build; produced for teaching). Your output will differ:
T-1003 en: estimated 664 reported 741 output 212 (reasoning 171)
T-1006 en: estimated 668 reported 745 output 198 (reasoning 158)
...
reported / estimated: min 1.112 max 1.118The shape is what to look for: a stable ratio a bit above 1 (template overhead), and output far larger than the 30-token label because a reasoning model spends hidden reasoning tokens first. Use your measured ratio as the safety margin in your budget.