Part A: The Project and the Toolkit
By the end of this module, you'll have:
- Run a real language model (TinyLM, 1.07 million parameters) on your own CPU, read its next-token probabilities, and generated a support reply one token at a time.
- Read a real training log, pointed at the step where the model started to overfit, and measured how much of its output is memorized text.
- Read a model spec (parameters, layers, context window, vocabulary) for TinyLM and for published open-weight models, and turned parameter counts into gigabytes of memory.
- Measured four limits yourself: tokenization hiding letters and digits, sensitivity to phrasing, calibration, and hallucination outside the training data.
- Built a selection framework (capability floor, latency, cost, privacy, control) and applied it to the Brightlane support assistant, plus a one-page measured fact sheet for a model.
- Set up the
supportdeskrepository with the dataset and three helper files (data.py,llm.py,stand_in.py) that every later module imports.
Prerequisites: Working Python (functions, classes, dataclasses, running scripts from a terminal). No machine learning background. A free API key (Groq or Gemini) or a local Ollama install is optional for this module: every example runs without one.
Where we are: This is the first module. Before we prompt, retrieve, fine-tune, or evaluate anything, we need an honest picture of the object itself: what a large language model computes, how it was made, what it is reliably good and bad at, and how to choose one.
How this module is organized
| Part | What it covers |
|---|---|
| Part A: The Project and the Toolkit | The Brightlane support desk, the ticket dataset, and the helpers data.py, llm.py, and stand_in.py; a first triage call |
| Part B: Next-Token Prediction | TinyLM, next-token probabilities, generating a reply token by token, why one objective yields many skills |
| Part C: What Pretraining Learned | The training log, loss, overfitting, memorization, and hallucination as a property of the objective |
| Part D: Reading a Model Spec | Parameters, layers, context window, vocabulary; checking published configs; memory arithmetic |
| Part E: The Lifecycle | Pretraining, mid-training, supervised fine-tuning, preference optimization, reasoning training; base vs tuned models; open weights vs open source |
| Part F: Capabilities and Limits, Measured | Tokens vs letters and digits, phrasing sensitivity, calibration, knowledge cutoffs, self-knowledge |
| Part G: The Model Landscape | Model tiers, proprietary vs open weights, reasoning vs standard, small on-device models, benchmarks, a selection framework |
Setting up
The course has one code repository, supportdesk. Every module adds example scripts to examples/ and tests to tests/, and imports the shared helpers under supportdesk/. All commands run from the repository root with PYTHONPATH=. so that import supportdesk works without installing anything.
cd supportdesk
python3.11 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
export PYTHONPATH=.
python -m pytest -q tests/test_m01_foundations.pyCode explained
- In simple words: make an isolated Python environment, install the pinned libraries, and run this module's tests to prove the setup works.
- What happens:
venvcreates a private Python 3.11 environment so course libraries do not collide with anything else on your machine.requirements.txtpins exact versions (openai 3.16.2, pydantic 2.13.5, tiktoken 0.14.0, tokenizers 0.23.2, torch 2.14.0, numpy 2.4.6, pytest 9.1.1).PYTHONPATH=.lets Python find thesupportdeskpackage in the current folder. The tests load the dataset, load TinyLM, and check a few behaviors you will see in this module. - Comes out: pytest prints one dot per passing test. Timing depends on your machine.
........ [100%]
8 passed in 13.28sThe examples in this module are complete scripts in examples/. Run them in the order they appear; later ones reuse ideas (not variables) from earlier ones. PyTorch examples call torch.set_num_threads(1): a 1-million-parameter model does not benefit from more threads, and one thread keeps timing steady on a busy laptop.
Part A: The Project and the Toolkit
The Brightlane support desk
Brightlane is a fictional project-management SaaS with four plans: Free, Team (12 USD per user per month), Business (24 USD), and Enterprise (custom pricing). Its support desk receives tickets in several languages. Maya, the support lead, wants an assistant that triages incoming tickets, answers from the help center, and drafts replies that her team reviews before sending. Over the course you will build that assistant, then evaluate it, secure it, adapt it, and operate it.
Two data sources ship with the repository:
data/tickets.jsonl: 72 hand-written tickets. Each has a gold label, meaning the correct answer a human decided in advance: a category, a priority, which help-center article answers it, and whether it is answerable from the help center at all. The tickets are split into a dev set (48 tickets you may look at while building) and a test set (24 tickets you keep for the final check, so you do not tune to them by accident).data/kb/*.md: 12 help-center articles. The file name without.mdis the article id, for examplebilling-refunds.
Here is one line of tickets.jsonl, pretty-printed:
{
"id": "T-1001",
"subject": "Charged twice this month",
"body": "Hi, my card was charged 288 USD twice on 3 September for the Team plan (invoice INV-2026-004512). Please refund the duplicate.",
"customer_tier": "team",
"language": "en",
"split": "dev",
"gold": {"category": "billing", "priority": "high", "kb_article": "billing-refunds", "answerable": true}
}Code explained
- In simple words: one ticket is one JSON object on one line, like a row in a spreadsheet with a nested "answer key".
- What happens:
subjectandbodyare what the customer wrote.customer_tieris their plan.languageis an ISO code (en, es, de, ja, hi).splitsays whether the ticket belongs to dev or test.goldholds the labels:categoryis one of six values,priorityone of four,kb_articlenames the article that answers it (ornullwhen no article does), andanswerablesays whether the help center covers it. - Comes out: nothing runs here; this is the data format. In the file each ticket sits on a single line (the JSON Lines format), which makes it easy to append and stream.
Now load everything and look at the shape of the data.
"""Module 1: a first look at the Brightlane dataset the whole course uses."""
from collections import Counter
from supportdesk.data import CATEGORIES, PRIORITIES, get_article, load_articles, load_tickets
tickets = load_tickets()
dev, test = load_tickets("dev"), load_tickets("test")
articles = load_articles()
print(f"tickets: {len(tickets)} (dev {len(dev)}, test {len(test)})")
print("languages:", dict(Counter(t.language for t in tickets).most_common()))
print("categories:", {c: sum(t.gold["category"] == c for t in tickets) for c in CATEGORIES})
print("priorities:", {p: sum(t.gold["priority"] == p for t in tickets) for p in PRIORITIES})
print("answerable from the help center:", sum(t.gold["answerable"] for t in tickets))
print(f"articles: {len(articles)}:", ", ".join(a.id for a in articles))
first = tickets[0]
print("\n--- one ticket, as an agent reads it ---")
print(first.id, first.customer_tier, first.language)
print(first.text)
print("gold:", first.gold)
article = get_article(first.gold["kb_article"])
print("\n--- the article that answers it ---")
print(article.title, article.tags)
print(article.body[:160] + "...")Code explained
- In simple words: count what is in the dataset and print one ticket next to the article that answers it, the way a new support agent would study the queue on day one.
- What happens:
load_tickets()returns all 72 tickets asTicketobjects;load_tickets("dev")andload_tickets("test")filter by split.Countertallies languages. The category and priority loops count gold labels using the canonicalCATEGORIESandPRIORITIEStuples, so a typo in a label would show up as a missing count.ticket.textjoins subject and body.get_articlefetches the gold article by id. - Comes out: the dataset is small and uneven: 64 of 72 tickets are English, only 4 are urgent, and 10 cannot be answered from the help center. Keep those numbers in mind. With 72 tickets, one ticket is 1.4 percentage points of accuracy, so small differences you measure later will often be noise.
tickets: 72 (dev 48, test 24)
languages: {'en': 64, 'de': 3, 'es': 2, 'ja': 2, 'hi': 1}
categories: {'billing': 16, 'cancellation': 6, 'account_access': 17, 'bug': 9, 'how_to': 18, 'feature_request': 6}
priorities: {'low': 24, 'normal': 31, 'high': 13, 'urgent': 4}
answerable from the help center: 62
articles: 12: account-login, account-sso, billing-invoices, billing-plans, billing-refunds, boards-automations, data-privacy, exports-data, feature-requests, integrations-slack, mobile-app, status-incidents
--- one ticket, as an agent reads it ---
T-1001 team en
Subject: Charged twice this month
Hi, my card was charged 288 USD twice on 3 September for the Team plan (invoice INV-2026-004512). Please refund the duplicate.
gold: {'category': 'billing', 'priority': 'high', 'kb_article': 'billing-refunds', 'answerable': True}
--- the article that answers it ---
Refunds and cancellations ('billing', 'cancellation')
You can cancel at any time from Settings > Billing > Cancel plan. Monthly plans stay active until the end of the current billing period and are not refunded pro...The shared helpers
Three files are introduced in this module and imported by every later one. You do not need to memorize them. Skim each, read the explanation, and come back when a later module uses a function.
supportdesk/data.py
"""Load the Brightlane support dataset: tickets with gold labels and help-center articles."""
from __future__ import annotations
import json
from dataclasses import dataclass
from pathlib import Path
DATA_DIR = Path(__file__).resolve().parents[1] / "data"
CATEGORIES = ("billing", "cancellation", "account_access", "bug", "how_to", "feature_request")
PRIORITIES = ("low", "normal", "high", "urgent")
@dataclass(frozen=True)
class Ticket:
id: str
subject: str
body: str
customer_tier: str
language: str
split: str
gold: dict
@property
def text(self) -> str:
"""Subject and body as one string, the way an agent reads a ticket."""
return f"Subject: {self.subject}\n\n{self.body}"
@dataclass(frozen=True)
class Article:
id: str
title: str
tags: tuple[str, ...]
body: str
def load_tickets(split: str | None = None) -> list[Ticket]:
"""All tickets, or only the 'dev' or 'test' split. Order is stable."""
tickets = []
with (DATA_DIR / "tickets.jsonl").open(encoding="utf-8") as f:
for line in f:
t = Ticket(**json.loads(line))
if split is None or t.split == split:
tickets.append(t)
return tickets
def load_articles() -> list[Article]:
"""Help-center articles from data/kb, sorted by id."""
articles = []
for path in sorted((DATA_DIR / "kb").glob("*.md")):
text = path.read_text(encoding="utf-8")
_, header, body = text.split("---\n", 2)
meta = dict(line.split(": ", 1) for line in header.strip().splitlines())
tags = tuple(t.strip() for t in meta.get("tags", "").strip("[]").split(",") if t.strip())
articles.append(Article(path.stem, meta["title"], tags, body.strip()))
return articles
def get_article(article_id: str) -> Article:
for article in load_articles():
if article.id == article_id:
return article
raise KeyError(f"No article {article_id!r}")Code explained
- In simple words:
data.pyturns the files indata/into Python objects so no example ever parses JSON or Markdown by hand. - What happens:
DATA_DIR,CATEGORIES,PRIORITIES: the data folder (found relative to this file, so it works from any working directory) and the fixed label sets. Later modules build schemas and metrics from these tuples, so there is one source of truth.Ticket: a frozen dataclass (its fields cannot be changed after creation, which prevents a script from accidentally editing gold labels). Thetextproperty formats subject and body the way an agent reads them.Article: the same idea for a help-center article: id, title, tags, and body.load_tickets(split): reads the JSON Lines file line by line, builds aTicketfrom each, and keeps only the requested split. Order is stable, so results are reproducible.load_articles(): reads everykb/*.mdfile. Each file starts with a small header between---lines (called front matter); the function splits it off, parsestitleandtags, and keeps the rest as the body.get_article(article_id): finds one article by id and raises a clearKeyErrorif it does not exist.- Comes out: nothing when imported; the functions return the objects you saw in the previous example.
supportdesk/llm.py
"""One helper for three LLM providers: Groq, Gemini, and Ollama.
All three expose an OpenAI-compatible Chat Completions endpoint, so a single
client library (`openai`) talks to each of them. Choose a provider with the
LLM_PROVIDER environment variable (groq, gemini, ollama) and override the
model with LLM_MODEL. Keys come from GROQ_API_KEY and GEMINI_API_KEY; Ollama
runs locally and needs no key.
Every call returns a ChatResult with the text, any tool calls, token usage,
and latency, because the course measures cost and speed on every example.
"""
from __future__ import annotations
import json
import os
import time
from collections.abc import Iterator
from dataclasses import dataclass, field
from typing import Any
from openai import OpenAI
PROVIDERS: dict[str, dict[str, str]] = {
"groq": {
"base_url": "https://api.groq.com/openai/v1",
"key_env": "GROQ_API_KEY",
"default_model": "openai/gpt-oss-120b",
},
"gemini": {
"base_url": "https://generativelanguage.googleapis.com/v1beta/openai/",
"key_env": "GEMINI_API_KEY",
"default_model": "gemini-3.5-flash",
},
"ollama": {
"base_url": "http://localhost:11434/v1",
"key_env": "",
"default_model": "qwen3:8b",
},
}
@dataclass
class ToolCall:
id: str
name: str
arguments: dict[str, Any]
raw_arguments: str = ""
@dataclass
class Usage:
input_tokens: int = 0
output_tokens: int = 0
cached_tokens: int = 0
reasoning_tokens: int = 0
@dataclass
class ChatResult:
text: str
tool_calls: list[ToolCall] = field(default_factory=list)
usage: Usage = field(default_factory=Usage)
latency_ms: float = 0.0
finish_reason: str = ""
model: str = ""
provider: str = ""
def as_message(self) -> dict[str, Any]:
"""The assistant message to append to the conversation history."""
message: dict[str, Any] = {"role": "assistant", "content": self.text or ""}
if self.tool_calls:
message["tool_calls"] = [
{
"id": call.id,
"type": "function",
"function": {"name": call.name, "arguments": call.raw_arguments or json.dumps(call.arguments)},
}
for call in self.tool_calls
]
return message
def resolve(provider: str | None = None, model: str | None = None) -> tuple[str, str]:
"""Pick the provider and model from arguments first, then the environment, then defaults."""
name = (provider or os.environ.get("LLM_PROVIDER", "groq")).lower()
if name not in PROVIDERS:
raise ValueError(f"Unknown LLM_PROVIDER {name!r}. Choose one of {sorted(PROVIDERS)}.")
return name, model or os.environ.get("LLM_MODEL") or PROVIDERS[name]["default_model"]
def make_client(provider: str) -> OpenAI:
"""Build an OpenAI-compatible client for one provider, reading its key from the environment."""
settings = PROVIDERS[provider]
if settings["key_env"]:
api_key = os.environ.get(settings["key_env"], "")
if not api_key:
raise RuntimeError(f"Set {settings['key_env']} in your environment to use {provider}.")
base_url = settings["base_url"]
else:
api_key = "ollama"
base_url = os.environ.get("OLLAMA_BASE_URL", settings["base_url"])
return OpenAI(api_key=api_key, base_url=base_url, max_retries=2, timeout=120.0)
def _usage(raw: Any) -> Usage:
if raw is None:
return Usage()
prompt_details = getattr(raw, "prompt_tokens_details", None)
completion_details = getattr(raw, "completion_tokens_details", None)
return Usage(
input_tokens=raw.prompt_tokens or 0,
output_tokens=raw.completion_tokens or 0,
cached_tokens=(getattr(prompt_details, "cached_tokens", 0) or 0) if prompt_details else 0,
reasoning_tokens=(getattr(completion_details, "reasoning_tokens", 0) or 0) if completion_details else 0,
)
def _parse_arguments(raw: str) -> dict[str, Any]:
try:
value = json.loads(raw or "{}")
except json.JSONDecodeError:
return {"_unparseable": raw}
return value if isinstance(value, dict) else {"_value": value}
def chat(
messages: list[dict[str, Any]],
*,
provider: str | None = None,
model: str | None = None,
temperature: float | None = 0.0,
top_p: float | None = None,
max_tokens: int | None = None,
stop: list[str] | None = None,
seed: int | None = None,
tools: list[dict[str, Any]] | None = None,
tool_choice: str | dict[str, Any] | None = None,
response_format: dict[str, Any] | None = None,
reasoning_effort: str | None = None,
extra: dict[str, Any] | None = None,
) -> ChatResult:
"""Send one Chat Completions request and return text, tool calls, usage, and latency.
Parameters left as None are not sent, so each provider uses its own default.
`extra` passes provider-specific fields through unchanged.
"""
name, model_id = resolve(provider, model)
kwargs: dict[str, Any] = {"model": model_id, "messages": messages}
optional = {
"temperature": temperature, "top_p": top_p, "max_tokens": max_tokens, "stop": stop,
"seed": seed, "tools": tools, "tool_choice": tool_choice,
"response_format": response_format, "reasoning_effort": reasoning_effort,
}
kwargs.update({k: v for k, v in optional.items() if v is not None})
if extra:
kwargs["extra_body"] = extra
started = time.perf_counter()
response = make_client(name).chat.completions.create(**kwargs)
latency_ms = (time.perf_counter() - started) * 1000
choice = response.choices[0]
calls = [
ToolCall(c.id, c.function.name, _parse_arguments(c.function.arguments), c.function.arguments or "")
for c in (choice.message.tool_calls or [])
if c.type == "function"
]
return ChatResult(
text=choice.message.content or "",
tool_calls=calls,
usage=_usage(response.usage),
latency_ms=round(latency_ms, 1),
finish_reason=choice.finish_reason or "",
model=model_id,
provider=name,
)
@dataclass
class StreamStats:
ttft_ms: float = 0.0
total_ms: float = 0.0
chunks: int = 0
def stream_chat(
messages: list[dict[str, Any]],
stats: StreamStats | None = None,
*,
provider: str | None = None,
model: str | None = None,
temperature: float | None = 0.0,
max_tokens: int | None = None,
) -> Iterator[str]:
"""Yield text pieces as they arrive. Fills `stats` with time to first token and total time."""
name, model_id = resolve(provider, model)
kwargs: dict[str, Any] = {"model": model_id, "messages": messages, "stream": True}
if temperature is not None:
kwargs["temperature"] = temperature
if max_tokens is not None:
kwargs["max_tokens"] = max_tokens
stats = stats if stats is not None else StreamStats()
started = time.perf_counter()
for chunk in make_client(name).chat.completions.create(**kwargs):
if not chunk.choices:
continue
piece = chunk.choices[0].delta.content or ""
if piece:
if stats.chunks == 0:
stats.ttft_ms = round((time.perf_counter() - started) * 1000, 1)
stats.chunks += 1
yield piece
stats.total_ms = round((time.perf_counter() - started) * 1000, 1)Code explained
- In simple words:
llm.pyis one telephone that can dial three different companies. You always speak the same language (a list of messages), and it always hands back the same kind of answer (aChatResult). - What happens:
PROVIDERS: the three supported providers. Groq and Gemini are hosted services with free tiers; Ollama runs models on your own machine. All three speak the OpenAI-compatible Chat Completions format (a de facto standard request shape: a model name plus a list of{"role", "content"}messages), so one client library covers them. The default models areopenai/gpt-oss-120bon Groq,gemini-3.5-flashon Gemini, andqwen3:8bon Ollama.ToolCall,Usage,ChatResult: plain dataclasses for what comes back.Usagerecords tokens (the units models read and write, roughly word pieces; Module 2 covers them in depth) split into input, output, cached, and reasoning tokens, because cost is billed per token.ChatResult.as_message()converts a reply back into a message you can append to the conversation.resolve(provider, model): picks the provider and model from function arguments, then theLLM_PROVIDERandLLM_MODELenvironment variables, then the defaults. Unknown providers fail loudly.make_client(provider): builds anopenai.OpenAIclient pointed at the provider's URL, with the key read from the environment (never from code). It raises a clear error if the key is missing.max_retries=2retries transient network errors._usageand_parse_arguments: small private helpers._usagecopes with providers that omit usage details._parse_argumentsnever raises on malformed tool-call JSON; it wraps the raw text so the caller can decide what to do (Module 6 builds on this).chat(messages, ...): sends one request. Parameters left asNoneare not sent, so each provider keeps its own default. It times the call, extracts text, tool calls, usage, and finish reason, and returns aChatResult.temperaturedefaults to 0.0 (the least random setting; Module 3 explains why that still is not fully deterministic).StreamStatsandstream_chat(...): the streaming version, which yields text as it arrives and records time to first token. Module 3 uses it.- Comes out: nothing when imported. Calling
chatneeds a key or a running Ollama server.
supportdesk/stand_in.py
"""A scripted stand-in for `llm.chat`, for running examples and tests without an API key.
It has the same call signature as `supportdesk.llm.chat` and returns the same
ChatResult type. It is NOT a language model: it replays scripted replies or
applies a rule you give it. Use it to test plumbing (parsing, loops, retries,
budgets), never to measure model quality.
"""
from __future__ import annotations
import itertools
from collections.abc import Callable
from typing import Any
from supportdesk.llm import ChatResult, Usage
from supportdesk.tokens import count_messages, count_tokens
Responder = Callable[[list[dict[str, Any]], dict[str, Any]], ChatResult | str]
class ScriptedLLM:
"""Replays a list of replies in order, or calls `responder(messages, kwargs)`."""
_ids = itertools.count(1)
def __init__(self, replies: list[ChatResult | str] | None = None, responder: Responder | None = None,
model: str = "scripted-stand-in") -> None:
if (replies is None) == (responder is None):
raise ValueError("Pass exactly one of `replies` or `responder`.")
self.replies = list(replies or [])
self.responder = responder
self.model = model
self.calls: list[dict[str, Any]] = []
def __call__(self, messages: list[dict[str, Any]], **kwargs: Any) -> ChatResult:
self.calls.append({"messages": [dict(m) for m in messages], **kwargs})
if self.responder is not None:
reply = self.responder(messages, kwargs)
elif self.replies:
reply = self.replies.pop(0)
else:
raise RuntimeError("ScriptedLLM ran out of scripted replies.")
result = ChatResult(text=reply) if isinstance(reply, str) else reply
if result.usage.input_tokens == 0:
result.usage = Usage(input_tokens=count_messages(messages), output_tokens=count_tokens(result.text or ""))
result.model = result.model or self.model
result.provider = result.provider or "stand-in"
result.finish_reason = result.finish_reason or ("tool_calls" if result.tool_calls else "stop")
return resultCode explained
- In simple words:
ScriptedLLMis a crash-test dummy. It has the same shape as a model call, so you can test everything around the model without paying for, or waiting on, a real one. It has no intelligence. - What happens:
ScriptedLLM(replies=...)replays a list of replies in order.ScriptedLLM(responder=...)instead calls a function you write, which receives the messages and keyword arguments and returns text or aChatResult. Passing both, or neither, is an error.__call__makes the object callable with the same signature asllm.chat, so code written aschat(messages, temperature=0)works with either. It records every call inself.calls(handy for tests), fills in estimated token usage with the offline tokenizer fromtokens.pywhen you did not script it, and setsprovider="stand-in"so logs never confuse it with a real model.- When the scripted replies run out, it raises instead of inventing text.
- Comes out: a
ChatResult, exactly likellm.chat.
Your first call: a triage question
Now use the helpers together. The script below asks a model to triage one real ticket. If a key is configured, it calls the real provider through llm.chat. If not, it falls back to ScriptedLLM so the rest of the code still runs.
"""Module 1: the first call to a hosted model, and the same call through the stand-in.
With a key set (for example GROQ_API_KEY, or LLM_PROVIDER=ollama with Ollama running),
this sends a real request through supportdesk.llm.chat. Without one, it falls back to
ScriptedLLM, which is NOT a model: it replays a scripted reply so the plumbing runs.
"""
import os
from supportdesk import llm
from supportdesk.data import CATEGORIES, load_tickets
from supportdesk.stand_in import ScriptedLLM
ticket = next(t for t in load_tickets() if t.id == "T-1008")
messages = [
{"role": "system", "content": "You triage support tickets for Brightlane, a project-management SaaS. "
f"Reply with one category from: {', '.join(CATEGORIES)}. "
"Then one short sentence explaining why."},
{"role": "user", "content": ticket.text},
]
provider, model = llm.resolve()
key_env = llm.PROVIDERS[provider]["key_env"]
use_real = provider == "ollama" or bool(os.environ.get(key_env))
if use_real:
chat = llm.chat
print(f"calling {provider} / {model}")
else:
chat = ScriptedLLM(replies=["account_access\nThe customer is locked out after failed password attempts."])
print(f"no {key_env} set: using ScriptedLLM (scripted reply, not model output)")
result = chat(messages, temperature=0.0, max_tokens=200)
print("text: ", result.text.replace("\n", " | "))
print("provider/model:", result.provider, "/", result.model)
print("finish_reason: ", result.finish_reason)
print("usage: ", result.usage)
print("latency_ms: ", result.latency_ms)
print("gold category: ", ticket.gold["category"])Code explained
- In simple words: build a two-message conversation (instructions plus the ticket), send it to whichever "model" is available, and print everything the result object carries.
- What happens: the system message sets the job and the allowed categories (taken from
CATEGORIES, not typed by hand). The user message is the ticket text for T-1008, a customer who is locked out.llm.resolve()reports which provider and model would be used. The script checks the provider's key variable; Ollama needs no key. Both paths callchat(messages, temperature=0.0, max_tokens=200)with identical arguments, which is the point of the shared signature. - Comes out: this build has no API key, so the output below is real plumbing output from
ScriptedLLM. The text is the scripted reply, not a model's judgment. The usage numbers are real counts from the offline tokenizer (90 input tokens, 13 output tokens) and latency is 0 because nothing went over the network.
no GROQ_API_KEY set: using ScriptedLLM (scripted reply, not model output)
text: account_access | The customer is locked out after failed password attempts.
provider/model: stand-in / scripted-stand-in
finish_reason: stop
usage: Usage(input_tokens=90, output_tokens=13, cached_tokens=0, reasoning_tokens=0)
latency_ms: 0.0
gold category: account_accessWith a Groq key set (export GROQ_API_KEY=...), the same script calls openai/gpt-oss-120b. The following is an illustrative sample run (not captured in this build; produced for teaching). Your output will differ, including the wording, the token counts, and the latency:
calling groq / openai/gpt-oss-120b
text: account_access | The customer is locked out after too many failed password attempts and needs access restored.
provider/model: groq / openai/gpt-oss-120b
finish_reason: stop
usage: Usage(input_tokens=131, output_tokens=74, cached_tokens=0, reasoning_tokens=48)
latency_ms: 640.2
gold category: account_accessCode explained
- In simple words: the shape a real reply takes, so you know what to look for when you run it yourself.
- What happens: a real provider counts input tokens with its own tokenizer and chat formatting, so the input count differs from the stand-in's estimate.
gpt-oss-120bis a reasoning model: it writes hidden "thinking" tokens before the answer, and they show up inreasoning_tokensand are billed as output tokens. - Comes out: labelled illustrative. Run the script with your key to get real numbers; you will use them in Module 2 to compute cost per ticket.
Two failures you will almost certainly hit, and how to read them:
| What you see | What it means | Fix |
|---|---|---|
RuntimeError: Set GROQ_API_KEY in your environment to use groq. | make_client found no key for the selected provider | export GROQ_API_KEY=..., or export LLM_PROVIDER=gemini with GEMINI_API_KEY, or export LLM_PROVIDER=ollama |
openai.APIConnectionError: Connection error. with LLM_PROVIDER=ollama | Nothing is listening on localhost:11434 | Start Ollama (ollama serve) and pull the model (ollama pull qwen3:8b), or set OLLAMA_BASE_URL |
Both are real errors captured in this build. Read the last line of the traceback first: llm.py is written so that line names the missing piece.
Across the course you will choose between three ways of running an example:
| Situation | Use this | Why |
|---|---|---|
| You want to know how well a model does a task | llm.chat with a real provider | Only a real model produces evidence about quality |
| You are testing loops, parsers, retries, budgets, or routing | ScriptedLLM | Deterministic, free, instant; failures you see are in your code, not the model |
| You want to see a mechanism (probabilities, sampling, training) with real numbers | TinyLM | You can open the model up; no hosted API exposes its internals this way |