Part A: Getting reliable structure
Module 6: Structured Outputs and Tool Use
By the end of this module, you'll have:
- A tolerant JSON parser, measured on 24 hostile model outputs, that raises the parse rate from 21% to 88% and logs every repair it makes.
- A repair loop with an attempt ceiling and a salvage step, so a malformed answer becomes a valid object, a safe partial result, or a clean hand-off to a human, never a crash.
- A strict
response_formatbuilt from the course's pydantic schemas, plus a clear picture of which of Groq, Gemini, and Ollama enforce it. - A tool-calling loop for the Brightlane assistant with three tools (
search_kb,get_invoice,issue_refund), parallel calls ordered bytool_call_id,tool_choicecontrol, and errors returned to the model instead of raised. - A guard layer that treats model output as untrusted input: argument validation, permission scoping, timeouts, a process sandbox, and human approval before any refund. A refund larger than its invoice is rejected before a person is even asked.
Prerequisites: Modules 1 to 5. You need llm.chat and ScriptedLLM (Module 1), token counting with tokens.py and cost math with pricing.py (Module 2), and the prompt-as-code habits of Modules 4 and 5. kb_search.py is formally introduced in Module 7; here we only call KBSearch().search() as a tool.
Where we are: In Module 5 you chained calls and ran self-correction loops, and every step still passed plain text to the next. That works until a program, not a person, has to read the answer. This module makes the assistant's output something code can trust: typed triage objects, and tool calls that are checked before anything happens in Brightlane's billing system.
How this module is organized
| Part | What it covers |
|---|---|
| Setup | The working copy, and how the examples run |
| Part A: Getting reliable structure | Why free text breaks code, schema design, native structured output vs prompted JSON, pydantic validation, a hostile-output parser, repair loops, salvage, schema size and depth |
| Part B: Tool and function calling | What the model really sees, descriptions and parameters, the execution loop, errors as data, parallel calls, tool_choice, too many tools |
| Part C: Model output as untrusted input | Argument validation, permission scoping, timeouts, sandboxing, human approval for privileged actions, tests |
| Module Lab | Triage all 48 dev tickets through the full pipeline, then route invoice tickets through the tool loop with an approval queue |
| Project Milestone, Interview Questions, Other Tools, Coming Up | Wrap-up |
Setup
All examples run in order from the repository root, in one shell session. They produce real output without an API key: the parsers, validators, loops, and guards are deterministic code, and every "model" in the offline runs is either ScriptedLLM or a keyword baseline. Neither is a language model. They test plumbing, and every place they appear says so. The scripts that call a hosted model run through llm.chat and need a key; their output is marked as illustrative.
source /home/claude/venv/bin/activate # or your own venv with requirements.txt installed
cd supportdesk
export PYTHONPATH=.
python -c "import pydantic, openai, httpx; print(pydantic.VERSION, openai.__version__, httpx.__version__)"
Code explained
- In simple words: switch on the course environment, stand in the repo, and let Python find both
supportdeskandexamples. - What happens:
PYTHONPATH=.makessupportdesk.*importable and lets module scripts import each other asexamples.m06_...(a namespace package, no__init__.pyneeded). The last line prints the versions this module was built with.httpxcomes withopenai; we use it once to capture requests offline. - Comes out:
Part A: Getting reliable structure
Why free text breaks downstream systems
Brightlane's help desk routes each ticket to a queue with a response deadline. The router is ordinary Python written by someone who expects clean data. A downstream system is any code that consumes the model's answer: a router, a database write, a billing call. It cannot "read between the lines" the way Maya can.
"""Why free text breaks downstream code: a real router fed three kinds of model output."""
from __future__ import annotations
import json
import re
import traceback
# The support desk's router: which queue a ticket goes to and how fast it must be answered.
QUEUES = {"billing": "finance-queue", "cancellation": "retention-queue", "account_access": "access-queue",
"bug": "engineering-queue", "how_to": "tier1-queue", "feature_request": "product-queue"}
SLA_HOURS = {"urgent": 1, "high": 4, "normal": 24, "low": 72}
def route(triage: dict) -> str:
"""Downstream code written by someone who expects clean data."""
queue = QUEUES[triage["category"]]
hours = SLA_HOURS[triage["priority"]]
owner = "human" if triage["needs_human"] else "assistant"
return f"{queue} within {hours}h, handled by {owner}"
FREE_TEXT = """Category: Billing
Priority: High (money is wrong)
This customer needs a human because they want a refund."""
LOOKS_LIKE_JSON = """Sure! Here is the triage:
{"category": "billing", "priority": "high", "needs_human": true}"""
VALID = '{"category": "billing", "priority": "high", "needs_human": true}'
def try_route(label: str, text: str) -> None:
print(f"--- {label}")
try:
print("OK:", route(json.loads(text)))
except Exception as exc: # show the failure the way a log would
print("CRASH:", "".join(traceback.format_exception_only(exc)).strip())
try_route("valid JSON", VALID)
try_route("JSON with a friendly preamble", LOOKS_LIKE_JSON)
try_route("free text", FREE_TEXT)
# The usual "quick fix": scrape the free text with a regex. It runs, then fails later.
print("--- free text scraped with a regex")
fields = {k.lower(): v.strip() for k, v in re.findall(r"^(\w+):\s*(.+)$", FREE_TEXT, re.MULTILINE)}
print("scraped:", fields)
try:
print("OK:", route(fields))
except Exception as exc:
print("CRASH:", "".join(traceback.format_exception_only(exc)).strip())
Code explained
- In simple words: we feed one honest router three answers a model might plausibly give, and watch which ones it survives.
- What happens:
route()indexes two dictionaries and reads a boolean.try_routerunsjson.loadsfirst, as most first versions do. The last block tries the classic quick fix, a regex that scrapesKey: valuelines out of free text. - Comes out:
--- valid JSON
OK: finance-queue within 4h, handled by human
--- JSON with a friendly preamble
CRASH: json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)
--- free text
CRASH: json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)
--- free text scraped with a regex
scraped: {'category': 'Billing', 'priority': 'High (money is wrong)'}
CRASH: KeyError: 'Billing'
Only strict JSON survives. The preamble breaks json.loads at character 0. The regex "fix" is worse: it runs without error, produces 'Billing' with a capital B and 'High (money is wrong)', drops needs_human entirely, and fails one step later with a KeyError. In production that is the dangerous kind of failure, because the error surfaces far from its cause. The rest of Part A removes each failure in turn: say exactly what shape you want, ask the provider to enforce it, parse defensively, validate, repair, and salvage.
Schema design: required, optional, enumerated, bounded fields
A schema is a machine-readable description of the shape an answer must have. The course defines its schemas once, as pydantic models, in supportdesk/schemas.py. Pydantic is a Python library that turns type hints into validators: you declare fields, and it checks data against them and produces a JSON Schema. JSON Schema is the standard format providers accept for structured output and tool parameters. This module introduces the file.
supportdesk/schemas.py
"""Typed output schemas for the support assistant, validated with pydantic."""
from __future__ import annotations
from typing import Literal
from pydantic import BaseModel, ConfigDict, Field
Category = Literal["billing", "cancellation", "account_access", "bug", "how_to", "feature_request"]
Priority = Literal["low", "normal", "high", "urgent"]
class Triage(BaseModel):
"""What the assistant decides about one incoming ticket."""
model_config = ConfigDict(extra="forbid")
category: Category = Field(description="The single best category for the ticket.")
priority: Priority = Field(description="urgent: many users blocked or security risk now; high: one user blocked or money wrong; normal: needs an answer; low: question or idea.")
language: str = Field(pattern=r"^[a-z]{2}$", description="ISO 639-1 code of the customer's language, for example 'en'.")
summary: str = Field(min_length=5, max_length=200, description="One sentence in English describing the request.")
needs_human: bool = Field(description="True if an agent must act (refund, account change, legal, security) or the help center cannot answer it.")
class DraftReply(BaseModel):
"""A reply the assistant proposes; a human agent reviews it before sending."""
model_config = ConfigDict(extra="forbid")
reply: str = Field(min_length=1, max_length=2000)
cited_articles: list[str] = Field(default_factory=list, max_length=3, description="Help-center article ids the reply relies on.")
confidence: Literal["low", "medium", "high"]
def triage_json_schema() -> dict:
"""The JSON Schema for Triage, ready for a provider's structured-output mode."""
return Triage.model_json_schema()
Code explained
- In simple words: two forms the assistant fills in, with the rules for each box written right on the form.
- What happens:
CategoryandPriorityareLiteraltypes: enumerations, fields that may only take one of a fixed list of strings. They reuse the labels fromsupportdesk/data.py, so gold labels and model output use the same vocabulary.Triageis what the assistant decides about one ticket.model_config = ConfigDict(extra="forbid")rejects unknown keys, which turns into"additionalProperties": falsein the schema. Every field is required (no default).languageis bounded by a regular expression (two lowercase letters),summaryby a length range of 5 to 200. Thedescriptionstrings are not comments: they are sent to the model and are part of the prompt. The priority description is a small rubric, because "high" means nothing unless you define it.DraftReplyis the reply a human reviews.cited_articlesis optional: it has a default (an empty list), so the model may leave it out. It is also bounded (max_length=3on a list means at most three items).triage_json_schema()returns the JSON Schema forTriage, ready to hand to a provider.
- Comes out: nothing on import. The next block prints the real schema.
python -c "import json; from supportdesk.schemas import triage_json_schema; print(json.dumps(triage_json_schema(), indent=2))"
Code explained
- In simple words: ask pydantic what it will actually send.
- What happens:
Triage.model_json_schema()walks the fields and emits JSON Schema (draft 2020-12 style). Printing it is the only reliable way to know what the model will see; do not guess from the Python class. - Comes out:
{
"additionalProperties": false,
"description": "What the assistant decides about one incoming ticket.",
"properties": {
"category": {
"description": "The single best category for the ticket.",
"enum": [
"billing",
"cancellation",
"account_access",
"bug",
"how_to",
"feature_request"
],
"title": "Category",
"type": "string"
},
"priority": {
"description": "urgent: many users blocked or security risk now; high: one user blocked or money wrong; normal: needs an answer; low: question or idea.",
"enum": [
"low",
"normal",
"high",
"urgent"
],
"title": "Priority",
"type": "string"
},
"language": {
"description": "ISO 639-1 code of the customer's language, for example 'en'.",
"pattern": "^[a-z]{2}$",
"title": "Language",
"type": "string"
},
"summary": {
"description": "One sentence in English describing the request.",
"maxLength": 200,
"minLength": 5,
"title": "Summary",
"type": "string"
},
"needs_human": {
"description": "True if an agent must act (refund, account change, legal, security) or the help center cannot answer it.",
"title": "Needs Human",
"type": "boolean"
}
},
"required": [
"category",
"priority",
"language",
"summary",
"needs_human"
],
"title": "Triage",
"type": "object"
}
Read it the way the model will: each property has a type, enums list every allowed value, bounds appear as pattern, minLength, maxLength, and required lists all five fields. Pydantic also adds title keys, which cost tokens and carry no information here; we strip them later.
A few design rules for the fields you add:
| Situation | Use this | Why |
|---|---|---|
| A value from a known, small set (category, priority) | An enum (Literal[...]) | The model cannot invent "Billing" or "critical"; with strict mode the provider enforces it, and validation catches it everywhere else |
| A value your code must always have | Required, no default | A missing field should fail loudly, not silently become a default |
| A value that is often absent (cited articles) | Optional with a safe default | Forces nothing the model cannot know; in strict mode make it nullable (below) |
| Free text shown to a person (summary, reply) | str with min_length and max_length | Stops empty strings and runaway essays |
| A number that moves money or quotas | Numeric bounds (gt, le) plus a business check in code | A schema bound is coarse; the real limit (the invoice amount) lives in your database |
| A classification the model should justify | Put a short reasoning field before the label, or keep reasoning out of the schema | Research (below) finds strict formats can hurt reasoning; measure on your task |
Now look at how optional fields interact with provider strict modes.
"""Look at the schemas the assistant uses, and turn one into a strict response_format."""
from __future__ import annotations
import json
from examples.m06_structured import response_format_for, validate_output
from supportdesk.schemas import DraftReply, Triage, triage_json_schema
schema = triage_json_schema()
print("Triage required:", schema["required"])
print("Triage enums:", {k: v["enum"] for k, v in schema["properties"].items() if "enum" in v})
print("Triage bounds:", {k: {b: v[b] for b in ("pattern", "minLength", "maxLength") if b in v}
for k, v in schema["properties"].items() if any(b in v for b in ("pattern", "minLength"))})
print("\nDraftReply required (pydantic):", DraftReply.model_json_schema()["required"])
fmt = response_format_for(DraftReply)
print("DraftReply response_format for strict mode:")
print(json.dumps(fmt, indent=2))
# A strict-mode reply must include every field, so optional ones arrive as null.
from_model = {"reply": "You can export a board from Board menu > Export.", "cited_articles": None, "confidence": "high"}
print("\nvalidated:", validate_output(from_model, DraftReply))
Code explained
- In simple words: print the schema's rules, then reshape
DraftReplyfor a provider that insists every field be present. - What happens: the first prints pull the required list, enums, and bounds out of the
Triageschema.response_format_for(DraftReply)(fromexamples/m06_structured.py, shown in full in the next section) builds aresponse_formatin strict mode, a provider setting that constrains generation to the schema. Groq documents two strict-mode rules: every field must be listed inrequired, and every object must setadditionalProperties: false; optional values are expressed as a union withnull. The helper does that reshaping. The last lines show the reverse trip: a strict-mode answer with"cited_articles": nullvalidates back into aDraftReplywhosecited_articlesis[]. - Comes out:
Native structured-output modes vs prompted JSON
There are two ways to get JSON from a model:
- Prompted JSON: you describe the format in the prompt ("reply with only a JSON object matching this schema") and hope. Every provider supports it. Nothing enforces it.
- Native structured output: you send the schema in the request's
response_format, and the provider uses it during generation. With constrained decoding, the provider masks out any next token that would break the schema, so the output is valid JSON of the right shape by construction. With a best-effort mode, the schema guides the model but invalid output is still possible.
What each course provider documents, checked on 21 September 2026 (these pages change often, so check them again):
| Provider (OpenAI-compatible endpoint) | response_format with json_schema | Strict, constrained output | Notes from the docs |
|---|---|---|---|
| Groq | Yes | Yes on GPT-OSS 20B and 120B (the course default) and one Qwen model; other models best-effort | "Streaming and tool use are not currently supported with Structured Outputs." Best-effort mode can return HTTP 400 "Generated JSON does not match the expected schema." (Groq structured outputs) |
| Gemini | Yes, through the OpenAI library; Google's example passes a pydantic class to client.beta.chat.completions.parse | The compatibility page does not document strict. The native API's structured output lists supported keywords (enum, minimum, maximum, format, anyOf, $ref, additionalProperties); minLength and pattern are not in that list | The compatibility layer is labeled beta. The native docs warn "Very large or deeply nested schemas may be rejected" and "always validate values in your application." (Gemini OpenAI compatibility, Gemini structured output) |
| Ollama | Yes: "Structured outputs work through the OpenAI-compatible API via response_format" | strict is not listed among supported fields | Also supports format with a JSON Schema on its native API. Note that tool_choice is listed as not supported (Ollama OpenAI compatibility, Ollama structured outputs) |
Two consequences follow. First, even the best case guarantees syntax and shape, not truth: a strictly valid triage can still say billing for a login problem. Second, for two of three providers, keywords like pattern or maxLength may be ignored during generation. Either way, you validate on your side every time.
Here is the helper file that the rest of Part A uses. It is long, so it sits in a collapsible block; each function is explained after it.
examples/m06_structured.py
"""Module 6 helpers: turn model text into validated objects, or fail loudly.
Four layers, each usable on its own:
response_format_for build a json_schema response_format for a provider
parse_json a tolerant parser for hostile model output
generate_structured call, parse, validate, and repair with an attempt ceiling
salvage keep the fields that are valid when the whole object is not
"""
from __future__ import annotations
import copy
import json
import re
from dataclasses import dataclass, field
from typing import Any, Callable
from pydantic import BaseModel, ValidationError
from supportdesk.llm import ChatResult, Usage
# ---------------------------------------------------------------------------
# 1. Native structured output: a json_schema response_format
# ---------------------------------------------------------------------------
def strict_schema(model_cls: type[BaseModel]) -> dict[str, Any]:
"""The model's JSON Schema, reshaped for strict mode.
Strict mode (Groq documents it for gpt-oss) needs every property listed in
`required` and `additionalProperties: false` on every object. A field that
was optional becomes nullable instead: the model must write it, but may
write null. `validate_output` turns those nulls back into defaults.
"""
schema = copy.deepcopy(model_cls.model_json_schema())
def fix(node: Any) -> None:
if isinstance(node, dict):
if isinstance(node.get("title"), str): # a schema title, not a property named "title"
node.pop("title")
if node.get("type") == "object" and "properties" in node:
was_required = set(node.get("required", []))
for name, prop in node["properties"].items():
if name not in was_required:
prop.pop("default", None)
wrapped: dict[str, Any] = {"anyOf": [prop, {"type": "null"}]}
if "description" in prop: # keep the description where the model reads it
wrapped = {"description": prop.pop("description"), **wrapped}
node["properties"][name] = wrapped
node["required"] = list(node["properties"])
node["additionalProperties"] = False
for value in node.values():
fix(value)
elif isinstance(node, list):
for value in node:
fix(value)
fix(schema)
return schema
def response_format_for(model_cls: type[BaseModel], strict: bool = True) -> dict[str, Any]:
"""The `response_format` value for llm.chat(..., response_format=...)."""
return {
"type": "json_schema",
"json_schema": {"name": model_cls.__name__, "schema": strict_schema(model_cls), "strict": strict},
}
# ---------------------------------------------------------------------------
# 2. A tolerant parser, built in stages so each stage can be measured
# ---------------------------------------------------------------------------
FENCE = re.compile(r"```(?:json|JSON)?\s*\n?(.*?)(?:```|$)", re.DOTALL)
SMART_QUOTES = str.maketrans({"\u201c": '"', "\u201d": '"', "\u2018": "'", "\u2019": "'"})
class ParseError(ValueError):
"""Raised when no JSON object can be recovered from the text."""
@dataclass
class Parsed:
value: dict[str, Any]
repairs: list[str] = field(default_factory=list) # what the parser had to fix, for logging
def strip_fences(text: str) -> tuple[str, bool]:
"""Return the inside of the first ```json fence, if there is one."""
match = FENCE.search(text)
return (match.group(1), True) if match else (text, False)
def extract_object(text: str) -> tuple[str, bool]:
"""Cut from the first '{' to its matching '}', skipping braces inside strings.
If the text ends before the object closes, return everything from '{' on:
the truncation stage may still be able to close it.
"""
start = text.find("{")
if start == -1:
raise ParseError("no '{' in output")
depth, in_str, quote, escaped = 0, False, "", False
for i in range(start, len(text)):
ch = text[i]
if in_str:
if escaped:
escaped = False
elif ch == "\\":
escaped = True
elif ch == quote:
in_str = False
elif ch in "\"'":
in_str, quote = True, ch
elif ch == "{":
depth += 1
elif ch == "}":
depth -= 1
if depth == 0:
return text[start:i + 1], (start > 0 or i + 1 < len(text.rstrip()))
return text[start:], start > 0
def normalize(text: str) -> tuple[str, list[str]]:
"""Fix the JSON-adjacent syntax models write: single quotes, trailing commas,
Python literals, and smart quotes. Works token by token, so text inside
double-quoted strings is never touched."""
repairs: list[str] = []
if any(c in text for c in "\u201c\u201d\u2018\u2019"):
text = text.translate(SMART_QUOTES)
repairs.append("smart_quotes")
out: list[str] = []
i, n = 0, len(text)
while i < n:
ch = text[i]
if ch == '"': # copy a JSON string unchanged
j = i + 1
while j < n and text[j] != '"':
j += 2 if text[j] == "\\" else 1
out.append(text[i:j + 1])
i = j + 1
elif ch == "'": # 'single quoted' -> "double quoted"
j = i + 1
while j < n and text[j] != "'":
j += 2 if text[j] == "\\" else 1
out.append(json.dumps(text[i + 1:j].replace("\\'", "'")))
if "single_quotes" not in repairs:
repairs.append("single_quotes")
i = j + 1
elif ch == ",": # drop a comma that closes nothing
j = i + 1
while j < n and text[j] in " \t\r\n":
j += 1
if j < n and text[j] in "}]":
if "trailing_comma" not in repairs:
repairs.append("trailing_comma")
else:
out.append(ch)
i += 1
elif text.startswith(("True", "False", "None"), i) and not (i and text[i - 1].isalnum()):
word = next(w for w in ("True", "False", "None") if text.startswith(w, i))
out.append({"True": "true", "False": "false", "None": "null"}[word])
if "python_literals" not in repairs:
repairs.append("python_literals")
i += len(word)
else:
out.append(ch)
i += 1
return "".join(out), repairs
def close_truncated(text: str) -> tuple[str, bool]:
"""Close an object that was cut off mid-stream (for example by max_tokens).
Closes an open string, drops a dangling key or comma, then appends the
closing brackets in the right order. The value may be incomplete, so the
caller must treat a closed object as suspect.
"""
stack: list[str] = []
in_str, escaped = False, False
for ch in text:
if in_str:
if escaped:
escaped = False
elif ch == "\\":
escaped = True
elif ch == '"':
in_str = False
elif ch == '"':
in_str = True
elif ch in "{[":
stack.append("}" if ch == "{" else "]")
elif ch in "}]" and stack:
stack.pop()
if not stack and not in_str:
return text, False
fixed = text + ('"' if in_str else "")
fixed = re.sub(r',\s*"[^"]*"\s*:?\s*$', "", fixed) # a key with no value yet
fixed = re.sub(r'[,:]\s*$', "", fixed.rstrip()) # a dangling comma or colon
return fixed + "".join(reversed(stack)), True
PARSER_STAGES = ("plain", "fences", "extract", "normalize", "truncation")
def parse_json(text: str, stages: tuple[str, ...] = PARSER_STAGES) -> Parsed:
"""Recover one JSON object from model output, applying only the listed stages."""
repairs: list[str] = []
candidate = text.strip()
if "fences" in stages:
candidate, fenced = strip_fences(candidate)
if fenced:
repairs.append("fences")
if "extract" in stages:
candidate, cut = extract_object(candidate)
if cut:
repairs.append("surrounding_text")
if "normalize" in stages:
candidate, fixes = normalize(candidate)
repairs += fixes
if "truncation" in stages:
candidate, closed = close_truncated(candidate)
if closed:
repairs.append("closed_truncated")
try:
value = json.loads(candidate)
except json.JSONDecodeError as exc:
raise ParseError(f"{exc.msg} at char {exc.pos}") from exc
if not isinstance(value, dict):
raise ParseError(f"expected an object, got {type(value).__name__}")
return Parsed(value, repairs)
# ---------------------------------------------------------------------------
# 3. Validation, repair loop with a ceiling, and salvage
# ---------------------------------------------------------------------------
def validate_output(data: dict[str, Any], model_cls: type[BaseModel]) -> BaseModel:
"""Validate with pydantic after turning strict-mode nulls back into defaults."""
cleaned = {
k: v for k, v in data.items()
if not (v is None and k in model_cls.model_fields and not model_cls.model_fields[k].is_required())
}
return model_cls.model_validate(cleaned)
def describe_errors(exc: ValidationError) -> list[str]:
"""Short, model-readable error lines such as "priority: Input should be ..."."""
lines = []
for err in exc.errors():
where = ".".join(str(p) for p in err["loc"]) or "(root)"
lines.append(f"{where}: {err['msg']}")
return lines
@dataclass
class StructuredResult:
value: BaseModel | None
attempts: int
errors: list[list[str]] # the errors of each failed attempt
repairs: list[str] # parser repairs on the attempt that succeeded
usage: Usage
last_data: dict[str, Any] | None # the last parsed object, for salvage
def generate_structured(
llm: Callable[..., ChatResult],
messages: list[dict[str, Any]],
model_cls: type[BaseModel],
*,
max_attempts: int = 3,
**chat_kwargs: Any,
) -> StructuredResult:
"""Ask, parse, validate. On failure, show the model its errors and ask again,
at most `max_attempts` calls in total. Never loops forever, never raises
for bad output: the caller decides what a failure means."""
history = list(messages)
usage = Usage()
all_errors: list[list[str]] = []
last_data: dict[str, Any] | None = None
for attempt in range(1, max_attempts + 1):
result = llm(history, **chat_kwargs)
usage.input_tokens += result.usage.input_tokens
usage.output_tokens += result.usage.output_tokens
try:
parsed = parse_json(result.text)
last_data = parsed.value
value = validate_output(parsed.value, model_cls)
return StructuredResult(value, attempt, all_errors, parsed.repairs, usage, last_data)
except ParseError as exc:
errors = [f"(root): could not parse JSON: {exc}"]
except ValidationError as exc:
errors = describe_errors(exc)
all_errors.append(errors)
history += [
{"role": "assistant", "content": result.text},
{"role": "user", "content": "Your reply did not match the schema:\n- " + "\n- ".join(errors)
+ "\nReply again with only the corrected JSON object."},
]
return StructuredResult(None, max_attempts, all_errors, [], usage, last_data)
def salvage(data: dict[str, Any], model_cls: type[BaseModel]) -> tuple[dict[str, Any], dict[str, str]]:
"""Split a parsed object into fields that pass validation and fields that do not.
Returns (valid_fields, problems). A field is valid if the full validation
reported no error located at it. Unknown keys and missing required fields
are listed in problems.
"""
try:
return validate_output(data, model_cls).model_dump(), {}
except ValidationError as exc:
problems: dict[str, str] = {}
for err in exc.errors():
name = str(err["loc"][0]) if err["loc"] else "(root)"
problems.setdefault(name, err["msg"])
valid = {k: v for k, v in data.items() if k in model_cls.model_fields and k not in problems}
return valid, problems
Code explained
- In simple words: four tools for one job: ask for the right shape, recover JSON from messy text, check it, and when it is wrong, either ask again or keep what is usable.
- What happens:
strict_schema(model_cls)copies pydantic's schema and walks it recursively. It removestitlekeys, lists every property inrequired, setsadditionalProperties: falseon every object, and wraps each formerly optional property inanyOf: [original, null], keeping itsdescriptionon the outside where the model reads it. It also descends into$defs, so nested models get the same treatment.response_format_for(model_cls, strict=True)wraps that schema in the{"type": "json_schema", "json_schema": {...}}envelope thatllm.chat(..., response_format=...)passes straight through to the provider.ParseErroris the one exception the parser raises.Parsedcarries the recovered object plusrepairs, a list of what was fixed. Logging repairs matters: a sudden rise inclosed_truncatedtells youmax_tokensis too low long before anyone reads a bad reply.strip_fencesreturns the inside of the first Markdown code fence (with or withoutjson), and tolerates a missing closing fence.extract_objectfinds the first{and walks forward counting braces, but skips braces inside quoted strings, so a summary containing{{name}}does not end the object early. Text before and after the object (a preamble or a sign-off) is dropped. If the object never closes, it returns the tail for the truncation stage.normalizerewrites the almost-JSON that models write: smart quotes become straight quotes,'single quoted'strings become"double quoted"viajson.dumps(so inner characters are escaped correctly), trailing commas before}or]are dropped, and Python'sTrue,False,Nonebecometrue,false,null. It scans character by character and copies double-quoted strings untouched, so an apostrophe inside"Customer's card"is safe.close_truncatedhandles output cut off mid-stream. It tracks open brackets and strings, closes an open string, drops a dangling key, comma, or colon, and appends the missing closers in reverse order. The result may be missing fields or hold a half-written value, so the repair is logged asclosed_truncatedand treated as suspect later.parse_json(text, stages)runs the enabled stages in order and thenjson.loads. Making the stages a parameter is what lets us measure each one.validate_outputremovesnullvalues for fields that have defaults (the strict-mode round trip) and callsmodel_validate.describe_errorsturns a pydanticValidationErrorinto short lines likepriority: Input should be ..., which are what the repair loop sends back to the model.generate_structuredis the repair loop: call, parse, validate; on failure, append the bad answer and a message listing the errors, and try again, at mostmax_attemptscalls. It returns aStructuredResultwith the value (orNone), attempt count, all errors, total usage, and the last parsed object for salvage. It never raises for bad model output.salvagevalidates the whole object, reads which fields the errors point at, and returns the fields that passed plus a dictionary of problems.
- Comes out: nothing on its own; the next examples run it.
First, see what native structured output looks like on the wire. examples/m06_wire.py swaps the HTTP transport under the openai client for a fake one that records the request and returns a canned reply, so the request is real (built by llm.chat) and only the reply is invented.
examples/m06_wire.py
"""See the exact JSON that llm.chat sends, without an API key or network.
`capture_wire` swaps the HTTP transport under the openai client for a fake one
that records each request body and answers with a canned response. The request
is real (built by llm.chat and the openai library); only the reply is canned.
"""
from __future__ import annotations
import json
from collections.abc import Iterator
from contextlib import contextmanager
from typing import Any
import httpx
from openai import OpenAI
import supportdesk.llm as llm
@contextmanager
def capture_wire(canned_message: dict[str, Any], finish_reason: str = "stop") -> Iterator[list[dict[str, Any]]]:
"""Record request bodies sent by llm.chat; reply with `canned_message` every time."""
bodies: list[dict[str, Any]] = []
def handler(request: httpx.Request) -> httpx.Response:
bodies.append(json.loads(request.content))
return httpx.Response(200, json={
"id": "chatcmpl-canned", "object": "chat.completion", "created": 0, "model": bodies[-1]["model"],
"choices": [{"index": 0, "message": {"role": "assistant", **canned_message}, "finish_reason": finish_reason}],
"usage": {"prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0},
})
def fake_client(provider: str) -> OpenAI:
return OpenAI(api_key="not-a-real-key", base_url=llm.PROVIDERS[provider]["base_url"],
http_client=httpx.Client(transport=httpx.MockTransport(handler)))
original = llm.make_client
llm.make_client = fake_client
try:
yield bodies
finally:
llm.make_client = original
Code explained
- In simple words: a tap on the wire between
llm.chatand the provider, so we can read exactly what leaves your machine. - What happens:
capture_wireis a context manager. Inside it,supportdesk.llm.make_clientis replaced byfake_client, which builds a realOpenAIclient whosehttp_clientuseshttpx.MockTransport(handler). Thehandlerdecodes and stores each request body, then returns a minimal Chat Completions response containingcanned_message. On exit the realmake_clientis restored. No key or network is needed, and the canonicalllm.pyis not edited. - Comes out: nothing by itself; it yields the list of captured request bodies.
"""The request body llm.chat sends for native structured output (captured offline)."""
from __future__ import annotations
import json
from examples.m06_native import native_messages
from examples.m06_structured import response_format_for
from examples.m06_wire import capture_wire
from supportdesk.llm import chat
from supportdesk.schemas import Triage
canned = {"content": '{"category":"billing","priority":"high","language":"en",'
'"summary":"Customer was charged twice for the Team plan.","needs_human":true}'}
with capture_wire(canned) as bodies:
result = chat(native_messages("Subject: Charged twice\n\nMy card was charged 288 USD twice."),
provider="groq", response_format=response_format_for(Triage), max_tokens=2000)
body = bodies[0]
print(json.dumps({k: v for k, v in body.items() if k != "response_format"}, indent=2))
print("response_format.type:", body["response_format"]["type"])
print("response_format.json_schema keys:", list(body["response_format"]["json_schema"]))
print("\nChatResult.text:", result.text)
print("finish_reason:", result.finish_reason, "| provider:", result.provider, "| model:", result.model)
Code explained
- In simple words: send one native structured-output request through
llm.chatand read the body that would go to Groq. - What happens:
native_messages(fromexamples/m06_native.py, next block) builds a system and user message.chat(..., response_format=response_format_for(Triage), max_tokens=2000)runs the real request code. The script prints the body without the long schema, then theresponse_formatenvelope's keys, then theChatResultthatllm.chatbuilt from the canned reply. - Comes out:
{
"messages": [
{
"role": "system",
"content": "You triage Brightlane support tickets. Brightlane is a project-management SaaS."
},
{
"role": "user",
"content": "Subject: Charged twice\n\nMy card was charged 288 USD twice."
}
],
"model": "openai/gpt-oss-120b",
"max_tokens": 2000,
"temperature": 0.0
}
response_format.type: json_schema
response_format.json_schema keys: ['name', 'schema', 'strict']
ChatResult.text: {"category":"billing","priority":"high","language":"en","summary":"Customer was charged twice for the Team plan.","needs_human":true}
finish_reason: stop | provider: groq | model: openai/gpt-oss-120b
llm.chat only sends parameters you set (temperature defaults to 0.0 in the helper), and response_format goes through unchanged. The text is the canned reply, not model output. max_tokens=2000 is deliberately generous: gpt-oss is a reasoning model, and on providers that count reasoning tokens against the output limit, a tight limit is the most common cause of truncated JSON.
The live comparison script sends the same tickets both ways.
"""Native structured output vs prompted JSON for ticket triage, through llm.chat.
Needs a provider key (or a local Ollama). Usage:
LLM_PROVIDER=groq GROQ_API_KEY=... python examples/m06_native.py
"""
from __future__ import annotations
import json
import os
import sys
from pydantic import ValidationError
from examples.m06_structured import ParseError, parse_json, response_format_for, validate_output
from supportdesk.data import load_tickets
from supportdesk.llm import PROVIDERS, chat, resolve
from supportdesk.schemas import Triage, triage_json_schema
SYSTEM = "You triage Brightlane support tickets. Brightlane is a project-management SaaS."
def native_messages(ticket_text: str) -> list[dict]:
return [{"role": "system", "content": SYSTEM}, {"role": "user", "content": ticket_text}]
def prompted_messages(ticket_text: str) -> list[dict]:
schema = json.dumps(triage_json_schema(), separators=(",", ":"))
return [{"role": "system", "content": f"{SYSTEM}\nReply with only a JSON object that matches this JSON Schema, "
f"with no other text:\n{schema}"}, {"role": "user", "content": ticket_text}]
def triage_once(ticket_text: str, mode: str) -> tuple[str, str]:
"""Return (outcome, detail) for one ticket in 'native' or 'prompted' mode."""
if mode == "native":
result = chat(native_messages(ticket_text), response_format=response_format_for(Triage), max_tokens=2000)
else:
result = chat(prompted_messages(ticket_text), max_tokens=2000)
try:
parsed = parse_json(result.text)
triage = validate_output(parsed.value, Triage)
return "valid", f"{triage.category}/{triage.priority} repairs={parsed.repairs} out_tokens={result.usage.output_tokens}"
except (ParseError, ValidationError) as exc:
return "invalid", str(exc).splitlines()[0]
if __name__ == "__main__":
provider, model = resolve()
key_env = PROVIDERS[provider]["key_env"]
if key_env and not os.environ.get(key_env):
sys.exit(f"Set {key_env} (or LLM_PROVIDER=ollama) to run this example against {provider}/{model}.")
print(f"provider={provider} model={model}")
for ticket in load_tickets("dev")[:5]:
for mode in ("native", "prompted"):
outcome, detail = triage_once(ticket.text, mode)
print(f"{ticket.id} {mode:<8} {outcome:<7} {detail} gold={ticket.gold['category']}/{ticket.gold['priority']}")
Code explained
- In simple words: triage five real dev tickets twice, once with the schema enforced by the provider and once with the schema only described in the prompt, and check both with the same parser and validator.
- What happens:
native_messageshas no format instructions at all; the schema travels inresponse_format.prompted_messagespastes the compact schema into the system prompt instead.triage_oncecallschat, thenparse_jsonandvalidate_output, and reports validity, the parser repairs needed, and output tokens. Without a key the script stops with a clear message. - Comes out: without a key (what this build can run):
Set GROQ_API_KEY (or LLM_PROVIDER=ollama) to run this example against groq/openai/gpt-oss-120b.
With a key, each line shows one ticket and mode. Illustrative sample run (not captured in this build; produced for teaching). Your output will differ.
Copy
provider=groq model=openai/gpt-oss-120b
T-1001 native valid billing/high repairs=[] out_tokens=... gold=billing/high
T-1001 prompted valid billing/high repairs=[] out_tokens=... gold=billing/high
T-1002 native valid billing/normal repairs=[] out_tokens=... gold=billing/normal
What to look for in your run: the repairs column for the prompted mode (fences and preambles show up there first), and any invalid lines. Five tickets is a smoke test, not a comparison; the lab's --live flag runs all 48.
| Situation | Use this | Why |
|---|---|---|
| Provider documents strict mode for your model (Groq gpt-oss) | Native json_schema with strict: true, plus local validation | Shape is guaranteed; validation still catches semantic bounds and provider drift |
Provider supports json_schema but not documented strict (Gemini compatibility layer, Ollama) | Native json_schema, tolerant parser, validation, repair loop | Shape is likely but not guaranteed; your code must not assume it |
| You also need tools in the same call on Groq | Tools, and put the structured answer in a tool's arguments or a separate call | Groq documents that structured outputs and tool use cannot be combined |
| Provider or model has no structured output | Prompted JSON, tolerant parser, repair loop | Works everywhere; costs more tokens and retries |
| Reasoning-heavy task where format seems to hurt quality | Let the model reason in free text, then extract to JSON in a second, cheap call | Separates thinking from formatting; measure both on your eval set |
Validation with typed models
Parsing answers "is this JSON?". Validation answers "is this JSON the object my code expects?". Pydantic does it in one call, Triage.model_validate(data), which either returns a typed object or raises ValidationError with one entry per problem. You saw those errors in the repair message above and will see them again below.
One behavior to know: by default pydantic runs in lax mode, which coerces obvious conversions. In the hostile corpus below, "needs_human": "yes" validates as True. That is convenient for booleans, but for a field like an amount you may prefer an error. Triage.model_validate(data, strict=True) turns coercion off for one call; ConfigDict(strict=True) turns it off for a model. The course's schemas keep lax mode and rely on enums and bounds for the fields that matter.
Parsing hostile output: fences, preambles, trailing commas, truncation
Even with native modes you will parse text from models that do not support them, from best-effort modes, from logs, and from outputs cut off by max_tokens. To build the parser honestly, we need a test set. examples/m06_hostile.py holds 24 outputs for a Triage object, each imitating a failure seen in practice, and a scoreboard that turns on parser stages one at a time.
examples/m06_hostile.py
"""A corpus of 24 hostile model outputs for a Triage object, and the parser scoreboard.
Each case imitates a failure seen in real model output. Run it to see how many
cases each parser stage recovers, and how many then pass Triage validation.
"""
from __future__ import annotations
from pydantic import ValidationError
from examples.m06_structured import PARSER_STAGES, ParseError, parse_json, validate_output
from supportdesk.schemas import Triage
GOOD = '{"category": "billing", "priority": "high", "language": "en", "summary": "Customer was charged twice for the Team plan.", "needs_human": true}'
HOSTILE: dict[str, str] = {
"clean": GOOD,
"clean_pretty": '{\n "category": "how_to",\n "priority": "normal",\n "language": "de",\n "summary": "Asks how to export a board to CSV.",\n "needs_human": false\n}',
"fence_json": "```json\n" + GOOD + "\n```",
"fence_bare": "```\n" + GOOD + "\n```",
"preamble": "Sure! Here is the triage for this ticket:\n\n" + GOOD,
"postamble": GOOD + "\n\nLet me know if you'd like me to draft a reply as well.",
"preamble_fence_post": "Here's the JSON you asked for:\n```json\n" + GOOD + "\n```\nI marked it high because money is wrong.",
"trailing_comma": '{"category": "billing", "priority": "high", "language": "en", "summary": "Customer was charged twice.", "needs_human": true,}',
"trailing_comma_pretty": '{\n "category": "bug",\n "priority": "urgent",\n "language": "en",\n "summary": "Boards fail to load for the whole team.",\n "needs_human": true,\n}',
"single_quotes": "{'category': 'billing', 'priority': 'high', 'language': 'en', 'summary': 'Customer was charged twice.', 'needs_human': true}",
"python_dict": "{'category': 'bug', 'priority': 'high', 'language': 'es', 'summary': 'Mobile app crashes on login.', 'needs_human': False}",
"smart_quotes": "{\u201ccategory\u201d: \u201cbilling\u201d, \u201cpriority\u201d: \u201chigh\u201d, \u201clanguage\u201d: \u201cen\u201d, \u201csummary\u201d: \u201cCustomer was charged twice.\u201d, \u201cneeds_human\u201d: true}",
"braces_in_string": 'Result: {"category": "how_to", "priority": "low", "language": "en", "summary": "Asks what {{name}} means in templates.", "needs_human": false} Done.',
"truncated_in_value": '{"category": "billing", "priority": "high", "language": "en", "summary": "Customer was charged tw',
"truncated_after_comma": '{"category": "billing", "priority": "high", "language": "en", "summary": "Customer was charged twice.",',
"truncated_in_key": '{"category": "account_access", "priority": "high", "language": "ja", "summary": "Cannot log in after SSO change.", "needs_hu',
"fence_truncated": "```json\n{\"category\": \"cancellation\", \"priority\": \"normal\", \"language\": \"en\", \"summary\": \"Wants to cancel the annual plan",
"wrong_enum": '{"category": "Billing", "priority": "critical", "language": "en", "summary": "Customer was charged twice.", "needs_human": true}',
"wrong_types": '{"category": "billing", "priority": "high", "language": "English", "summary": "Charged twice.", "needs_human": "yes"}',
"extra_field": '{"category": "billing", "priority": "high", "language": "en", "summary": "Customer was charged twice.", "needs_human": true, "confidence": 0.9}',
"array_wrapped": '[' + GOOD + ']',
"prose_only": "This looks like a billing problem with high priority; a human should refund the duplicate charge.",
"refusal": "I'm sorry, but I can't help with that request.",
"apostrophe_in_single": "{'category': 'billing', 'priority': 'high', 'language': 'en', 'summary': 'Customer's card was charged twice.', 'needs_human': true}",
}
LADDER = [PARSER_STAGES[:i] for i in range(1, len(PARSER_STAGES) + 1)]
def score(stages: tuple[str, ...]) -> tuple[int, int, list[str]]:
"""How many cases parse to an object, and how many of those validate as Triage."""
parsed = valid = 0
failures = []
for name, text in HOSTILE.items():
try:
data = parse_json(text, stages).value
except ParseError:
failures.append(name)
continue
parsed += 1
try:
validate_output(data, Triage)
valid += 1
except ValidationError:
pass
return parsed, valid, failures
if __name__ == "__main__":
n = len(HOSTILE)
print(f"{n} hostile outputs\n")
print(f"{'parser stages':<50} {'parsed':>10} {'valid Triage':>14}")
for stages in LADDER:
parsed, valid, _ = score(stages)
print(f"{' + '.join(stages):<50} {parsed:>3}/{n} {parsed / n:>4.0%} {valid:>6}/{n} {valid / n:>4.0%}")
print("\nstill unparseable with every stage:", score(PARSER_STAGES)[2])
print("\nrepairs logged per case (full parser):")
for name, text in HOSTILE.items():
try:
print(f" {name:<22} {parse_json(text).repairs}")
except ParseError as exc:
print(f" {name:<22} ParseError: {exc}")
Code explained
- In simple words: a zoo of broken answers, and a referee that counts how many each version of the parser can rescue.
- What happens:
HOSTILEmaps a case name to a raw output: clean JSON (two cases), code fences, preambles and postambles, trailing commas, single quotes, a Python dict, smart quotes, braces inside a string, four kinds of truncation, three schema violations (wrong enum values, wrong types, an extra field), an array around the object, prose, a refusal, and one deliberately nasty case, a single-quoted string containing an apostrophe.LADDERlists the stage sets, from plainjson.loadsup to all five stages.scorecounts, for one stage set, how many cases parse into an object and how many of those also validate asTriage. - Comes out:
24 hostile outputs
parser stages parsed valid Triage
plain 5/24 21% 2/24 8%
plain + fences 8/24 33% 5/24 21%
plain + fences + extract 12/24 50% 9/24 38%
plain + fences + extract + normalize 17/24 71% 14/24 58%
plain + fences + extract + normalize + truncation 21/24 88% 14/24 58%
still unparseable with every stage: ['prose_only', 'refusal', 'apostrophe_in_single']
repairs logged per case (full parser):
clean []
clean_pretty []
fence_json ['fences']
fence_bare ['fences']
preamble ['surrounding_text']
postamble ['surrounding_text']
preamble_fence_post ['fences']
trailing_comma ['trailing_comma']
trailing_comma_pretty ['trailing_comma']
single_quotes ['single_quotes']
python_dict ['single_quotes', 'python_literals']
smart_quotes ['smart_quotes']
braces_in_string ['surrounding_text']
truncated_in_value ['closed_truncated']
truncated_after_comma ['closed_truncated']
truncated_in_key ['closed_truncated']
fence_truncated ['fences', 'closed_truncated']
wrong_enum []
wrong_types []
extra_field []
array_wrapped ['surrounding_text']
prose_only ParseError: no '{' in output
refusal ParseError: no '{' in output
apostrophe_in_single ParseError: Expecting ',' delimiter at char 83
Here is the same result as a before-and-after table. Each row adds one stage to the row above.
| Parser stages | Parsed (n=24) | Valid Triage (n=24) | What the new stage rescued |
|---|---|---|---|
json.loads only | 5 (21%) | 2 (8%) | Clean JSON; three parse but break the schema |
| + strip fences | 8 (33%) | 5 (21%) | Fenced output |
| + extract object | 12 (50%) | 9 (38%) | Preambles, postambles, text around braces, the array wrapper |
| + normalize | 17 (71%) | 14 (58%) | Trailing commas, single quotes, Python literals, smart quotes |
| + close truncation | 21 (88%) | 14 (58%) | Four truncated outputs parse, but none validates: all are missing fields |
Read this carefully, because it is easy to over-claim. These rates describe a corpus I constructed, one case per failure type, not the frequency of each failure from any real model. The ordering is what transfers: each stage rescues a distinct class, and no stage broke a case an earlier stage handled (the test suite checks that). The truncation stage raises the parse rate but not the valid rate, which is the correct outcome: a truncated object is incomplete, and pretending otherwise would be worse. Those four go to salvage.
Diagnosing the three that still fail, read the log line, then decide:
prose_onlyandrefusal:ParseError: no '{' in output. There is no JSON to recover; a parser should not invent one. These go to the repair loop, and if the model keeps refusing, to a human.apostrophe_in_single:Expecting ',' delimiter at char 83. The model used single quotes and also wroteCustomer's. The apostrophe ends the string early, and no local rule can tell an apostrophe from a closing quote in general. Fixing it would need guesswork that breaks other cases, so we let the repair loop ask again. Knowing when to stop adding heuristics is part of the design.
Repair loops with an attempt ceiling
A repair loop sends the model its own invalid answer plus the validation errors and asks for a corrected answer. An attempt ceiling is the maximum number of calls it may make. Without one, a model that keeps failing turns one ticket into an unbounded bill. ScriptedLLM stands in for the model here: it replays scripted replies, so this shows the loop's plumbing, not how often a real model fixes itself.
"""The repair loop with an attempt ceiling, driven by ScriptedLLM (plumbing only, not a model)."""
from __future__ import annotations
from examples.m06_structured import generate_structured
from supportdesk.schemas import Triage
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages
messages = [
{"role": "system", "content": "Triage the ticket. Reply with only a JSON object for the Triage schema."},
{"role": "user", "content": "Subject: Charged twice\n\nMy card was charged 288 USD twice on 3 September."},
]
BROKEN = ('Here you go: {"category": "billing", "priority": "critical", "language": "English", '
'"summary": "Charged twice for the Team plan."}')
FIXED = ('{"category": "billing", "priority": "high", "language": "en", '
'"summary": "Charged twice for the Team plan.", "needs_human": true}')
print("=== Case 1: broken, then fixed")
llm = ScriptedLLM(replies=[BROKEN, FIXED])
result = generate_structured(llm, messages, Triage, max_attempts=3)
print("value:", result.value)
print("attempts:", result.attempts, "| errors per failed attempt:", result.errors)
print("tokens spent:", result.usage.input_tokens, "in,", result.usage.output_tokens, "out")
print("the correction message the loop sent on attempt 2:")
print(" " + llm.calls[1]["messages"][-1]["content"].replace("\n", "\n "))
print("\n=== Case 2: never fixed, the ceiling stops it")
stubborn = ScriptedLLM(responder=lambda msgs, kw: BROKEN)
result = generate_structured(stubborn, messages, Triage, max_attempts=3)
print("value:", result.value, "| attempts:", result.attempts, "| calls made:", len(stubborn.calls))
print("input tokens per call (estimate):", [count_messages(c["messages"]) for c in stubborn.calls])
print("tokens spent:", result.usage.input_tokens, "in,", result.usage.output_tokens, "out")
print("last parsed object kept for salvage:", result.last_data)
Code explained
- In simple words: the first reply is broken in three ways, the loop explains the problems, and the second reply is fixed; then a stubborn "model" shows the ceiling at work.
- What happens: in case 1,
ScriptedLLM(replies=[BROKEN, FIXED])returns a reply with a preamble,priority: "critical",language: "English", and noneeds_human, then a correct object.generate_structuredparses the first (the preamble is removed), validation fails, and the loop appends the error message you see printed. In case 2, aresponderreturns the broken reply forever.count_messagesestimates each call's input tokens from the recorded messages. - Comes out:
=== Case 1: broken, then fixed
value: category='billing' priority='high' language='en' summary='Charged twice for the Team plan.' needs_human=True
attempts: 2 | errors per failed attempt: [["priority: Input should be 'low', 'normal', 'high' or 'urgent'", "language: String should match pattern '^[a-z]{2}$'", 'needs_human: Field required']]
tokens spent: 196 in, 72 out
the correction message the loop sent on attempt 2:
Your reply did not match the schema:
- priority: Input should be 'low', 'normal', 'high' or 'urgent'
- language: String should match pattern '^[a-z]{2}$'
- needs_human: Field required
Reply again with only the corrected JSON object.
=== Case 2: never fixed, the ceiling stops it
value: None | attempts: 3 | calls made: 3
input tokens per call (estimate): [47, 149, 251]
tokens spent: 447 in, 105 out
last parsed object kept for salvage: {'category': 'billing', 'priority': 'critical', 'language': 'English', 'summary': 'Charged twice for the Team plan.'}
Case 1 took two calls. Case 2 stopped at exactly three calls, and each call was bigger than the last (47, then 149, then 251 estimated input tokens) because each retry carries the whole failed exchange. That growth is why the ceiling is small: with a real model, the third attempt rarely succeeds where two failed, while it always costs the most. A ceiling of 2 or 3 is typical; measure your own model's success rate per attempt before raising it. The loop keeps last_data so the caller can still salvage something.
| Situation | Use this | Why |
|---|---|---|
| Validation errors on a field or two | Repair loop with the specific error lines, ceiling 2 to 3 | Targeted feedback is cheap and usually enough |
| No JSON at all (refusal, prose) twice in a row | Stop and escalate to a human | Retrying the same prompt rarely changes a refusal |
Output truncated (closed_truncated in repairs) | Raise max_tokens or shrink the schema, then retry once | The cause is the limit, not the model's understanding |
| Strict native mode on a supported model | Ceiling of 1 or 2, still validate | Shape errors should not happen; semantic errors still can |
Partial results and salvaging what is valid
When the ceiling is hit, or the output was truncated, you often still have useful fields. Salvaging means keeping the fields that pass validation and making a deliberate, safe decision about the rest, instead of either crashing or trusting everything.
"""Salvage: keep the valid fields of an invalid object, then apply a safe fallback policy."""
from __future__ import annotations
from examples.m06_hostile import HOSTILE
from examples.m06_structured import ParseError, parse_json, salvage
from supportdesk.schemas import Triage
SAFE_DEFAULTS = {"priority": "high", "needs_human": True} # unknown means: a person looks, soon
def triage_or_fallback(data: dict) -> tuple[str, dict, dict]:
"""Return (status, triage_fields, problems). status is valid, salvaged, or unusable."""
valid, problems = salvage(data, Triage)
if not problems:
return "valid", valid, {}
if "category" not in valid: # without a category we cannot route at all
return "unusable", valid, problems
filled = {**valid, **{k: v for k, v in SAFE_DEFAULTS.items() if k not in valid}}
filled["needs_human"] = True # anything salvaged gets human review
return "salvaged", filled, problems
counts = {"valid": 0, "salvaged": 0, "unusable": 0, "unparseable": 0}
for name, text in HOSTILE.items():
try:
parsed = parse_json(text)
except ParseError:
counts["unparseable"] += 1
continue
status, fields, problems = triage_or_fallback(parsed.value)
if "closed_truncated" in parsed.repairs and status == "valid":
status = "salvaged" # a closed truncation is never fully trusted
counts[status] += 1
if status != "valid":
route = {k: fields.get(k) for k in ("category", "priority", "needs_human")}
print(f"{name:<22} {status:<9} route with {route}")
print(f"{'':<32} problems={problems}")
print("\n", counts, f"(n={len(HOSTILE)})")
Code explained
- In simple words: for every hostile output that parses but does not fully validate, keep the good parts and fill the gaps in the safest direction.
- What happens:
triage_or_fallbackcallssalvage. If nothing is wrong, the object is valid. Ifcategoryitself is bad, the ticket cannot be routed, so it is unusable. Otherwise missing or invalid fields getSAFE_DEFAULTS(priorityhigh, a human looks), and anything salvaged is forced toneeds_human = True. An object whose repairs includeclosed_truncatedis never counted as fully valid, even if every field happened to validate. - Comes out:
truncated_in_value salvaged route with {'category': 'billing', 'priority': 'high', 'needs_human': True}
problems={'needs_human': 'Field required'}
truncated_after_comma salvaged route with {'category': 'billing', 'priority': 'high', 'needs_human': True}
problems={'needs_human': 'Field required'}
truncated_in_key salvaged route with {'category': 'account_access', 'priority': 'high', 'needs_human': True}
problems={'needs_human': 'Field required'}
fence_truncated salvaged route with {'category': 'cancellation', 'priority': 'normal', 'needs_human': True}
problems={'needs_human': 'Field required'}
wrong_enum unusable route with {'category': None, 'priority': None, 'needs_human': True}
problems={'category': "Input should be 'billing', 'cancellation', 'account_access', 'bug', 'how_to' or 'feature_request'", 'priority': "Input should be 'low', 'normal', 'high' or 'urgent'"}
wrong_types salvaged route with {'category': 'billing', 'priority': 'high', 'needs_human': True}
problems={'language': "String should match pattern '^[a-z]{2}$'"}
extra_field salvaged route with {'category': 'billing', 'priority': 'high', 'needs_human': True}
problems={'confidence': 'Extra inputs are not permitted'}
{'valid': 14, 'salvaged': 6, 'unusable': 1, 'unparseable': 3} (n=24)
Of 24 outputs, 14 are fully valid, 6 are salvaged into a safe route, 1 is unusable (both enum fields wrong), and 3 never parsed. The wrong_types case also shows lax mode at work: "needs_human": "yes" did not appear as a problem because pydantic coerced it to True. The policy is the important design choice: salvage fails toward more human attention, never less. A salvaged ticket with a wrong priority costs Maya a minute; a salvaged ticket silently marked needs_human: false could leave a refund request unanswered.
Schema size, nesting depth, and their effect on accuracy
A schema is prompt text. In native mode the provider renders it into the model's context; in prompted mode you paste it yourself. Either way it costs tokens on every call, and bigger or deeper schemas give the model more to get wrong. The cost side is easy to measure here.
"""What a schema costs in tokens: formatting, descriptions, width, and nesting depth."""
from __future__ import annotations
import json
from typing import Any
from pydantic import BaseModel, ConfigDict, create_model
from examples.m06_structured import strict_schema
from supportdesk.schemas import Triage
from supportdesk.tokens import count_tokens
def tokens(schema: dict[str, Any], indent: int | None = None) -> int:
text = json.dumps(schema, indent=indent, separators=None if indent else (",", ":"))
return count_tokens(text, "o200k_base")
def without(schema: Any, keys: set[str]) -> Any:
if isinstance(schema, dict):
return {k: without(v, keys) for k, v in schema.items() if k not in keys}
if isinstance(schema, list):
return [without(v, keys) for v in schema]
return schema
base = Triage.model_json_schema()
print("Triage schema, o200k_base tokens")
print(f" pydantic output, indent=2 {tokens(base, 2):>5}")
print(f" compact separators {tokens(base):>5}")
print(f" strict (titles removed) {tokens(strict_schema(Triage)):>5}")
print(f" strict, descriptions removed {tokens(without(strict_schema(Triage), {'description'})):>5}")
# Width: the same kind of field, repeated.
print("\nWidth: flat objects with N string fields (each with a 10-word description)")
for n in (5, 10, 20, 40):
fields = {f"field_{i}": (str, ...) for i in range(n)}
model = create_model(f"Wide{n}", __config__=ConfigDict(extra="forbid"), **fields)
schema = strict_schema(model)
for prop in schema["properties"].values():
prop["description"] = "One short sentence the support agent will read before replying."
print(f" {n:>2} fields: {tokens(schema):>5} tokens")
# Depth: the same five Triage leaves, wrapped in 0 to 3 extra levels of objects.
def nest(model: type[BaseModel], depth: int) -> type[BaseModel]:
for level in range(depth):
model = create_model(f"Level{level}", __config__=ConfigDict(extra="forbid"), inner=(model, ...))
return model
print("\nDepth: Triage's five fields wrapped in extra object levels")
for depth in range(4):
schema = strict_schema(nest(Triage, depth))
example = json.loads(Triage(category="billing", priority="high", language="en",
summary="Customer was charged twice.", needs_human=True).model_dump_json())
for _ in range(depth):
example = {"inner": example}
print(f" depth {depth + 1}: schema {tokens(schema):>4} tokens (uses $defs: {'$defs' in schema}), "
f"one answer {count_tokens(json.dumps(example), 'o200k_base'):>3} tokens")
Code explained
- In simple words: count what the schema itself costs, and how that grows with more fields and deeper nesting.
- What happens:
tokens()serializes a schema and counts it with theo200k_basetokenizer fromsupportdesk/tokens.py. The first block compares four versions of theTriageschema. The width block builds flat models with 5 to 40 string fields using pydantic'screate_model, each with the same 10-word description. The depth block wrapsTriagein 0 to 3 extra objects and also counts the tokens of one matching answer. - Comes out:
Triage schema, o200k_base tokens
pydantic output, indent=2 381
compact separators 248
strict (titles removed) 222
strict, descriptions removed 113
Width: flat objects with N string fields (each with a 10-word description)
5 fields: 146 tokens
10 fields: 276 tokens
20 fields: 536 tokens
40 fields: 1056 tokens
Depth: Triage's five fields wrapped in extra object levels
depth 1: schema 222 tokens (uses $defs: False), one answer 34 tokens
depth 2: schema 257 tokens (uses $defs: True), one answer 37 tokens
depth 3: schema 290 tokens (uses $defs: True), one answer 41 tokens
depth 4: schema 323 tokens (uses $defs: True), one answer 44 tokens
Formatting alone matters: the indented schema costs 381 tokens, the compact one 248, and removing pydantic's titles brings it to 222. Descriptions are about half of what remains (222 vs 113 without them), and they are usually worth it, because they carry the rubric. Width is linear, about 26 tokens per described field. Each nesting level adds about 33 schema tokens and 3 to 4 answer tokens here; pydantic switches to $defs and $ref as soon as there is a nested model, and those references are exactly what some providers limit. Counts are for the JSON text; each provider renders schemas into its own prompt format, so treat them as estimates and confirm with the usage a real response reports.
The accuracy side cannot be measured without a real model, and published evidence is mixed, so here is what the research says and what it does not:
- Tam et al., "Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models" (2024, arXiv:2408.02442), report "a significant decline in LLMs' reasoning abilities under format restrictions," with stricter constraints degrading reasoning tasks more.
- Geng et al., "JSONSchemaBench" (2025, arXiv:2501.10868), test six constrained-decoding systems, including OpenAI's and Gemini's, on 10,000 real-world schemas and find that no framework handles every schema feature and complexity level; coverage of complex schemas varies widely.
- Gemini's docs warn that very large or deeply nested schemas may be rejected.
None of these gives a number you can apply to Brightlane's triage. So measure it, with a harness that compares the flat Triage against a two-level NestedTriage holding the same five facts.
"""Does a nested schema change accuracy? A harness to measure it on the dev tickets.
LLM_PROVIDER=groq GROQ_API_KEY=... PYTHONPATH=. python examples/m06_schema_eval.py
PYTHONPATH=. python examples/m06_schema_eval.py --dry (keyword baseline: checks the harness only)
"""
from __future__ import annotations
import json
import sys
from pydantic import BaseModel, ConfigDict
from examples.m06_lab import keyword_triage, wilson
from examples.m06_structured import generate_structured, response_format_for
from supportdesk.data import load_tickets
from supportdesk.schemas import Category, Priority, Triage
from supportdesk.stand_in import ScriptedLLM
class Routing(BaseModel):
model_config = ConfigDict(extra="forbid")
category: Category
priority: Priority
needs_human: bool
class Customer(BaseModel):
model_config = ConfigDict(extra="forbid")
language: str
summary: str
class NestedTriage(BaseModel):
"""The same five facts as Triage, grouped two levels deep."""
model_config = ConfigDict(extra="forbid")
routing: Routing
customer: Customer
def category_of(value: BaseModel) -> str:
return value.routing.category if isinstance(value, NestedTriage) else value.category
def run(schema_cls: type[BaseModel], llm, tickets) -> tuple[int, int, int]:
valid = correct = out_tokens = 0
for t in tickets:
messages = [{"role": "system", "content": "Triage this Brightlane support ticket as JSON."},
{"role": "user", "content": t.text}]
res = generate_structured(llm, messages, schema_cls, max_attempts=1,
response_format=response_format_for(schema_cls), max_tokens=2000)
out_tokens += res.usage.output_tokens
if res.value is not None:
valid += 1
correct += category_of(res.value) == t.gold["category"]
return valid, correct, out_tokens
if __name__ == "__main__":
tickets = load_tickets("dev")
if "--dry" in sys.argv:
by_text = {t.text: t for t in tickets}
def dry(messages, kwargs):
d = keyword_triage(by_text[messages[1]["content"]])
if kwargs["response_format"]["json_schema"]["name"] == "NestedTriage":
d = {"routing": {k: d[k] for k in ("category", "priority", "needs_human")},
"customer": {k: d[k] for k in ("language", "summary")}}
return json.dumps(d)
llm = ScriptedLLM(responder=dry)
else:
from supportdesk.llm import chat as llm
n = len(tickets)
for schema_cls in (Triage, NestedTriage):
valid, correct, out_tokens = run(schema_cls, llm, tickets)
lo, hi = wilson(correct, n)
print(f"{schema_cls.__name__:<13} valid {valid}/{n} category accuracy {correct}/{n} = {correct / n:.0%} "
f"(95% CI {lo:.0%} to {hi:.0%}) output tokens {out_tokens}")
Code explained
- In simple words: the same tickets, the same facts, two shapes; count valid answers, correct categories, and output tokens for each.
- What happens:
NestedTriagegroups the fields intoroutingandcustomer.runsends each dev ticket throughgenerate_structuredwith one attempt (so the first answer is what gets scored) and the strictresponse_formatfor that schema.--dryswaps in aScriptedLLMwhose answers come from the lab's keyword baseline, reshaped for the nested schema, so you can check the harness runs end to end without a key. - Comes out: the dry run (keyword baseline, not a model):