CourseLarge Language Models · Module 6: Structured Outputs and Tool Use · part 27 of 80
Part 27 · Module 6: Structured Outputs and Tool Use

Part A: Getting reliable structure

45 min read·22 Sept 2026

Module 6: Structured Outputs and Tool Use

By the end of this module, you'll have:

  • A tolerant JSON parser, measured on 24 hostile model outputs, that raises the parse rate from 21% to 88% and logs every repair it makes.
  • A repair loop with an attempt ceiling and a salvage step, so a malformed answer becomes a valid object, a safe partial result, or a clean hand-off to a human, never a crash.
  • A strict response_format built from the course's pydantic schemas, plus a clear picture of which of Groq, Gemini, and Ollama enforce it.
  • A tool-calling loop for the Brightlane assistant with three tools (search_kb, get_invoice, issue_refund), parallel calls ordered by tool_call_id, tool_choice control, and errors returned to the model instead of raised.
  • A guard layer that treats model output as untrusted input: argument validation, permission scoping, timeouts, a process sandbox, and human approval before any refund. A refund larger than its invoice is rejected before a person is even asked.

Prerequisites: Modules 1 to 5. You need llm.chat and ScriptedLLM (Module 1), token counting with tokens.py and cost math with pricing.py (Module 2), and the prompt-as-code habits of Modules 4 and 5. kb_search.py is formally introduced in Module 7; here we only call KBSearch().search() as a tool.

Where we are: In Module 5 you chained calls and ran self-correction loops, and every step still passed plain text to the next. That works until a program, not a person, has to read the answer. This module makes the assistant's output something code can trust: typed triage objects, and tool calls that are checked before anything happens in Brightlane's billing system.

How this module is organized

PartWhat it covers
SetupThe working copy, and how the examples run
Part A: Getting reliable structureWhy free text breaks code, schema design, native structured output vs prompted JSON, pydantic validation, a hostile-output parser, repair loops, salvage, schema size and depth
Part B: Tool and function callingWhat the model really sees, descriptions and parameters, the execution loop, errors as data, parallel calls, tool_choice, too many tools
Part C: Model output as untrusted inputArgument validation, permission scoping, timeouts, sandboxing, human approval for privileged actions, tests
Module LabTriage all 48 dev tickets through the full pipeline, then route invoice tickets through the tool loop with an approval queue
Project Milestone, Interview Questions, Other Tools, Coming UpWrap-up

Setup

All examples run in order from the repository root, in one shell session. They produce real output without an API key: the parsers, validators, loops, and guards are deterministic code, and every "model" in the offline runs is either ScriptedLLM or a keyword baseline. Neither is a language model. They test plumbing, and every place they appear says so. The scripts that call a hosted model run through llm.chat and need a key; their output is marked as illustrative.

bash
source /home/claude/venv/bin/activate          # or your own venv with requirements.txt installed
cd supportdesk
export PYTHONPATH=.
python -c "import pydantic, openai, httpx; print(pydantic.VERSION, openai.__version__, httpx.__version__)"

Code explained

  • In simple words: switch on the course environment, stand in the repo, and let Python find both supportdesk and examples.
  • What happens: PYTHONPATH=. makes supportdesk.* importable and lets module scripts import each other as examples.m06_... (a namespace package, no __init__.py needed). The last line prints the versions this module was built with. httpx comes with openai; we use it once to capture requests offline.
  • Comes out:

Part A: Getting reliable structure

Why free text breaks downstream systems

Brightlane's help desk routes each ticket to a queue with a response deadline. The router is ordinary Python written by someone who expects clean data. A downstream system is any code that consumes the model's answer: a router, a database write, a billing call. It cannot "read between the lines" the way Maya can.

python
"""Why free text breaks downstream code: a real router fed three kinds of model output."""
from __future__ import annotations

import json
import re
import traceback

# The support desk's router: which queue a ticket goes to and how fast it must be answered.
QUEUES = {"billing": "finance-queue", "cancellation": "retention-queue", "account_access": "access-queue",
          "bug": "engineering-queue", "how_to": "tier1-queue", "feature_request": "product-queue"}
SLA_HOURS = {"urgent": 1, "high": 4, "normal": 24, "low": 72}


def route(triage: dict) -> str:
    """Downstream code written by someone who expects clean data."""
    queue = QUEUES[triage["category"]]
    hours = SLA_HOURS[triage["priority"]]
    owner = "human" if triage["needs_human"] else "assistant"
    return f"{queue} within {hours}h, handled by {owner}"


FREE_TEXT = """Category: Billing
Priority: High (money is wrong)
This customer needs a human because they want a refund."""

LOOKS_LIKE_JSON = """Sure! Here is the triage:
{"category": "billing", "priority": "high", "needs_human": true}"""

VALID = '{"category": "billing", "priority": "high", "needs_human": true}'


def try_route(label: str, text: str) -> None:
    print(f"--- {label}")
    try:
        print("OK:", route(json.loads(text)))
    except Exception as exc:  # show the failure the way a log would
        print("CRASH:", "".join(traceback.format_exception_only(exc)).strip())


try_route("valid JSON", VALID)
try_route("JSON with a friendly preamble", LOOKS_LIKE_JSON)
try_route("free text", FREE_TEXT)

# The usual "quick fix": scrape the free text with a regex. It runs, then fails later.
print("--- free text scraped with a regex")
fields = {k.lower(): v.strip() for k, v in re.findall(r"^(\w+):\s*(.+)$", FREE_TEXT, re.MULTILINE)}
print("scraped:", fields)
try:
    print("OK:", route(fields))
except Exception as exc:
    print("CRASH:", "".join(traceback.format_exception_only(exc)).strip())

Code explained

  • In simple words: we feed one honest router three answers a model might plausibly give, and watch which ones it survives.
  • What happens: route() indexes two dictionaries and reads a boolean. try_route runs json.loads first, as most first versions do. The last block tries the classic quick fix, a regex that scrapes Key: value lines out of free text.
  • Comes out:
text
--- valid JSON
OK: finance-queue within 4h, handled by human
--- JSON with a friendly preamble
CRASH: json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)
--- free text
CRASH: json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)
--- free text scraped with a regex
scraped: {'category': 'Billing', 'priority': 'High (money is wrong)'}
CRASH: KeyError: 'Billing'

Only strict JSON survives. The preamble breaks json.loads at character 0. The regex "fix" is worse: it runs without error, produces 'Billing' with a capital B and 'High (money is wrong)', drops needs_human entirely, and fails one step later with a KeyError. In production that is the dangerous kind of failure, because the error surfaces far from its cause. The rest of Part A removes each failure in turn: say exactly what shape you want, ask the provider to enforce it, parse defensively, validate, repair, and salvage.

Schema design: required, optional, enumerated, bounded fields

A schema is a machine-readable description of the shape an answer must have. The course defines its schemas once, as pydantic models, in supportdesk/schemas.py. Pydantic is a Python library that turns type hints into validators: you declare fields, and it checks data against them and produces a JSON Schema. JSON Schema is the standard format providers accept for structured output and tool parameters. This module introduces the file.

supportdesk/schemas.py

python
"""Typed output schemas for the support assistant, validated with pydantic."""
from __future__ import annotations

from typing import Literal

from pydantic import BaseModel, ConfigDict, Field

Category = Literal["billing", "cancellation", "account_access", "bug", "how_to", "feature_request"]
Priority = Literal["low", "normal", "high", "urgent"]


class Triage(BaseModel):
    """What the assistant decides about one incoming ticket."""

    model_config = ConfigDict(extra="forbid")

    category: Category = Field(description="The single best category for the ticket.")
    priority: Priority = Field(description="urgent: many users blocked or security risk now; high: one user blocked or money wrong; normal: needs an answer; low: question or idea.")
    language: str = Field(pattern=r"^[a-z]{2}$", description="ISO 639-1 code of the customer's language, for example 'en'.")
    summary: str = Field(min_length=5, max_length=200, description="One sentence in English describing the request.")
    needs_human: bool = Field(description="True if an agent must act (refund, account change, legal, security) or the help center cannot answer it.")


class DraftReply(BaseModel):
    """A reply the assistant proposes; a human agent reviews it before sending."""

    model_config = ConfigDict(extra="forbid")

    reply: str = Field(min_length=1, max_length=2000)
    cited_articles: list[str] = Field(default_factory=list, max_length=3, description="Help-center article ids the reply relies on.")
    confidence: Literal["low", "medium", "high"]


def triage_json_schema() -> dict:
    """The JSON Schema for Triage, ready for a provider's structured-output mode."""
    return Triage.model_json_schema()

Code explained

  • In simple words: two forms the assistant fills in, with the rules for each box written right on the form.
  • What happens:
    • Category and Priority are Literal types: enumerations, fields that may only take one of a fixed list of strings. They reuse the labels from supportdesk/data.py, so gold labels and model output use the same vocabulary.
    • Triage is what the assistant decides about one ticket. model_config = ConfigDict(extra="forbid") rejects unknown keys, which turns into "additionalProperties": false in the schema. Every field is required (no default). language is bounded by a regular expression (two lowercase letters), summary by a length range of 5 to 200. The description strings are not comments: they are sent to the model and are part of the prompt. The priority description is a small rubric, because "high" means nothing unless you define it.
    • DraftReply is the reply a human reviews. cited_articles is optional: it has a default (an empty list), so the model may leave it out. It is also bounded (max_length=3 on a list means at most three items).
    • triage_json_schema() returns the JSON Schema for Triage, ready to hand to a provider.
  • Comes out: nothing on import. The next block prints the real schema.
bash
python -c "import json; from supportdesk.schemas import triage_json_schema; print(json.dumps(triage_json_schema(), indent=2))"

Code explained

  • In simple words: ask pydantic what it will actually send.
  • What happens: Triage.model_json_schema() walks the fields and emits JSON Schema (draft 2020-12 style). Printing it is the only reliable way to know what the model will see; do not guess from the Python class.
  • Comes out:
json
{
  "additionalProperties": false,
  "description": "What the assistant decides about one incoming ticket.",
  "properties": {
    "category": {
      "description": "The single best category for the ticket.",
      "enum": [
        "billing",
        "cancellation",
        "account_access",
        "bug",
        "how_to",
        "feature_request"
      ],
      "title": "Category",
      "type": "string"
    },
    "priority": {
      "description": "urgent: many users blocked or security risk now; high: one user blocked or money wrong; normal: needs an answer; low: question or idea.",
      "enum": [
        "low",
        "normal",
        "high",
        "urgent"
      ],
      "title": "Priority",
      "type": "string"
    },
    "language": {
      "description": "ISO 639-1 code of the customer's language, for example 'en'.",
      "pattern": "^[a-z]{2}$",
      "title": "Language",
      "type": "string"
    },
    "summary": {
      "description": "One sentence in English describing the request.",
      "maxLength": 200,
      "minLength": 5,
      "title": "Summary",
      "type": "string"
    },
    "needs_human": {
      "description": "True if an agent must act (refund, account change, legal, security) or the help center cannot answer it.",
      "title": "Needs Human",
      "type": "boolean"
    }
  },
  "required": [
    "category",
    "priority",
    "language",
    "summary",
    "needs_human"
  ],
  "title": "Triage",
  "type": "object"
}

Read it the way the model will: each property has a type, enums list every allowed value, bounds appear as pattern, minLength, maxLength, and required lists all five fields. Pydantic also adds title keys, which cost tokens and carry no information here; we strip them later.

A few design rules for the fields you add:

SituationUse thisWhy
A value from a known, small set (category, priority)An enum (Literal[...])The model cannot invent "Billing" or "critical"; with strict mode the provider enforces it, and validation catches it everywhere else
A value your code must always haveRequired, no defaultA missing field should fail loudly, not silently become a default
A value that is often absent (cited articles)Optional with a safe defaultForces nothing the model cannot know; in strict mode make it nullable (below)
Free text shown to a person (summary, reply)str with min_length and max_lengthStops empty strings and runaway essays
A number that moves money or quotasNumeric bounds (gt, le) plus a business check in codeA schema bound is coarse; the real limit (the invoice amount) lives in your database
A classification the model should justifyPut a short reasoning field before the label, or keep reasoning out of the schemaResearch (below) finds strict formats can hurt reasoning; measure on your task

Now look at how optional fields interact with provider strict modes.

python
"""Look at the schemas the assistant uses, and turn one into a strict response_format."""
from __future__ import annotations

import json

from examples.m06_structured import response_format_for, validate_output
from supportdesk.schemas import DraftReply, Triage, triage_json_schema

schema = triage_json_schema()
print("Triage required:", schema["required"])
print("Triage enums:", {k: v["enum"] for k, v in schema["properties"].items() if "enum" in v})
print("Triage bounds:", {k: {b: v[b] for b in ("pattern", "minLength", "maxLength") if b in v}
                         for k, v in schema["properties"].items() if any(b in v for b in ("pattern", "minLength"))})

print("\nDraftReply required (pydantic):", DraftReply.model_json_schema()["required"])
fmt = response_format_for(DraftReply)
print("DraftReply response_format for strict mode:")
print(json.dumps(fmt, indent=2))

# A strict-mode reply must include every field, so optional ones arrive as null.
from_model = {"reply": "You can export a board from Board menu > Export.", "cited_articles": None, "confidence": "high"}
print("\nvalidated:", validate_output(from_model, DraftReply))

Code explained

  • In simple words: print the schema's rules, then reshape DraftReply for a provider that insists every field be present.
  • What happens: the first prints pull the required list, enums, and bounds out of the Triage schema. response_format_for(DraftReply) (from examples/m06_structured.py, shown in full in the next section) builds a response_format in strict mode, a provider setting that constrains generation to the schema. Groq documents two strict-mode rules: every field must be listed in required, and every object must set additionalProperties: false; optional values are expressed as a union with null. The helper does that reshaping. The last lines show the reverse trip: a strict-mode answer with "cited_articles": null validates back into a DraftReply whose cited_articles is [].
  • Comes out:

Native structured-output modes vs prompted JSON

There are two ways to get JSON from a model:

  • Prompted JSON: you describe the format in the prompt ("reply with only a JSON object matching this schema") and hope. Every provider supports it. Nothing enforces it.
  • Native structured output: you send the schema in the request's response_format, and the provider uses it during generation. With constrained decoding, the provider masks out any next token that would break the schema, so the output is valid JSON of the right shape by construction. With a best-effort mode, the schema guides the model but invalid output is still possible.

What each course provider documents, checked on 21 September 2026 (these pages change often, so check them again):

Provider (OpenAI-compatible endpoint)response_format with json_schemaStrict, constrained outputNotes from the docs
GroqYesYes on GPT-OSS 20B and 120B (the course default) and one Qwen model; other models best-effort"Streaming and tool use are not currently supported with Structured Outputs." Best-effort mode can return HTTP 400 "Generated JSON does not match the expected schema." (Groq structured outputs)
GeminiYes, through the OpenAI library; Google's example passes a pydantic class to client.beta.chat.completions.parseThe compatibility page does not document strict. The native API's structured output lists supported keywords (enum, minimum, maximum, format, anyOf, $ref, additionalProperties); minLength and pattern are not in that listThe compatibility layer is labeled beta. The native docs warn "Very large or deeply nested schemas may be rejected" and "always validate values in your application." (Gemini OpenAI compatibility, Gemini structured output)
OllamaYes: "Structured outputs work through the OpenAI-compatible API via response_format"strict is not listed among supported fieldsAlso supports format with a JSON Schema on its native API. Note that tool_choice is listed as not supported (Ollama OpenAI compatibility, Ollama structured outputs)

Two consequences follow. First, even the best case guarantees syntax and shape, not truth: a strictly valid triage can still say billing for a login problem. Second, for two of three providers, keywords like pattern or maxLength may be ignored during generation. Either way, you validate on your side every time.

Here is the helper file that the rest of Part A uses. It is long, so it sits in a collapsible block; each function is explained after it.

examples/m06_structured.py

python
"""Module 6 helpers: turn model text into validated objects, or fail loudly.

Four layers, each usable on its own:
  response_format_for  build a json_schema response_format for a provider
  parse_json           a tolerant parser for hostile model output
  generate_structured  call, parse, validate, and repair with an attempt ceiling
  salvage              keep the fields that are valid when the whole object is not
"""
from __future__ import annotations

import copy
import json
import re
from dataclasses import dataclass, field
from typing import Any, Callable

from pydantic import BaseModel, ValidationError

from supportdesk.llm import ChatResult, Usage

# ---------------------------------------------------------------------------
# 1. Native structured output: a json_schema response_format
# ---------------------------------------------------------------------------


def strict_schema(model_cls: type[BaseModel]) -> dict[str, Any]:
    """The model's JSON Schema, reshaped for strict mode.

    Strict mode (Groq documents it for gpt-oss) needs every property listed in
    `required` and `additionalProperties: false` on every object. A field that
    was optional becomes nullable instead: the model must write it, but may
    write null. `validate_output` turns those nulls back into defaults.
    """
    schema = copy.deepcopy(model_cls.model_json_schema())

    def fix(node: Any) -> None:
        if isinstance(node, dict):
            if isinstance(node.get("title"), str):      # a schema title, not a property named "title"
                node.pop("title")
            if node.get("type") == "object" and "properties" in node:
                was_required = set(node.get("required", []))
                for name, prop in node["properties"].items():
                    if name not in was_required:
                        prop.pop("default", None)
                        wrapped: dict[str, Any] = {"anyOf": [prop, {"type": "null"}]}
                        if "description" in prop:            # keep the description where the model reads it
                            wrapped = {"description": prop.pop("description"), **wrapped}
                        node["properties"][name] = wrapped
                node["required"] = list(node["properties"])
                node["additionalProperties"] = False
            for value in node.values():
                fix(value)
        elif isinstance(node, list):
            for value in node:
                fix(value)

    fix(schema)
    return schema


def response_format_for(model_cls: type[BaseModel], strict: bool = True) -> dict[str, Any]:
    """The `response_format` value for llm.chat(..., response_format=...)."""
    return {
        "type": "json_schema",
        "json_schema": {"name": model_cls.__name__, "schema": strict_schema(model_cls), "strict": strict},
    }


# ---------------------------------------------------------------------------
# 2. A tolerant parser, built in stages so each stage can be measured
# ---------------------------------------------------------------------------

FENCE = re.compile(r"```(?:json|JSON)?\s*\n?(.*?)(?:```|$)", re.DOTALL)
SMART_QUOTES = str.maketrans({"\u201c": '"', "\u201d": '"', "\u2018": "'", "\u2019": "'"})


class ParseError(ValueError):
    """Raised when no JSON object can be recovered from the text."""


@dataclass
class Parsed:
    value: dict[str, Any]
    repairs: list[str] = field(default_factory=list)   # what the parser had to fix, for logging


def strip_fences(text: str) -> tuple[str, bool]:
    """Return the inside of the first ```json fence, if there is one."""
    match = FENCE.search(text)
    return (match.group(1), True) if match else (text, False)


def extract_object(text: str) -> tuple[str, bool]:
    """Cut from the first '{' to its matching '}', skipping braces inside strings.

    If the text ends before the object closes, return everything from '{' on:
    the truncation stage may still be able to close it.
    """
    start = text.find("{")
    if start == -1:
        raise ParseError("no '{' in output")
    depth, in_str, quote, escaped = 0, False, "", False
    for i in range(start, len(text)):
        ch = text[i]
        if in_str:
            if escaped:
                escaped = False
            elif ch == "\\":
                escaped = True
            elif ch == quote:
                in_str = False
        elif ch in "\"'":
            in_str, quote = True, ch
        elif ch == "{":
            depth += 1
        elif ch == "}":
            depth -= 1
            if depth == 0:
                return text[start:i + 1], (start > 0 or i + 1 < len(text.rstrip()))
    return text[start:], start > 0


def normalize(text: str) -> tuple[str, list[str]]:
    """Fix the JSON-adjacent syntax models write: single quotes, trailing commas,
    Python literals, and smart quotes. Works token by token, so text inside
    double-quoted strings is never touched."""
    repairs: list[str] = []
    if any(c in text for c in "\u201c\u201d\u2018\u2019"):
        text = text.translate(SMART_QUOTES)
        repairs.append("smart_quotes")
    out: list[str] = []
    i, n = 0, len(text)
    while i < n:
        ch = text[i]
        if ch == '"':                                   # copy a JSON string unchanged
            j = i + 1
            while j < n and text[j] != '"':
                j += 2 if text[j] == "\\" else 1
            out.append(text[i:j + 1])
            i = j + 1
        elif ch == "'":                                 # 'single quoted' -> "double quoted"
            j = i + 1
            while j < n and text[j] != "'":
                j += 2 if text[j] == "\\" else 1
            out.append(json.dumps(text[i + 1:j].replace("\\'", "'")))
            if "single_quotes" not in repairs:
                repairs.append("single_quotes")
            i = j + 1
        elif ch == ",":                                 # drop a comma that closes nothing
            j = i + 1
            while j < n and text[j] in " \t\r\n":
                j += 1
            if j < n and text[j] in "}]":
                if "trailing_comma" not in repairs:
                    repairs.append("trailing_comma")
            else:
                out.append(ch)
            i += 1
        elif text.startswith(("True", "False", "None"), i) and not (i and text[i - 1].isalnum()):
            word = next(w for w in ("True", "False", "None") if text.startswith(w, i))
            out.append({"True": "true", "False": "false", "None": "null"}[word])
            if "python_literals" not in repairs:
                repairs.append("python_literals")
            i += len(word)
        else:
            out.append(ch)
            i += 1
    return "".join(out), repairs


def close_truncated(text: str) -> tuple[str, bool]:
    """Close an object that was cut off mid-stream (for example by max_tokens).

    Closes an open string, drops a dangling key or comma, then appends the
    closing brackets in the right order. The value may be incomplete, so the
    caller must treat a closed object as suspect.
    """
    stack: list[str] = []
    in_str, escaped = False, False
    for ch in text:
        if in_str:
            if escaped:
                escaped = False
            elif ch == "\\":
                escaped = True
            elif ch == '"':
                in_str = False
        elif ch == '"':
            in_str = True
        elif ch in "{[":
            stack.append("}" if ch == "{" else "]")
        elif ch in "}]" and stack:
            stack.pop()
    if not stack and not in_str:
        return text, False
    fixed = text + ('"' if in_str else "")
    fixed = re.sub(r',\s*"[^"]*"\s*:?\s*$', "", fixed)    # a key with no value yet
    fixed = re.sub(r'[,:]\s*$', "", fixed.rstrip())         # a dangling comma or colon
    return fixed + "".join(reversed(stack)), True


PARSER_STAGES = ("plain", "fences", "extract", "normalize", "truncation")


def parse_json(text: str, stages: tuple[str, ...] = PARSER_STAGES) -> Parsed:
    """Recover one JSON object from model output, applying only the listed stages."""
    repairs: list[str] = []
    candidate = text.strip()
    if "fences" in stages:
        candidate, fenced = strip_fences(candidate)
        if fenced:
            repairs.append("fences")
    if "extract" in stages:
        candidate, cut = extract_object(candidate)
        if cut:
            repairs.append("surrounding_text")
    if "normalize" in stages:
        candidate, fixes = normalize(candidate)
        repairs += fixes
    if "truncation" in stages:
        candidate, closed = close_truncated(candidate)
        if closed:
            repairs.append("closed_truncated")
    try:
        value = json.loads(candidate)
    except json.JSONDecodeError as exc:
        raise ParseError(f"{exc.msg} at char {exc.pos}") from exc
    if not isinstance(value, dict):
        raise ParseError(f"expected an object, got {type(value).__name__}")
    return Parsed(value, repairs)


# ---------------------------------------------------------------------------
# 3. Validation, repair loop with a ceiling, and salvage
# ---------------------------------------------------------------------------


def validate_output(data: dict[str, Any], model_cls: type[BaseModel]) -> BaseModel:
    """Validate with pydantic after turning strict-mode nulls back into defaults."""
    cleaned = {
        k: v for k, v in data.items()
        if not (v is None and k in model_cls.model_fields and not model_cls.model_fields[k].is_required())
    }
    return model_cls.model_validate(cleaned)


def describe_errors(exc: ValidationError) -> list[str]:
    """Short, model-readable error lines such as "priority: Input should be ..."."""
    lines = []
    for err in exc.errors():
        where = ".".join(str(p) for p in err["loc"]) or "(root)"
        lines.append(f"{where}: {err['msg']}")
    return lines


@dataclass
class StructuredResult:
    value: BaseModel | None
    attempts: int
    errors: list[list[str]]            # the errors of each failed attempt
    repairs: list[str]                 # parser repairs on the attempt that succeeded
    usage: Usage
    last_data: dict[str, Any] | None   # the last parsed object, for salvage


def generate_structured(
    llm: Callable[..., ChatResult],
    messages: list[dict[str, Any]],
    model_cls: type[BaseModel],
    *,
    max_attempts: int = 3,
    **chat_kwargs: Any,
) -> StructuredResult:
    """Ask, parse, validate. On failure, show the model its errors and ask again,
    at most `max_attempts` calls in total. Never loops forever, never raises
    for bad output: the caller decides what a failure means."""
    history = list(messages)
    usage = Usage()
    all_errors: list[list[str]] = []
    last_data: dict[str, Any] | None = None
    for attempt in range(1, max_attempts + 1):
        result = llm(history, **chat_kwargs)
        usage.input_tokens += result.usage.input_tokens
        usage.output_tokens += result.usage.output_tokens
        try:
            parsed = parse_json(result.text)
            last_data = parsed.value
            value = validate_output(parsed.value, model_cls)
            return StructuredResult(value, attempt, all_errors, parsed.repairs, usage, last_data)
        except ParseError as exc:
            errors = [f"(root): could not parse JSON: {exc}"]
        except ValidationError as exc:
            errors = describe_errors(exc)
        all_errors.append(errors)
        history += [
            {"role": "assistant", "content": result.text},
            {"role": "user", "content": "Your reply did not match the schema:\n- " + "\n- ".join(errors)
             + "\nReply again with only the corrected JSON object."},
        ]
    return StructuredResult(None, max_attempts, all_errors, [], usage, last_data)


def salvage(data: dict[str, Any], model_cls: type[BaseModel]) -> tuple[dict[str, Any], dict[str, str]]:
    """Split a parsed object into fields that pass validation and fields that do not.

    Returns (valid_fields, problems). A field is valid if the full validation
    reported no error located at it. Unknown keys and missing required fields
    are listed in problems.
    """
    try:
        return validate_output(data, model_cls).model_dump(), {}
    except ValidationError as exc:
        problems: dict[str, str] = {}
        for err in exc.errors():
            name = str(err["loc"][0]) if err["loc"] else "(root)"
            problems.setdefault(name, err["msg"])
        valid = {k: v for k, v in data.items() if k in model_cls.model_fields and k not in problems}
        return valid, problems

Code explained

  • In simple words: four tools for one job: ask for the right shape, recover JSON from messy text, check it, and when it is wrong, either ask again or keep what is usable.
  • What happens:
    • strict_schema(model_cls) copies pydantic's schema and walks it recursively. It removes title keys, lists every property in required, sets additionalProperties: false on every object, and wraps each formerly optional property in anyOf: [original, null], keeping its description on the outside where the model reads it. It also descends into $defs, so nested models get the same treatment.
    • response_format_for(model_cls, strict=True) wraps that schema in the {"type": "json_schema", "json_schema": {...}} envelope that llm.chat(..., response_format=...) passes straight through to the provider.
    • ParseError is the one exception the parser raises. Parsed carries the recovered object plus repairs, a list of what was fixed. Logging repairs matters: a sudden rise in closed_truncated tells you max_tokens is too low long before anyone reads a bad reply.
    • strip_fences returns the inside of the first Markdown code fence (with or without json), and tolerates a missing closing fence.
    • extract_object finds the first { and walks forward counting braces, but skips braces inside quoted strings, so a summary containing {{name}} does not end the object early. Text before and after the object (a preamble or a sign-off) is dropped. If the object never closes, it returns the tail for the truncation stage.
    • normalize rewrites the almost-JSON that models write: smart quotes become straight quotes, 'single quoted' strings become "double quoted" via json.dumps (so inner characters are escaped correctly), trailing commas before } or ] are dropped, and Python's True, False, None become true, false, null. It scans character by character and copies double-quoted strings untouched, so an apostrophe inside "Customer's card" is safe.
    • close_truncated handles output cut off mid-stream. It tracks open brackets and strings, closes an open string, drops a dangling key, comma, or colon, and appends the missing closers in reverse order. The result may be missing fields or hold a half-written value, so the repair is logged as closed_truncated and treated as suspect later.
    • parse_json(text, stages) runs the enabled stages in order and then json.loads. Making the stages a parameter is what lets us measure each one.
    • validate_output removes null values for fields that have defaults (the strict-mode round trip) and calls model_validate. describe_errors turns a pydantic ValidationError into short lines like priority: Input should be ..., which are what the repair loop sends back to the model.
    • generate_structured is the repair loop: call, parse, validate; on failure, append the bad answer and a message listing the errors, and try again, at most max_attempts calls. It returns a StructuredResult with the value (or None), attempt count, all errors, total usage, and the last parsed object for salvage. It never raises for bad model output.
    • salvage validates the whole object, reads which fields the errors point at, and returns the fields that passed plus a dictionary of problems.
  • Comes out: nothing on its own; the next examples run it.

First, see what native structured output looks like on the wire. examples/m06_wire.py swaps the HTTP transport under the openai client for a fake one that records the request and returns a canned reply, so the request is real (built by llm.chat) and only the reply is invented.

examples/m06_wire.py

python
"""See the exact JSON that llm.chat sends, without an API key or network.

`capture_wire` swaps the HTTP transport under the openai client for a fake one
that records each request body and answers with a canned response. The request
is real (built by llm.chat and the openai library); only the reply is canned.
"""
from __future__ import annotations

import json
from collections.abc import Iterator
from contextlib import contextmanager
from typing import Any

import httpx
from openai import OpenAI

import supportdesk.llm as llm


@contextmanager
def capture_wire(canned_message: dict[str, Any], finish_reason: str = "stop") -> Iterator[list[dict[str, Any]]]:
    """Record request bodies sent by llm.chat; reply with `canned_message` every time."""
    bodies: list[dict[str, Any]] = []

    def handler(request: httpx.Request) -> httpx.Response:
        bodies.append(json.loads(request.content))
        return httpx.Response(200, json={
            "id": "chatcmpl-canned", "object": "chat.completion", "created": 0, "model": bodies[-1]["model"],
            "choices": [{"index": 0, "message": {"role": "assistant", **canned_message}, "finish_reason": finish_reason}],
            "usage": {"prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0},
        })

    def fake_client(provider: str) -> OpenAI:
        return OpenAI(api_key="not-a-real-key", base_url=llm.PROVIDERS[provider]["base_url"],
                      http_client=httpx.Client(transport=httpx.MockTransport(handler)))

    original = llm.make_client
    llm.make_client = fake_client
    try:
        yield bodies
    finally:
        llm.make_client = original

Code explained

  • In simple words: a tap on the wire between llm.chat and the provider, so we can read exactly what leaves your machine.
  • What happens: capture_wire is a context manager. Inside it, supportdesk.llm.make_client is replaced by fake_client, which builds a real OpenAI client whose http_client uses httpx.MockTransport(handler). The handler decodes and stores each request body, then returns a minimal Chat Completions response containing canned_message. On exit the real make_client is restored. No key or network is needed, and the canonical llm.py is not edited.
  • Comes out: nothing by itself; it yields the list of captured request bodies.
python
"""The request body llm.chat sends for native structured output (captured offline)."""
from __future__ import annotations

import json

from examples.m06_native import native_messages
from examples.m06_structured import response_format_for
from examples.m06_wire import capture_wire
from supportdesk.llm import chat
from supportdesk.schemas import Triage

canned = {"content": '{"category":"billing","priority":"high","language":"en",'
                     '"summary":"Customer was charged twice for the Team plan.","needs_human":true}'}
with capture_wire(canned) as bodies:
    result = chat(native_messages("Subject: Charged twice\n\nMy card was charged 288 USD twice."),
                  provider="groq", response_format=response_format_for(Triage), max_tokens=2000)

body = bodies[0]
print(json.dumps({k: v for k, v in body.items() if k != "response_format"}, indent=2))
print("response_format.type:", body["response_format"]["type"])
print("response_format.json_schema keys:", list(body["response_format"]["json_schema"]))
print("\nChatResult.text:", result.text)
print("finish_reason:", result.finish_reason, "| provider:", result.provider, "| model:", result.model)

Code explained

  • In simple words: send one native structured-output request through llm.chat and read the body that would go to Groq.
  • What happens: native_messages (from examples/m06_native.py, next block) builds a system and user message. chat(..., response_format=response_format_for(Triage), max_tokens=2000) runs the real request code. The script prints the body without the long schema, then the response_format envelope's keys, then the ChatResult that llm.chat built from the canned reply.
  • Comes out:
text
{
  "messages": [
    {
      "role": "system",
      "content": "You triage Brightlane support tickets. Brightlane is a project-management SaaS."
    },
    {
      "role": "user",
      "content": "Subject: Charged twice\n\nMy card was charged 288 USD twice."
    }
  ],
  "model": "openai/gpt-oss-120b",
  "max_tokens": 2000,
  "temperature": 0.0
}
response_format.type: json_schema
response_format.json_schema keys: ['name', 'schema', 'strict']

ChatResult.text: {"category":"billing","priority":"high","language":"en","summary":"Customer was charged twice for the Team plan.","needs_human":true}
finish_reason: stop | provider: groq | model: openai/gpt-oss-120b

llm.chat only sends parameters you set (temperature defaults to 0.0 in the helper), and response_format goes through unchanged. The text is the canned reply, not model output. max_tokens=2000 is deliberately generous: gpt-oss is a reasoning model, and on providers that count reasoning tokens against the output limit, a tight limit is the most common cause of truncated JSON.

The live comparison script sends the same tickets both ways.

python
"""Native structured output vs prompted JSON for ticket triage, through llm.chat.

Needs a provider key (or a local Ollama). Usage:
  LLM_PROVIDER=groq GROQ_API_KEY=... python examples/m06_native.py
"""
from __future__ import annotations

import json
import os
import sys

from pydantic import ValidationError

from examples.m06_structured import ParseError, parse_json, response_format_for, validate_output
from supportdesk.data import load_tickets
from supportdesk.llm import PROVIDERS, chat, resolve
from supportdesk.schemas import Triage, triage_json_schema

SYSTEM = "You triage Brightlane support tickets. Brightlane is a project-management SaaS."


def native_messages(ticket_text: str) -> list[dict]:
    return [{"role": "system", "content": SYSTEM}, {"role": "user", "content": ticket_text}]


def prompted_messages(ticket_text: str) -> list[dict]:
    schema = json.dumps(triage_json_schema(), separators=(",", ":"))
    return [{"role": "system", "content": f"{SYSTEM}\nReply with only a JSON object that matches this JSON Schema, "
             f"with no other text:\n{schema}"}, {"role": "user", "content": ticket_text}]


def triage_once(ticket_text: str, mode: str) -> tuple[str, str]:
    """Return (outcome, detail) for one ticket in 'native' or 'prompted' mode."""
    if mode == "native":
        result = chat(native_messages(ticket_text), response_format=response_format_for(Triage), max_tokens=2000)
    else:
        result = chat(prompted_messages(ticket_text), max_tokens=2000)
    try:
        parsed = parse_json(result.text)
        triage = validate_output(parsed.value, Triage)
        return "valid", f"{triage.category}/{triage.priority} repairs={parsed.repairs} out_tokens={result.usage.output_tokens}"
    except (ParseError, ValidationError) as exc:
        return "invalid", str(exc).splitlines()[0]


if __name__ == "__main__":
    provider, model = resolve()
    key_env = PROVIDERS[provider]["key_env"]
    if key_env and not os.environ.get(key_env):
        sys.exit(f"Set {key_env} (or LLM_PROVIDER=ollama) to run this example against {provider}/{model}.")
    print(f"provider={provider} model={model}")
    for ticket in load_tickets("dev")[:5]:
        for mode in ("native", "prompted"):
            outcome, detail = triage_once(ticket.text, mode)
            print(f"{ticket.id} {mode:<8} {outcome:<7} {detail}  gold={ticket.gold['category']}/{ticket.gold['priority']}")

Code explained

  • In simple words: triage five real dev tickets twice, once with the schema enforced by the provider and once with the schema only described in the prompt, and check both with the same parser and validator.
  • What happens: native_messages has no format instructions at all; the schema travels in response_format. prompted_messages pastes the compact schema into the system prompt instead. triage_once calls chat, then parse_json and validate_output, and reports validity, the parser repairs needed, and output tokens. Without a key the script stops with a clear message.
  • Comes out: without a key (what this build can run):
text
Set GROQ_API_KEY (or LLM_PROVIDER=ollama) to run this example against groq/openai/gpt-oss-120b.

With a key, each line shows one ticket and mode. Illustrative sample run (not captured in this build; produced for teaching). Your output will differ.

Copy

text
provider=groq model=openai/gpt-oss-120b
T-1001 native   valid   billing/high repairs=[] out_tokens=...  gold=billing/high
T-1001 prompted valid   billing/high repairs=[] out_tokens=...  gold=billing/high
T-1002 native   valid   billing/normal repairs=[] out_tokens=...  gold=billing/normal

What to look for in your run: the repairs column for the prompted mode (fences and preambles show up there first), and any invalid lines. Five tickets is a smoke test, not a comparison; the lab's --live flag runs all 48.

SituationUse thisWhy
Provider documents strict mode for your model (Groq gpt-oss)Native json_schema with strict: true, plus local validationShape is guaranteed; validation still catches semantic bounds and provider drift
Provider supports json_schema but not documented strict (Gemini compatibility layer, Ollama)Native json_schema, tolerant parser, validation, repair loopShape is likely but not guaranteed; your code must not assume it
You also need tools in the same call on GroqTools, and put the structured answer in a tool's arguments or a separate callGroq documents that structured outputs and tool use cannot be combined
Provider or model has no structured outputPrompted JSON, tolerant parser, repair loopWorks everywhere; costs more tokens and retries
Reasoning-heavy task where format seems to hurt qualityLet the model reason in free text, then extract to JSON in a second, cheap callSeparates thinking from formatting; measure both on your eval set

Validation with typed models

Parsing answers "is this JSON?". Validation answers "is this JSON the object my code expects?". Pydantic does it in one call, Triage.model_validate(data), which either returns a typed object or raises ValidationError with one entry per problem. You saw those errors in the repair message above and will see them again below.

One behavior to know: by default pydantic runs in lax mode, which coerces obvious conversions. In the hostile corpus below, "needs_human": "yes" validates as True. That is convenient for booleans, but for a field like an amount you may prefer an error. Triage.model_validate(data, strict=True) turns coercion off for one call; ConfigDict(strict=True) turns it off for a model. The course's schemas keep lax mode and rely on enums and bounds for the fields that matter.

Parsing hostile output: fences, preambles, trailing commas, truncation

Even with native modes you will parse text from models that do not support them, from best-effort modes, from logs, and from outputs cut off by max_tokens. To build the parser honestly, we need a test set. examples/m06_hostile.py holds 24 outputs for a Triage object, each imitating a failure seen in practice, and a scoreboard that turns on parser stages one at a time.

examples/m06_hostile.py

python
"""A corpus of 24 hostile model outputs for a Triage object, and the parser scoreboard.

Each case imitates a failure seen in real model output. Run it to see how many
cases each parser stage recovers, and how many then pass Triage validation.
"""
from __future__ import annotations

from pydantic import ValidationError

from examples.m06_structured import PARSER_STAGES, ParseError, parse_json, validate_output
from supportdesk.schemas import Triage

GOOD = '{"category": "billing", "priority": "high", "language": "en", "summary": "Customer was charged twice for the Team plan.", "needs_human": true}'

HOSTILE: dict[str, str] = {
    "clean": GOOD,
    "clean_pretty": '{\n  "category": "how_to",\n  "priority": "normal",\n  "language": "de",\n  "summary": "Asks how to export a board to CSV.",\n  "needs_human": false\n}',
    "fence_json": "```json\n" + GOOD + "\n```",
    "fence_bare": "```\n" + GOOD + "\n```",
    "preamble": "Sure! Here is the triage for this ticket:\n\n" + GOOD,
    "postamble": GOOD + "\n\nLet me know if you'd like me to draft a reply as well.",
    "preamble_fence_post": "Here's the JSON you asked for:\n```json\n" + GOOD + "\n```\nI marked it high because money is wrong.",
    "trailing_comma": '{"category": "billing", "priority": "high", "language": "en", "summary": "Customer was charged twice.", "needs_human": true,}',
    "trailing_comma_pretty": '{\n  "category": "bug",\n  "priority": "urgent",\n  "language": "en",\n  "summary": "Boards fail to load for the whole team.",\n  "needs_human": true,\n}',
    "single_quotes": "{'category': 'billing', 'priority': 'high', 'language': 'en', 'summary': 'Customer was charged twice.', 'needs_human': true}",
    "python_dict": "{'category': 'bug', 'priority': 'high', 'language': 'es', 'summary': 'Mobile app crashes on login.', 'needs_human': False}",
    "smart_quotes": "{\u201ccategory\u201d: \u201cbilling\u201d, \u201cpriority\u201d: \u201chigh\u201d, \u201clanguage\u201d: \u201cen\u201d, \u201csummary\u201d: \u201cCustomer was charged twice.\u201d, \u201cneeds_human\u201d: true}",
    "braces_in_string": 'Result: {"category": "how_to", "priority": "low", "language": "en", "summary": "Asks what {{name}} means in templates.", "needs_human": false} Done.',
    "truncated_in_value": '{"category": "billing", "priority": "high", "language": "en", "summary": "Customer was charged tw',
    "truncated_after_comma": '{"category": "billing", "priority": "high", "language": "en", "summary": "Customer was charged twice.",',
    "truncated_in_key": '{"category": "account_access", "priority": "high", "language": "ja", "summary": "Cannot log in after SSO change.", "needs_hu',
    "fence_truncated": "```json\n{\"category\": \"cancellation\", \"priority\": \"normal\", \"language\": \"en\", \"summary\": \"Wants to cancel the annual plan",
    "wrong_enum": '{"category": "Billing", "priority": "critical", "language": "en", "summary": "Customer was charged twice.", "needs_human": true}',
    "wrong_types": '{"category": "billing", "priority": "high", "language": "English", "summary": "Charged twice.", "needs_human": "yes"}',
    "extra_field": '{"category": "billing", "priority": "high", "language": "en", "summary": "Customer was charged twice.", "needs_human": true, "confidence": 0.9}',
    "array_wrapped": '[' + GOOD + ']',
    "prose_only": "This looks like a billing problem with high priority; a human should refund the duplicate charge.",
    "refusal": "I'm sorry, but I can't help with that request.",
    "apostrophe_in_single": "{'category': 'billing', 'priority': 'high', 'language': 'en', 'summary': 'Customer's card was charged twice.', 'needs_human': true}",
}

LADDER = [PARSER_STAGES[:i] for i in range(1, len(PARSER_STAGES) + 1)]


def score(stages: tuple[str, ...]) -> tuple[int, int, list[str]]:
    """How many cases parse to an object, and how many of those validate as Triage."""
    parsed = valid = 0
    failures = []
    for name, text in HOSTILE.items():
        try:
            data = parse_json(text, stages).value
        except ParseError:
            failures.append(name)
            continue
        parsed += 1
        try:
            validate_output(data, Triage)
            valid += 1
        except ValidationError:
            pass
    return parsed, valid, failures


if __name__ == "__main__":
    n = len(HOSTILE)
    print(f"{n} hostile outputs\n")
    print(f"{'parser stages':<50} {'parsed':>10} {'valid Triage':>14}")
    for stages in LADDER:
        parsed, valid, _ = score(stages)
        print(f"{' + '.join(stages):<50} {parsed:>3}/{n} {parsed / n:>4.0%} {valid:>6}/{n} {valid / n:>4.0%}")
    print("\nstill unparseable with every stage:", score(PARSER_STAGES)[2])
    print("\nrepairs logged per case (full parser):")
    for name, text in HOSTILE.items():
        try:
            print(f"  {name:<22} {parse_json(text).repairs}")
        except ParseError as exc:
            print(f"  {name:<22} ParseError: {exc}")

Code explained

  • In simple words: a zoo of broken answers, and a referee that counts how many each version of the parser can rescue.
  • What happens: HOSTILE maps a case name to a raw output: clean JSON (two cases), code fences, preambles and postambles, trailing commas, single quotes, a Python dict, smart quotes, braces inside a string, four kinds of truncation, three schema violations (wrong enum values, wrong types, an extra field), an array around the object, prose, a refusal, and one deliberately nasty case, a single-quoted string containing an apostrophe. LADDER lists the stage sets, from plain json.loads up to all five stages. score counts, for one stage set, how many cases parse into an object and how many of those also validate as Triage.
  • Comes out:
text
24 hostile outputs

parser stages                                          parsed   valid Triage
plain                                                5/24  21%      2/24   8%
plain + fences                                       8/24  33%      5/24  21%
plain + fences + extract                            12/24  50%      9/24  38%
plain + fences + extract + normalize                17/24  71%     14/24  58%
plain + fences + extract + normalize + truncation   21/24  88%     14/24  58%

still unparseable with every stage: ['prose_only', 'refusal', 'apostrophe_in_single']

repairs logged per case (full parser):
  clean                  []
  clean_pretty           []
  fence_json             ['fences']
  fence_bare             ['fences']
  preamble               ['surrounding_text']
  postamble              ['surrounding_text']
  preamble_fence_post    ['fences']
  trailing_comma         ['trailing_comma']
  trailing_comma_pretty  ['trailing_comma']
  single_quotes          ['single_quotes']
  python_dict            ['single_quotes', 'python_literals']
  smart_quotes           ['smart_quotes']
  braces_in_string       ['surrounding_text']
  truncated_in_value     ['closed_truncated']
  truncated_after_comma  ['closed_truncated']
  truncated_in_key       ['closed_truncated']
  fence_truncated        ['fences', 'closed_truncated']
  wrong_enum             []
  wrong_types            []
  extra_field            []
  array_wrapped          ['surrounding_text']
  prose_only             ParseError: no '{' in output
  refusal                ParseError: no '{' in output
  apostrophe_in_single   ParseError: Expecting ',' delimiter at char 83

Here is the same result as a before-and-after table. Each row adds one stage to the row above.

Parser stagesParsed (n=24)Valid Triage (n=24)What the new stage rescued
json.loads only5 (21%)2 (8%)Clean JSON; three parse but break the schema
+ strip fences8 (33%)5 (21%)Fenced output
+ extract object12 (50%)9 (38%)Preambles, postambles, text around braces, the array wrapper
+ normalize17 (71%)14 (58%)Trailing commas, single quotes, Python literals, smart quotes
+ close truncation21 (88%)14 (58%)Four truncated outputs parse, but none validates: all are missing fields

Read this carefully, because it is easy to over-claim. These rates describe a corpus I constructed, one case per failure type, not the frequency of each failure from any real model. The ordering is what transfers: each stage rescues a distinct class, and no stage broke a case an earlier stage handled (the test suite checks that). The truncation stage raises the parse rate but not the valid rate, which is the correct outcome: a truncated object is incomplete, and pretending otherwise would be worse. Those four go to salvage.

Diagnosing the three that still fail, read the log line, then decide:

  • prose_only and refusal: ParseError: no '{' in output. There is no JSON to recover; a parser should not invent one. These go to the repair loop, and if the model keeps refusing, to a human.
  • apostrophe_in_single: Expecting ',' delimiter at char 83. The model used single quotes and also wrote Customer's. The apostrophe ends the string early, and no local rule can tell an apostrophe from a closing quote in general. Fixing it would need guesswork that breaks other cases, so we let the repair loop ask again. Knowing when to stop adding heuristics is part of the design.

Repair loops with an attempt ceiling

A repair loop sends the model its own invalid answer plus the validation errors and asks for a corrected answer. An attempt ceiling is the maximum number of calls it may make. Without one, a model that keeps failing turns one ticket into an unbounded bill. ScriptedLLM stands in for the model here: it replays scripted replies, so this shows the loop's plumbing, not how often a real model fixes itself.

python
"""The repair loop with an attempt ceiling, driven by ScriptedLLM (plumbing only, not a model)."""
from __future__ import annotations

from examples.m06_structured import generate_structured
from supportdesk.schemas import Triage
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages

messages = [
    {"role": "system", "content": "Triage the ticket. Reply with only a JSON object for the Triage schema."},
    {"role": "user", "content": "Subject: Charged twice\n\nMy card was charged 288 USD twice on 3 September."},
]
BROKEN = ('Here you go: {"category": "billing", "priority": "critical", "language": "English", '
          '"summary": "Charged twice for the Team plan."}')
FIXED = ('{"category": "billing", "priority": "high", "language": "en", '
         '"summary": "Charged twice for the Team plan.", "needs_human": true}')

print("=== Case 1: broken, then fixed")
llm = ScriptedLLM(replies=[BROKEN, FIXED])
result = generate_structured(llm, messages, Triage, max_attempts=3)
print("value:", result.value)
print("attempts:", result.attempts, "| errors per failed attempt:", result.errors)
print("tokens spent:", result.usage.input_tokens, "in,", result.usage.output_tokens, "out")
print("the correction message the loop sent on attempt 2:")
print("  " + llm.calls[1]["messages"][-1]["content"].replace("\n", "\n  "))

print("\n=== Case 2: never fixed, the ceiling stops it")
stubborn = ScriptedLLM(responder=lambda msgs, kw: BROKEN)
result = generate_structured(stubborn, messages, Triage, max_attempts=3)
print("value:", result.value, "| attempts:", result.attempts, "| calls made:", len(stubborn.calls))
print("input tokens per call (estimate):", [count_messages(c["messages"]) for c in stubborn.calls])
print("tokens spent:", result.usage.input_tokens, "in,", result.usage.output_tokens, "out")
print("last parsed object kept for salvage:", result.last_data)

Code explained

  • In simple words: the first reply is broken in three ways, the loop explains the problems, and the second reply is fixed; then a stubborn "model" shows the ceiling at work.
  • What happens: in case 1, ScriptedLLM(replies=[BROKEN, FIXED]) returns a reply with a preamble, priority: "critical", language: "English", and no needs_human, then a correct object. generate_structured parses the first (the preamble is removed), validation fails, and the loop appends the error message you see printed. In case 2, a responder returns the broken reply forever. count_messages estimates each call's input tokens from the recorded messages.
  • Comes out:
text
=== Case 1: broken, then fixed
value: category='billing' priority='high' language='en' summary='Charged twice for the Team plan.' needs_human=True
attempts: 2 | errors per failed attempt: [["priority: Input should be 'low', 'normal', 'high' or 'urgent'", "language: String should match pattern '^[a-z]{2}$'", 'needs_human: Field required']]
tokens spent: 196 in, 72 out
the correction message the loop sent on attempt 2:
  Your reply did not match the schema:
  - priority: Input should be 'low', 'normal', 'high' or 'urgent'
  - language: String should match pattern '^[a-z]{2}$'
  - needs_human: Field required
  Reply again with only the corrected JSON object.

=== Case 2: never fixed, the ceiling stops it
value: None | attempts: 3 | calls made: 3
input tokens per call (estimate): [47, 149, 251]
tokens spent: 447 in, 105 out
last parsed object kept for salvage: {'category': 'billing', 'priority': 'critical', 'language': 'English', 'summary': 'Charged twice for the Team plan.'}

Case 1 took two calls. Case 2 stopped at exactly three calls, and each call was bigger than the last (47, then 149, then 251 estimated input tokens) because each retry carries the whole failed exchange. That growth is why the ceiling is small: with a real model, the third attempt rarely succeeds where two failed, while it always costs the most. A ceiling of 2 or 3 is typical; measure your own model's success rate per attempt before raising it. The loop keeps last_data so the caller can still salvage something.

SituationUse thisWhy
Validation errors on a field or twoRepair loop with the specific error lines, ceiling 2 to 3Targeted feedback is cheap and usually enough
No JSON at all (refusal, prose) twice in a rowStop and escalate to a humanRetrying the same prompt rarely changes a refusal
Output truncated (closed_truncated in repairs)Raise max_tokens or shrink the schema, then retry onceThe cause is the limit, not the model's understanding
Strict native mode on a supported modelCeiling of 1 or 2, still validateShape errors should not happen; semantic errors still can

Partial results and salvaging what is valid

When the ceiling is hit, or the output was truncated, you often still have useful fields. Salvaging means keeping the fields that pass validation and making a deliberate, safe decision about the rest, instead of either crashing or trusting everything.

python
"""Salvage: keep the valid fields of an invalid object, then apply a safe fallback policy."""
from __future__ import annotations

from examples.m06_hostile import HOSTILE
from examples.m06_structured import ParseError, parse_json, salvage
from supportdesk.schemas import Triage

SAFE_DEFAULTS = {"priority": "high", "needs_human": True}   # unknown means: a person looks, soon


def triage_or_fallback(data: dict) -> tuple[str, dict, dict]:
    """Return (status, triage_fields, problems). status is valid, salvaged, or unusable."""
    valid, problems = salvage(data, Triage)
    if not problems:
        return "valid", valid, {}
    if "category" not in valid:                      # without a category we cannot route at all
        return "unusable", valid, problems
    filled = {**valid, **{k: v for k, v in SAFE_DEFAULTS.items() if k not in valid}}
    filled["needs_human"] = True                     # anything salvaged gets human review
    return "salvaged", filled, problems


counts = {"valid": 0, "salvaged": 0, "unusable": 0, "unparseable": 0}
for name, text in HOSTILE.items():
    try:
        parsed = parse_json(text)
    except ParseError:
        counts["unparseable"] += 1
        continue
    status, fields, problems = triage_or_fallback(parsed.value)
    if "closed_truncated" in parsed.repairs and status == "valid":
        status = "salvaged"                          # a closed truncation is never fully trusted
    counts[status] += 1
    if status != "valid":
        route = {k: fields.get(k) for k in ("category", "priority", "needs_human")}
        print(f"{name:<22} {status:<9} route with {route}")
        print(f"{'':<32} problems={problems}")
print("\n", counts, f"(n={len(HOSTILE)})")

Code explained

  • In simple words: for every hostile output that parses but does not fully validate, keep the good parts and fill the gaps in the safest direction.
  • What happens: triage_or_fallback calls salvage. If nothing is wrong, the object is valid. If category itself is bad, the ticket cannot be routed, so it is unusable. Otherwise missing or invalid fields get SAFE_DEFAULTS (priority high, a human looks), and anything salvaged is forced to needs_human = True. An object whose repairs include closed_truncated is never counted as fully valid, even if every field happened to validate.
  • Comes out:
text
truncated_in_value     salvaged  route with {'category': 'billing', 'priority': 'high', 'needs_human': True}
                                 problems={'needs_human': 'Field required'}
truncated_after_comma  salvaged  route with {'category': 'billing', 'priority': 'high', 'needs_human': True}
                                 problems={'needs_human': 'Field required'}
truncated_in_key       salvaged  route with {'category': 'account_access', 'priority': 'high', 'needs_human': True}
                                 problems={'needs_human': 'Field required'}
fence_truncated        salvaged  route with {'category': 'cancellation', 'priority': 'normal', 'needs_human': True}
                                 problems={'needs_human': 'Field required'}
wrong_enum             unusable  route with {'category': None, 'priority': None, 'needs_human': True}
                                 problems={'category': "Input should be 'billing', 'cancellation', 'account_access', 'bug', 'how_to' or 'feature_request'", 'priority': "Input should be 'low', 'normal', 'high' or 'urgent'"}
wrong_types            salvaged  route with {'category': 'billing', 'priority': 'high', 'needs_human': True}
                                 problems={'language': "String should match pattern '^[a-z]{2}$'"}
extra_field            salvaged  route with {'category': 'billing', 'priority': 'high', 'needs_human': True}
                                 problems={'confidence': 'Extra inputs are not permitted'}

 {'valid': 14, 'salvaged': 6, 'unusable': 1, 'unparseable': 3} (n=24)

Of 24 outputs, 14 are fully valid, 6 are salvaged into a safe route, 1 is unusable (both enum fields wrong), and 3 never parsed. The wrong_types case also shows lax mode at work: "needs_human": "yes" did not appear as a problem because pydantic coerced it to True. The policy is the important design choice: salvage fails toward more human attention, never less. A salvaged ticket with a wrong priority costs Maya a minute; a salvaged ticket silently marked needs_human: false could leave a refund request unanswered.

Schema size, nesting depth, and their effect on accuracy

A schema is prompt text. In native mode the provider renders it into the model's context; in prompted mode you paste it yourself. Either way it costs tokens on every call, and bigger or deeper schemas give the model more to get wrong. The cost side is easy to measure here.

python
"""What a schema costs in tokens: formatting, descriptions, width, and nesting depth."""
from __future__ import annotations

import json
from typing import Any

from pydantic import BaseModel, ConfigDict, create_model

from examples.m06_structured import strict_schema
from supportdesk.schemas import Triage
from supportdesk.tokens import count_tokens


def tokens(schema: dict[str, Any], indent: int | None = None) -> int:
    text = json.dumps(schema, indent=indent, separators=None if indent else (",", ":"))
    return count_tokens(text, "o200k_base")


def without(schema: Any, keys: set[str]) -> Any:
    if isinstance(schema, dict):
        return {k: without(v, keys) for k, v in schema.items() if k not in keys}
    if isinstance(schema, list):
        return [without(v, keys) for v in schema]
    return schema


base = Triage.model_json_schema()
print("Triage schema, o200k_base tokens")
print(f"  pydantic output, indent=2          {tokens(base, 2):>5}")
print(f"  compact separators                 {tokens(base):>5}")
print(f"  strict (titles removed)            {tokens(strict_schema(Triage)):>5}")
print(f"  strict, descriptions removed       {tokens(without(strict_schema(Triage), {'description'})):>5}")

# Width: the same kind of field, repeated.
print("\nWidth: flat objects with N string fields (each with a 10-word description)")
for n in (5, 10, 20, 40):
    fields = {f"field_{i}": (str, ...) for i in range(n)}
    model = create_model(f"Wide{n}", __config__=ConfigDict(extra="forbid"), **fields)
    schema = strict_schema(model)
    for prop in schema["properties"].values():
        prop["description"] = "One short sentence the support agent will read before replying."
    print(f"  {n:>2} fields: {tokens(schema):>5} tokens")


# Depth: the same five Triage leaves, wrapped in 0 to 3 extra levels of objects.
def nest(model: type[BaseModel], depth: int) -> type[BaseModel]:
    for level in range(depth):
        model = create_model(f"Level{level}", __config__=ConfigDict(extra="forbid"), inner=(model, ...))
    return model


print("\nDepth: Triage's five fields wrapped in extra object levels")
for depth in range(4):
    schema = strict_schema(nest(Triage, depth))
    example = json.loads(Triage(category="billing", priority="high", language="en",
                                summary="Customer was charged twice.", needs_human=True).model_dump_json())
    for _ in range(depth):
        example = {"inner": example}
    print(f"  depth {depth + 1}: schema {tokens(schema):>4} tokens (uses $defs: {'$defs' in schema}), "
          f"one answer {count_tokens(json.dumps(example), 'o200k_base'):>3} tokens")

Code explained

  • In simple words: count what the schema itself costs, and how that grows with more fields and deeper nesting.
  • What happens: tokens() serializes a schema and counts it with the o200k_base tokenizer from supportdesk/tokens.py. The first block compares four versions of the Triage schema. The width block builds flat models with 5 to 40 string fields using pydantic's create_model, each with the same 10-word description. The depth block wraps Triage in 0 to 3 extra objects and also counts the tokens of one matching answer.
  • Comes out:
text
Triage schema, o200k_base tokens
  pydantic output, indent=2            381
  compact separators                   248
  strict (titles removed)              222
  strict, descriptions removed         113

Width: flat objects with N string fields (each with a 10-word description)
   5 fields:   146 tokens
  10 fields:   276 tokens
  20 fields:   536 tokens
  40 fields:  1056 tokens

Depth: Triage's five fields wrapped in extra object levels
  depth 1: schema  222 tokens (uses $defs: False), one answer  34 tokens
  depth 2: schema  257 tokens (uses $defs: True), one answer  37 tokens
  depth 3: schema  290 tokens (uses $defs: True), one answer  41 tokens
  depth 4: schema  323 tokens (uses $defs: True), one answer  44 tokens

Formatting alone matters: the indented schema costs 381 tokens, the compact one 248, and removing pydantic's titles brings it to 222. Descriptions are about half of what remains (222 vs 113 without them), and they are usually worth it, because they carry the rubric. Width is linear, about 26 tokens per described field. Each nesting level adds about 33 schema tokens and 3 to 4 answer tokens here; pydantic switches to $defs and $ref as soon as there is a nested model, and those references are exactly what some providers limit. Counts are for the JSON text; each provider renders schemas into its own prompt format, so treat them as estimates and confirm with the usage a real response reports.

The accuracy side cannot be measured without a real model, and published evidence is mixed, so here is what the research says and what it does not:

  • Tam et al., "Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models" (2024, arXiv:2408.02442), report "a significant decline in LLMs' reasoning abilities under format restrictions," with stricter constraints degrading reasoning tasks more.
  • Geng et al., "JSONSchemaBench" (2025, arXiv:2501.10868), test six constrained-decoding systems, including OpenAI's and Gemini's, on 10,000 real-world schemas and find that no framework handles every schema feature and complexity level; coverage of complex schemas varies widely.
  • Gemini's docs warn that very large or deeply nested schemas may be rejected.

None of these gives a number you can apply to Brightlane's triage. So measure it, with a harness that compares the flat Triage against a two-level NestedTriage holding the same five facts.

python
"""Does a nested schema change accuracy? A harness to measure it on the dev tickets.

  LLM_PROVIDER=groq GROQ_API_KEY=... PYTHONPATH=. python examples/m06_schema_eval.py
  PYTHONPATH=. python examples/m06_schema_eval.py --dry     (keyword baseline: checks the harness only)
"""
from __future__ import annotations

import json
import sys

from pydantic import BaseModel, ConfigDict

from examples.m06_lab import keyword_triage, wilson
from examples.m06_structured import generate_structured, response_format_for
from supportdesk.data import load_tickets
from supportdesk.schemas import Category, Priority, Triage
from supportdesk.stand_in import ScriptedLLM


class Routing(BaseModel):
    model_config = ConfigDict(extra="forbid")
    category: Category
    priority: Priority
    needs_human: bool


class Customer(BaseModel):
    model_config = ConfigDict(extra="forbid")
    language: str
    summary: str


class NestedTriage(BaseModel):
    """The same five facts as Triage, grouped two levels deep."""
    model_config = ConfigDict(extra="forbid")
    routing: Routing
    customer: Customer


def category_of(value: BaseModel) -> str:
    return value.routing.category if isinstance(value, NestedTriage) else value.category


def run(schema_cls: type[BaseModel], llm, tickets) -> tuple[int, int, int]:
    valid = correct = out_tokens = 0
    for t in tickets:
        messages = [{"role": "system", "content": "Triage this Brightlane support ticket as JSON."},
                    {"role": "user", "content": t.text}]
        res = generate_structured(llm, messages, schema_cls, max_attempts=1,
                                  response_format=response_format_for(schema_cls), max_tokens=2000)
        out_tokens += res.usage.output_tokens
        if res.value is not None:
            valid += 1
            correct += category_of(res.value) == t.gold["category"]
    return valid, correct, out_tokens


if __name__ == "__main__":
    tickets = load_tickets("dev")
    if "--dry" in sys.argv:
        by_text = {t.text: t for t in tickets}

        def dry(messages, kwargs):
            d = keyword_triage(by_text[messages[1]["content"]])
            if kwargs["response_format"]["json_schema"]["name"] == "NestedTriage":
                d = {"routing": {k: d[k] for k in ("category", "priority", "needs_human")},
                     "customer": {k: d[k] for k in ("language", "summary")}}
            return json.dumps(d)
        llm = ScriptedLLM(responder=dry)
    else:
        from supportdesk.llm import chat as llm
    n = len(tickets)
    for schema_cls in (Triage, NestedTriage):
        valid, correct, out_tokens = run(schema_cls, llm, tickets)
        lo, hi = wilson(correct, n)
        print(f"{schema_cls.__name__:<13} valid {valid}/{n}  category accuracy {correct}/{n} = {correct / n:.0%} "
              f"(95% CI {lo:.0%} to {hi:.0%})  output tokens {out_tokens}")

Code explained

  • In simple words: the same tickets, the same facts, two shapes; count valid answers, correct categories, and output tokens for each.
  • What happens: NestedTriage groups the fields into routing and customer. run sends each dev ticket through generate_structured with one attempt (so the first answer is what gets scored) and the strict response_format for that schema. --dry swaps in a ScriptedLLM whose answers come from the lab's keyword baseline, reshaped for the nested schema, so you can check the harness runs end to end without a key.
  • Comes out: the dry run (keyword baseline, not a model):