CourseLarge Language Models · Module 6: Structured Outputs and Tool Use · part 28 of 80
Part 28 · Module 6: Structured Outputs and Tool Use

Part B: Tool and function calling

34 min read·22 Sept 2026

How tool calling works from the model's side

Tool calling (also called function calling) lets the model ask your program to run a function. The model never runs anything. It writes a request, in a special format, that says which tool and which arguments; your code decides whether and how to run it, and sends the result back as a new message.

From the model's side, a tool definition is just more text in its context. The provider takes the tools list from your request and renders it into the prompt using the model's chat template. The model then generates tokens, and if it decides to call a tool, it emits them in a tool-call format that the provider parses back into the structured tool_calls field you receive.

How Tool Calling Works

.

Here is the full request llm.chat sends with the three Brightlane tools, captured with capture_wire, and an approximation of the text a gpt-oss model is given.

python
"""What the model is actually given: the full request with tools, captured offline,
and an approximation of how a gpt-oss model sees those tools as text."""
from __future__ import annotations

import json

from examples.m06_tools import SYSTEM, TICKET, TOOLS
from examples.m06_wire import capture_wire
from supportdesk.llm import chat
from supportdesk.tokens import count_tokens

tools = [t.schema() for t in TOOLS.values()]
messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": TICKET}]

# The canned reply asks for one tool, the way a real provider response would.
canned = {"content": None, "tool_calls": [{"id": "call_a1", "type": "function", "function": {
    "name": "get_invoice", "arguments": "{\"invoice_id\": \"INV-2026-004512\"}"}}]}
with capture_wire(canned, finish_reason="tool_calls") as bodies:
    result = chat(messages, provider="groq", tools=tools, tool_choice="auto")

print("=== Request body sent by llm.chat")
print(json.dumps(bodies[0], indent=2))
print("\n=== Parsed by llm.chat")
print("finish_reason:", result.finish_reason)
print("tool_calls:", result.tool_calls)


def ts_type(prop: dict) -> str:
    if "enum" in prop:
        return " | ".join(json.dumps(v) for v in prop["enum"])
    return {"string": "string", "integer": "number", "number": "number", "boolean": "boolean"}[prop["type"]]


def render_harmony(tool_list: list) -> str:
    """Approximates the 'namespace functions' block of OpenAI's harmony format (used by gpt-oss).
    The provider does the real rendering; this is for seeing and counting, not for sending."""
    lines = ["# Tools", "", "## functions", "", "namespace functions {", ""]
    for t in tool_list:
        lines.append(f"// {t.description}")
        lines.append(f"type {t.name} = (_: {{")
        schema = t.args_model.model_json_schema()
        for name, prop in schema["properties"].items():
            if "description" in prop:
                lines.append(f"// {prop['description']}")
            optional = "" if name in schema.get("required", []) else "?"
            default = f" // default: {prop['default']}" if "default" in prop else ""
            lines.append(f"{name}{optional}: {ts_type(prop)},{default}")
        lines += ["}) => any;", ""]
    lines.append("} // namespace functions")
    return "\n".join(lines)


rendered = render_harmony(list(TOOLS.values()))
print("\n=== The same tools as a gpt-oss model reads them (approximation)")
print(rendered)
print("\ntokens (o200k_base): tools as JSON", count_tokens(json.dumps(tools)), "| rendered", count_tokens(rendered),
      "| system + ticket", count_tokens(SYSTEM + TICKET))

Code explained

  • In simple words: print exactly what leaves your machine when you offer tools, then print the same tools the way the model family behind the course default reads them.
  • What happens: tools comes from Tool.schema() in examples/m06_tools.py (shown in full in the next section). Inside capture_wire, chat(messages, provider="groq", tools=tools, tool_choice="auto") runs, and the canned reply contains one tool call, so we also see llm.chat parse tool_calls. render_harmony approximates the "namespace functions" block of OpenAI's harmony format, the chat format gpt-oss models are trained on, as documented in OpenAI's harmony guide: each tool becomes a TypeScript-like type with its description as a comment, optional fields get ?, and enums become "a" | "b". The provider does the real rendering; this function exists only so we can read and count it.
  • Comes out:

Writing tool descriptions the model can actually use

Since the description is the only documentation the model gets, write it like an API doc for a new colleague who reads fast:

  • Say what the tool does and when to use it: "Use it before any refund."
  • Say what it returns, so the model knows whether it answers the question.
  • Say the limits and the policy that applies: "Only for duplicate charges or annual plans cancelled within 14 days."
  • Say what to do on the unhappy path: "If not approved, tell the customer a teammate will review."
  • Put examples of formats in parameter descriptions: "like INV-2026-004512."

Keep names verb-first and unambiguous (get_invoice, not invoice), and do not offer two tools whose descriptions overlap; the model will split its choices between them.

Parameter design and enum constraints

Parameters are the tool's schema, so everything from the schema design section applies, with one addition: every argument is a decision the model can get wrong, so ask for as few as possible and constrain each one.

SituationUse thisWhy
A choice among known options (refund reason)An enum with a description of each valueThe model picks from a list instead of inventing a reason; the executor rejects anything else
An identifier (invoice number)A string with a pattern, and an example in the descriptionCatches "4512" before a database lookup
A quantity (amount, result count)Numeric bounds, plus a business check against real datale=10000 is a sanity net; "not more than this invoice" needs the database
Something your code already knows (customer id, user role)Not a parameter at all; take it from the session contextThe model cannot misuse what it cannot set
An optional knob with a good default (k)Optional with a default, nullable in strict modeThe model can ignore it

The last row is the most important security decision in this module. get_invoice takes only invoice_id. The customer it belongs to comes from ToolContext, which our code builds from the ticket's account. If customer_id were a parameter, a prompt injection in a ticket could ask for any customer's invoices.

The execution loop: call, execute, return result, continue

Here is the full tool module: fake billing database, argument models, tools, a safe executor, and the loop.

examples/m06_tools.py

python
"""Brightlane tools and a safe tool-calling loop.

Three tools:
  search_kb      read-only: BM25 search over the help center (supportdesk.kb_search)
  get_invoice    read-only: look up one invoice in a fake in-memory billing database
  issue_refund   privileged: moves money, so it needs a passing policy check AND human approval

Every call from the model goes through `execute`, which never raises: unknown
tools, bad arguments, permission problems, policy violations, timeouts, and
crashes all come back as a JSON error string the model can read and act on.

Run the demo:  PYTHONPATH=. python examples/m06_tools.py          (ScriptedLLM, no key)
               PYTHONPATH=. python examples/m06_tools.py --live   (llm.chat, needs a key)
"""
from __future__ import annotations

import json
import sys
import threading
import time
from concurrent.futures import ThreadPoolExecutor
from concurrent.futures import TimeoutError as FutureTimeout
from dataclasses import dataclass, field
from typing import Any, Callable, Literal

from pydantic import BaseModel, ConfigDict, Field, ValidationError

from examples.m06_structured import describe_errors, strict_schema, validate_output
from supportdesk.kb_search import KBSearch
from supportdesk.llm import ChatResult, ToolCall

INVOICE_ID = r"^INV-\d{4}-\d{6}$"

# ---------------------------------------------------------------------------
# A fake billing database (in memory, so every run starts from the same state)
# ---------------------------------------------------------------------------


@dataclass
class Invoice:
    invoice_id: str
    customer_id: str
    plan: str
    amount_usd: float
    issued: str
    refunded_usd: float = 0.0
    note: str = ""

    @property
    def refundable_usd(self) -> float:
        return round(self.amount_usd - self.refunded_usd, 2)


def make_billing_db() -> dict[str, Invoice]:
    return {inv.invoice_id: inv for inv in [
        Invoice("INV-2026-004512", "cus_acme", "team", 288.00, "2026-09-03"),
        Invoice("INV-2026-004513", "cus_acme", "team", 288.00, "2026-09-03", note="same card, same minute as INV-2026-004512"),
        Invoice("INV-2026-004871", "cus_sol", "team", 144.00, "2026-09-01"),
        Invoice("INV-2026-004872", "cus_sol", "team", 144.00, "2026-09-01", note="same card, same minute as INV-2026-004871"),
        Invoice("INV-2026-003990", "cus_acme", "business", 5760.00, "2026-07-15", note="annual plan"),
    ]}


# ---------------------------------------------------------------------------
# Argument models: the contract the model must satisfy before anything runs
# ---------------------------------------------------------------------------


class SearchKBArgs(BaseModel):
    model_config = ConfigDict(extra="forbid")
    query: str = Field(min_length=3, max_length=200, description="Keywords from the customer's question, in English.")
    k: int = Field(default=3, ge=1, le=5, description="How many articles to return (1 to 5).")


class GetInvoiceArgs(BaseModel):
    model_config = ConfigDict(extra="forbid")
    invoice_id: str = Field(pattern=INVOICE_ID, description="Invoice number exactly as written, like INV-2026-004512.")


class IssueRefundArgs(BaseModel):
    model_config = ConfigDict(extra="forbid")
    invoice_id: str = Field(pattern=INVOICE_ID, description="The invoice to refund, like INV-2026-004513.")
    amount_usd: float = Field(gt=0, le=10_000, description="Amount to refund in USD. Never more than the invoice's refundable amount.")
    reason: Literal["duplicate_charge", "annual_within_14_days"] = Field(
        description="duplicate_charge: the same charge was taken twice. annual_within_14_days: annual plan cancelled within 14 days of purchase or renewal.")


# ---------------------------------------------------------------------------
# Tools, context, and the executor
# ---------------------------------------------------------------------------


class ToolError(Exception):
    """An expected failure the model should hear about (not found, not allowed, too much)."""


@dataclass
class ToolContext:
    """What this conversation is allowed to touch. Built by our code, never by the model."""
    customer_id: str
    allowed_tools: frozenset[str]
    approver: Callable[[str, BaseModel], bool] = lambda name, args: False   # default: deny
    db: dict[str, Invoice] = field(default_factory=make_billing_db)
    audit: list[dict[str, Any]] = field(default_factory=list)


@dataclass(frozen=True)
class Tool:
    name: str
    description: str
    args_model: type[BaseModel]
    run: Callable[[Any, ToolContext], dict[str, Any]]
    precheck: Callable[[Any, ToolContext], None] | None = None   # policy check before approval
    privileged: bool = False
    timeout_s: float = 2.0

    def schema(self) -> dict[str, Any]:
        """The definition sent in the request's `tools` list."""
        return {"type": "function", "function": {
            "name": self.name, "description": self.description, "parameters": strict_schema(self.args_model)}}


_KB = KBSearch()


def _search_kb(args: SearchKBArgs, ctx: ToolContext) -> dict[str, Any]:
    hits = _KB.search(args.query, k=args.k)
    return {"results": [{"article_id": h.article_id, "title": h.title, "score": h.score} for h in hits]}


def _own_invoice(invoice_id: str, ctx: ToolContext) -> Invoice:
    invoice = ctx.db.get(invoice_id)
    if invoice is None or invoice.customer_id != ctx.customer_id:   # same message: do not reveal other customers' invoices
        raise ToolError(f"No invoice {invoice_id} for this customer. Ask the customer to confirm the number.")
    return invoice


def _get_invoice(args: GetInvoiceArgs, ctx: ToolContext) -> dict[str, Any]:
    inv = _own_invoice(args.invoice_id, ctx)
    return {"invoice_id": inv.invoice_id, "plan": inv.plan, "amount_usd": inv.amount_usd, "issued": inv.issued,
            "refunded_usd": inv.refunded_usd, "refundable_usd": inv.refundable_usd, "note": inv.note}


def _check_refund(args: IssueRefundArgs, ctx: ToolContext) -> None:
    inv = _own_invoice(args.invoice_id, ctx)
    if args.amount_usd > inv.refundable_usd:
        raise ToolError(f"Refund of {args.amount_usd:.2f} USD exceeds the refundable amount of "
                        f"{inv.refundable_usd:.2f} USD on {inv.invoice_id}.")


_REFUND_LOCK = threading.Lock()


def _issue_refund(args: IssueRefundArgs, ctx: ToolContext) -> dict[str, Any]:
    with _REFUND_LOCK:                           # check and update as one step, even for parallel calls
        _check_refund(args, ctx)                 # check again at execution time: state may have changed
        inv = ctx.db[args.invoice_id]
        inv.refunded_usd = round(inv.refunded_usd + args.amount_usd, 2)
    return {"status": "refunded", "invoice_id": inv.invoice_id, "amount_usd": args.amount_usd,
            "refundable_usd_now": inv.refundable_usd, "arrives": "5 to 10 business days, to the original payment method"}


TOOLS: dict[str, Tool] = {t.name: t for t in [
    Tool("search_kb", "Search Brightlane help-center articles. Use it before answering any how-to, billing, or policy "
         "question. Returns article ids and titles, best match first.", SearchKBArgs, _search_kb),
    Tool("get_invoice", "Look up one invoice of the current customer: amount, plan, date, and how much is still "
         "refundable. Use it before any refund.", GetInvoiceArgs, _get_invoice),
    Tool("issue_refund", "Refund part or all of one invoice to the original payment method. Only for duplicate charges "
         "or annual plans cancelled within 14 days. A human must approve it; if not approved, tell the customer a "
         "teammate will review.", IssueRefundArgs, _issue_refund, precheck=_check_refund, privileged=True),
]}

READ_ONLY = frozenset({"search_kb", "get_invoice"})
WITH_REFUNDS = READ_ONLY | {"issue_refund"}


@dataclass
class ToolOutcome:
    call_id: str
    name: str
    ok: bool
    content: str          # JSON string, exactly what goes back to the model
    ms: float


def _error(message: str, kind: str) -> str:
    return json.dumps({"error": message, "error_type": kind})


def execute(call: ToolCall, ctx: ToolContext, registry: dict[str, Tool] = TOOLS) -> ToolOutcome:
    """Run one tool call safely. Every failure becomes a JSON error string; nothing raises."""
    started = time.perf_counter()

    def done(ok: bool, content: str) -> ToolOutcome:
        ms = round((time.perf_counter() - started) * 1000, 1)
        ctx.audit.append({"call_id": call.id, "tool": call.name, "arguments": call.arguments, "ok": ok, "ms": ms})
        return ToolOutcome(call.id, call.name, ok, content, ms)

    tool = registry.get(call.name)
    if tool is None or call.name not in ctx.allowed_tools:
        return done(False, _error(f"Tool {call.name!r} is not available. Available: {sorted(ctx.allowed_tools)}.", "not_available"))
    if "_unparseable" in call.arguments:
        return done(False, _error("Arguments were not valid JSON. Send a JSON object.", "bad_arguments"))
    try:
        args = validate_output(call.arguments, tool.args_model)
    except ValidationError as exc:
        return done(False, _error("Invalid arguments: " + "; ".join(describe_errors(exc)), "bad_arguments"))
    try:
        if tool.precheck is not None:
            tool.precheck(args, ctx)
        if tool.privileged and not ctx.approver(tool.name, args):
            return done(False, _error("This action needs human approval and was not approved now. Do not retry "
                                      "it; tell the customer a teammate will review the request.", "not_approved"))
        pool = ThreadPoolExecutor(max_workers=1)
        future = pool.submit(tool.run, args, ctx)
        try:
            result = future.result(timeout=tool.timeout_s)
        finally:
            pool.shutdown(wait=False, cancel_futures=True)
        return done(True, json.dumps(result))
    except ToolError as exc:
        return done(False, _error(str(exc), "rejected"))
    except FutureTimeout:
        return done(False, _error(f"{tool.name} took longer than {tool.timeout_s}s and was abandoned. Try once more "
                                  "or continue without it.", "timeout"))
    except Exception as exc:  # a bug in our tool: log it, tell the model only that it failed
        return done(False, _error(f"{tool.name} failed internally ({type(exc).__name__}).", "internal"))


# ---------------------------------------------------------------------------
# The loop: call the model, execute its tool calls, return results, continue
# ---------------------------------------------------------------------------


@dataclass
class LoopResult:
    text: str
    messages: list[dict[str, Any]]
    turns: int
    stop_reason: str                      # "answered" or "turn_limit"
    outcomes: list[ToolOutcome] = field(default_factory=list)


def run_tool_loop(
    llm: Callable[..., ChatResult],
    messages: list[dict[str, Any]],
    ctx: ToolContext,
    *,
    registry: dict[str, Tool] = TOOLS,
    max_turns: int = 5,
    tool_choice: str | dict[str, Any] = "auto",
    verbose: bool = False,
    **chat_kwargs: Any,
) -> LoopResult:
    """Run the tool loop until the model answers in text or `max_turns` model calls are used.

    The last allowed turn is sent with tool_choice="none", so the model must answer.
    Tool calls in one turn run in parallel; results go back in the order of the calls,
    each tagged with its tool_call_id.
    """
    history = list(messages)
    offered = [registry[name].schema() for name in sorted(ctx.allowed_tools) if name in registry]
    outcomes: list[ToolOutcome] = []
    for turn in range(1, max_turns + 1):
        choice = "none" if turn == max_turns else (tool_choice if turn == 1 else "auto")
        result = llm(history, tools=offered, tool_choice=choice, **chat_kwargs)
        history.append(result.as_message())
        if not result.tool_calls:
            return LoopResult(result.text, history, turn, "answered", outcomes)
        if choice == "none":                          # the server ignored tool_choice: run nothing more
            break
        with ThreadPoolExecutor(max_workers=4) as pool:
            turn_outcomes = list(pool.map(lambda c: execute(c, ctx, registry), result.tool_calls))
        for outcome in turn_outcomes:                 # same order as result.tool_calls
            history.append({"role": "tool", "tool_call_id": outcome.call_id, "content": outcome.content})
            if verbose:
                status = "ok " if outcome.ok else "ERR"
                print(f"  turn {turn} {status} {outcome.name}({json.dumps(next(c.arguments for c in result.tool_calls if c.id == outcome.call_id))})")
                print(f"           -> {outcome.content[:110]}")
        outcomes += turn_outcomes
    return LoopResult("", history, max_turns, "turn_limit", outcomes)


# ---------------------------------------------------------------------------
# Demo
# ---------------------------------------------------------------------------

SYSTEM = ("You are Brightlane's support assistant. Use the tools to check facts before you answer. "
          "Refunds follow the help center policy. Keep replies short and friendly.")
TICKET = ("Subject: Charged twice this month\n\nHi, my card was charged 288 USD twice on 3 September for the Team "
          "plan (invoice INV-2026-004512). Please refund the duplicate.")


def call(call_id: str, name: str, **arguments: Any) -> ToolCall:
    return ToolCall(call_id, name, arguments, json.dumps(arguments))


def scripted_refund_run() -> list[ChatResult]:
    """What a good model might do on this ticket. Scripted: this is plumbing, not model output."""
    return [
        ChatResult(text="", tool_calls=[
            call("call_1", "search_kb", query="duplicate charge refund"),
            call("call_2", "get_invoice", invoice_id="INV-2026-004512"),
            call("call_3", "get_invoice", invoice_id="INV-2026-004513"),
        ]),
        ChatResult(text="", tool_calls=[
            call("call_4", "issue_refund", invoice_id="INV-2026-004513", amount_usd=288.0, reason="duplicate_charge")]),
        ChatResult(text="Sorry about the double charge. We refunded the duplicate 288 USD (INV-2026-004513); it "
                        "reaches your card in 5 to 10 business days. [billing-refunds]"),
    ]


def console_approver(name: str, args: BaseModel) -> bool:
    return input(f"Approve {name} {args.model_dump()}? [y/N] ").strip().lower() == "y"


if __name__ == "__main__":
    messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": TICKET}]
    if "--live" in sys.argv:
        from supportdesk.llm import chat
        ctx = ToolContext("cus_acme", WITH_REFUNDS, approver=console_approver)
        result = run_tool_loop(chat, messages, ctx, verbose=True, max_tokens=2000)
    else:
        from supportdesk.stand_in import ScriptedLLM
        ctx = ToolContext("cus_acme", WITH_REFUNDS, approver=lambda name, args: True)  # Maya clicks approve
        result = run_tool_loop(ScriptedLLM(replies=scripted_refund_run()), messages, ctx, verbose=True)
    print(f"\nstop: {result.stop_reason} after {result.turns} model calls")
    print("final reply:", result.text)
    print("INV-2026-004513 refunded_usd now:", ctx.db["INV-2026-004513"].refunded_usd)
    print("audit log:", [(a["tool"], a["ok"]) for a in ctx.audit])

Code explained

  • In simple words: three tools, a gatekeeper that checks every request before running it, and a loop that keeps going until the model answers in words or runs out of turns.
  • What happens:
    • Invoice and make_billing_db() are a fake billing database in memory. refundable_usd is the amount not yet refunded. Customer cus_acme has the two 288 USD Team invoices from ticket T-1001 (one is the duplicate); cus_sol has the pair from the Spanish ticket T-1031.
    • SearchKBArgs, GetInvoiceArgs, and IssueRefundArgs are the argument contracts: bounds on query and k, an invoice-number pattern, an amount bounded to (0, 10000], and an enum of the two refund reasons the help-center article allows. extra="forbid" rejects invented arguments.
    • ToolError marks expected failures (not found, too much) whose message is safe and useful for the model.
    • ToolContext is what this conversation may touch: the customer, the set of allowed tool names, an approver callback (default: deny), the database, and an audit list. Our code builds it; the model never sees or sets it.
    • Tool bundles a name, description, argument model, run function, optional precheck (a policy check that runs before approval), a privileged flag, and a timeout. schema() produces the definition for the request using strict_schema.
    • _search_kb wraps KBSearch.search. _own_invoice returns an invoice only if it belongs to the context's customer, and gives the same message for "does not exist" and "belongs to someone else", so the tool cannot be used to probe other customers' invoice numbers. _get_invoice returns the invoice fields. _check_refund rejects amounts above the refundable amount. _issue_refund repeats that check and updates the amount under a lock, so two parallel refund calls cannot both pass the check before either writes.
    • TOOLS is the registry; READ_ONLY and WITH_REFUNDS are the two permission scopes.
    • execute(call, ctx, registry) is the gatekeeper, in this order: tool exists and is allowed in this context; arguments were valid JSON; arguments validate against the model; policy precheck; human approval for privileged tools; run in a worker thread with future.result(timeout=...). Every failure returns a JSON string with error and error_type. An unexpected exception returns only its type name, never its message, so internal details do not leak into the model's context. Every call, success or failure, is appended to ctx.audit.
    • run_tool_loop offers only the tools in ctx.allowed_tools, calls the model, appends the assistant message (ChatResult.as_message() includes the tool_calls), and if there are tool calls, executes them concurrently with pool.map, which returns results in call order. Each result becomes a {"role": "tool", "tool_call_id": ..., "content": ...} message. The requested tool_choice applies to the first turn only (otherwise a forced tool would be called forever), later turns use "auto", and the last allowed turn uses "none". If a server ignores that and still returns tool calls, the loop stops without running them.
    • call() is a small helper for building ToolCall objects in scripts and tests; scripted_refund_run() scripts what a good agent might do on T-1001; console_approver asks at the terminal in live runs.
  • Comes out: running the file executes the demo below.
bash
python examples/m06_tools.py

Code explained

  • In simple words: run the T-1001 duplicate-charge ticket through the loop with a scripted model and an approver that says yes.
  • What happens: ScriptedLLM replays three replies from scripted_refund_run(): three tool calls at once, then a refund, then a final answer. This is plumbing, not model output. The approver lambda stands in for Maya clicking "approve". verbose=True prints each tool result (truncated to 110 characters).
  • Comes out:

To run the same loop against a real model, with you as the approver:

bash
LLM_PROVIDER=groq GROQ_API_KEY=your-key python examples/m06_tools.py --live

Code explained

  • In simple words: the same loop, with llm.chat deciding which tools to call and you approving refunds at the keyboard.
  • What happens: --live passes chat as the llm argument and console_approver as the approver; max_tokens=2000 leaves room for gpt-oss reasoning. Any provider works through LLM_PROVIDER (for Ollama, run a model that supports tools, and read the tool_choice note below).
  • Comes out: Illustrative sample run (not captured in this build; produced for teaching). Your output will differ.

Returning errors to the model instead of raising

If a tool raises, your loop crashes and the customer gets nothing. If a tool returns an error as its result, the model reads it and can recover: fix an argument, try another tool, or tell the customer what is missing. Groq's tool-use docs recommend the same pattern: return the error in the tool message so "the model can adjust its approach" (Groq local tool calling).

python
"""Errors go back to the model as data. Each bad call below comes back as a JSON error string."""
from __future__ import annotations

from examples.m06_tools import READ_ONLY, WITH_REFUNDS, ToolContext, call, execute
from supportdesk.llm import ToolCall

ctx = ToolContext("cus_acme", WITH_REFUNDS, approver=lambda name, args: True)
bad_calls = [
    ("made-up tool", call("c1", "delete_account", customer_id="cus_acme")),
    ("not JSON", ToolCall("c2", "get_invoice", {"_unparseable": "{invoice_id: 4512"}, "{invoice_id: 4512")),
    ("wrong format", call("c3", "get_invoice", invoice_id="4512")),
    ("extra argument", call("c4", "search_kb", query="refund", k=3, language="en")),
    ("out of range", call("c5", "search_kb", query="refund", k=50)),
    ("other customer", call("c6", "get_invoice", invoice_id="INV-2026-004871")),
    ("does not exist", call("c7", "get_invoice", invoice_id="INV-2026-999999")),
    ("enum violated", call("c8", "issue_refund", invoice_id="INV-2026-004513", amount_usd=288, reason="angry_customer")),
]
for label, bad in bad_calls:
    outcome = execute(bad, ctx)
    print(f"{label:<15} ok={outcome.ok!s:<5} {outcome.content}")

print("\nread-only session asks for a refund:")
readonly = ToolContext("cus_acme", READ_ONLY)
print(" ", execute(call("c9", "issue_refund", invoice_id="INV-2026-004513", amount_usd=288, reason="duplicate_charge"), readonly).content)

Code explained

  • In simple words: throw eight kinds of bad call at the gatekeeper and read what the model would be told.
  • What happens: each call(...) builds a ToolCall like the ones llm.chat returns. The "not JSON" case uses the _unparseable marker that llm.chat produces when the arguments string is not valid JSON. The last call uses a read-only context.
  • Comes out:

Parallel tool calls and result ordering

A model can request several tools in one turn, as in the T-1001 run. These are parallel tool calls: independent lookups your code may run concurrently. The contract with the API is about identity, not timing: each tool message must carry the tool_call_id of the call it answers, and every call in the assistant message needs a result before the next model call. OpenAI's Chat Completions API returns an error when an assistant message with tool_calls is not followed by a tool message for each id, and you should assume any compatible provider may do the same. Returning results in call order is the simple convention that keeps logs and tests readable.

python
"""Parallel tool calls: run them concurrently, return results in call order, matched by tool_call_id."""
from __future__ import annotations

import dataclasses
import json
import time
from concurrent.futures import ThreadPoolExecutor

from examples.m06_tools import TOOLS, WITH_REFUNDS, ToolContext, call, execute

finished: list[str] = []


def with_latency(tool, seconds: float):
    """Wrap a tool so it takes `seconds`, like a slow network API, and records when it finishes."""
    def run(args, ctx):
        time.sleep(seconds)
        out = tool.run(args, ctx)
        finished.append(tool.name + ":" + json.dumps(args.model_dump())[:40])
        return out
    return dataclasses.replace(tool, run=run)


registry = {"search_kb": with_latency(TOOLS["search_kb"], 0.30),
            "get_invoice": with_latency(TOOLS["get_invoice"], 0.10)}
ctx = ToolContext("cus_acme", WITH_REFUNDS)
calls = [call("call_1", "search_kb", query="duplicate charge refund"),
         call("call_2", "get_invoice", invoice_id="INV-2026-004512"),
         call("call_3", "get_invoice", invoice_id="INV-2026-004513")]

started = time.perf_counter()
sequential = [execute(c, ctx, registry) for c in calls]
seq_ms = (time.perf_counter() - started) * 1000

finished.clear()
started = time.perf_counter()
with ThreadPoolExecutor(max_workers=4) as pool:
    parallel = list(pool.map(lambda c: execute(c, ctx, registry), calls))
par_ms = (time.perf_counter() - started) * 1000

print(f"sequential: {seq_ms:.0f} ms   parallel: {par_ms:.0f} ms")
print("finished in this order:", finished)
print("tool messages appended in this order:")
for outcome in parallel:
    print(" ", {"role": "tool", "tool_call_id": outcome.call_id, "content": outcome.content[:60] + "..."})

Code explained

  • In simple words: make the lookups slow on purpose, run them one by one and then all at once, and check that the answers go back in the right order even when they finish in a different order.
  • What happens: with_latency wraps a tool so search_kb takes 0.30 s and each get_invoice 0.10 s, and records when each finishes. The sequential run calls execute three times in a row. The parallel run uses ThreadPoolExecutor.map, which yields results in input order regardless of completion order.
  • Comes out:

Tool choice control: auto, required, forced, none

tool_choice tells the provider how much freedom the model has:

  • "auto": the model decides whether to call tools or answer. The default.
  • "required": the model must call at least one tool.
  • Forced: {"type": "function", "function": {"name": "get_invoice"}} makes it call that specific tool.
  • "none": the model must answer in text, even though tools are defined.

Groq documents all four (Groq local tool calling). Gemini's native API has matching modes (AUTO, ANY, NONE, plus VALIDATED) (Gemini function calling). Ollama's OpenAI-compatibility page lists tool_choice as not supported, so a forced choice there may be silently ignored. Check the result instead of trusting the request.

python
"""tool_choice: auto, required, a forced function, none. What is sent, and how to check it was honored."""
from __future__ import annotations

from examples.m06_tools import SYSTEM, TICKET, TOOLS, READ_ONLY, ToolContext, call, run_tool_loop
from examples.m06_wire import capture_wire
from supportdesk.llm import ChatResult, chat
from supportdesk.stand_in import ScriptedLLM

tools = [TOOLS[name].schema() for name in sorted(READ_ONLY)]
messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": TICKET}]
CHOICES = {
    "auto": "auto",
    "required": "required",
    "forced": {"type": "function", "function": {"name": "get_invoice"}},
    "none": "none",
}
with capture_wire({"content": "ok"}) as bodies:
    for choice in CHOICES.values():
        chat(messages, provider="groq", tools=tools, tool_choice=choice)
for label, body in zip(CHOICES, bodies):
    print(f"{label:<9} tool_choice sent: {body['tool_choice']}")


def honored(result: ChatResult, choice) -> bool:
    """Did the reply obey tool_choice? Some servers accept the field and ignore it."""
    if choice == "none":
        return not result.tool_calls
    if choice == "required":
        return bool(result.tool_calls)
    if isinstance(choice, dict):
        return [c.name for c in result.tool_calls] == [choice["function"]["name"]]
    return True


print("\nA server that ignores tool_choice (ScriptedLLM playing that role):")
ignoring = ScriptedLLM(replies=[ChatResult("I think you were charged twice. Let me know!")])
reply = ignoring(messages, tools=tools, tool_choice=CHOICES["forced"])
print("  forced get_invoice, got text:", repr(reply.text[:40]), "| honored:", honored(reply, CHOICES["forced"]))

print("\nThe loop's last turn sends tool_choice='none', so a tool-happy model must answer:")


def tool_happy(msgs, kwargs):
    """Calls search_kb forever unless tool_choice is 'none'. Plumbing, not a model."""
    if kwargs.get("tool_choice") == "none":
        return ChatResult("Here is what I found so far: duplicate charges are refunded in 5 to 10 business days.")
    n = sum(1 for m in msgs if m["role"] == "tool")
    return ChatResult("", tool_calls=[call(f"loop_{n}", "search_kb", query="refund duplicate charge")])


llm = ScriptedLLM(responder=tool_happy)
result = run_tool_loop(llm, messages, ToolContext("cus_acme", READ_ONLY), max_turns=4)
print("  tool_choice per model call:", [c["tool_choice"] for c in llm.calls])
print("  stop:", result.stop_reason, "after", result.turns, "calls | final:", result.text[:60])

Code explained

  • In simple words: see the four settings on the wire, check whether a reply obeyed its setting, and watch the loop use "none" to force a final answer.
  • What happens: the first block captures four real request bodies with capture_wire. honored(result, choice) checks a reply against its tool_choice. A ScriptedLLM then plays a server that ignores a forced choice. Finally tool_happy is a responder that calls search_kb forever unless it receives tool_choice="none"; run_tool_loop with max_turns=4 gives it three free turns and then forces an answer.
  • Comes out:
text
auto      tool_choice sent: auto
required  tool_choice sent: required
forced    tool_choice sent: {'type': 'function', 'function': {'name': 'get_invoice'}}
none      tool_choice sent: none

A server that ignores tool_choice (ScriptedLLM playing that role):
  forced get_invoice, got text: 'I think you were charged twice. Let me k' | honored: False

The loop's last turn sends tool_choice='none', so a tool-happy model must answer:
  tool_choice per model call: ['auto', 'auto', 'auto', 'none']
  stop: answered after 4 calls | final: Here is what I found so far: duplicate charges are refunded

The loop sent auto three times and none on the last call, so it ended with an answer instead of a fourth tool call. Against a server that ignores tool_choice, the loop returns stop_reason="turn_limit" without running the extra calls (a test covers it), and your code should hand the ticket to a human.

SituationUse thisWhy
Normal assistant turns"auto"The model decides whether a lookup is needed
You know a lookup must happen first (an invoice number is in the ticket)Forced get_invoice on the first turn onlyGuarantees the fact is fetched; forcing it on later turns loops forever
Extraction dressed as a tool call (the arguments are the structured output)Forced single toolUseful on Groq, where structured outputs cannot be combined with tools
Last turn of a loop, or a "summarize now" step"none"Guarantees a text answer and bounds the loop
Provider that does not support tool_choice (Ollama per its docs)"auto" plus a local honored() check and a turn limitThe request field is not a guarantee

Too many tools: selection degradation and grouping strategies

Every tool you offer is sent on every call of every turn, and every tool is another option the model can pick by mistake. Measure the cost side first.

python
"""Too many tools: what 30 tool definitions cost, and two ways to offer fewer."""
from __future__ import annotations

import json
import math

from examples.m06_tools import TOOLS
from supportdesk.data import Article
from supportdesk.kb_search import KBSearch
from supportdesk.pricing import PRICES
from supportdesk.tokens import count_tokens

# 27 more tools a growing support desk might add (definitions only; none of them run here).
EXTRA = [
    ("billing", "list_invoices", "List the current customer's invoices, newest first.", {"limit": "integer"}),
    ("billing", "resend_invoice", "Email a copy of one invoice to the billing contact.", {"invoice_id": "string"}),
    ("billing", "update_tax_details", "Change the VAT number or company name on future invoices.", {"vat_number": "string", "company_name": "string"}),
    ("billing", "change_plan", "Move the workspace to another plan (Free, Team, Business).", {"plan": "string"}),
    ("billing", "get_subscription", "Show plan, seats, billing period, and renewal date.", {}),
    ("billing", "update_payment_method", "Send the customer a secure link to change their card.", {}),
    ("billing", "apply_discount", "Apply a discount code to the next invoice.", {"code": "string"}),
    ("cancellation", "cancel_subscription", "Cancel the plan at the end of the current billing period.", {"reason": "string"}),
    ("cancellation", "pause_subscription", "Pause billing for up to 3 months.", {"months": "integer"}),
    ("account_access", "send_password_reset", "Email a password reset link to a workspace member.", {"email": "string"}),
    ("account_access", "unlock_account", "Unlock an account locked after failed logins.", {"email": "string"}),
    ("account_access", "get_sso_config", "Show the workspace's SSO provider and enforcement settings.", {}),
    ("account_access", "reset_2fa", "Remove two-factor authentication so the member can enroll again.", {"email": "string"}),
    ("account_access", "list_members", "List workspace members with roles.", {"role": "string"}),
    ("account_access", "change_member_role", "Make a member an admin, member, or guest.", {"email": "string", "role": "string"}),
    ("bug", "get_status", "Current Brightlane incidents and component status.", {}),
    ("bug", "create_bug_report", "File a bug for engineering with steps to reproduce.", {"title": "string", "steps": "string"}),
    ("bug", "get_error_logs", "Recent client errors for this workspace.", {"hours": "integer"}),
    ("bug", "get_mobile_app_version", "Latest iOS and Android app versions and known issues.", {}),
    ("how_to", "get_article", "Return the full text of one help-center article.", {"article_id": "string"}),
    ("how_to", "start_export", "Start a CSV or JSON export of a board.", {"board_id": "string", "format": "string"}),
    ("how_to", "get_export_status", "Check whether an export has finished and get its link.", {"export_id": "string"}),
    ("how_to", "list_integrations", "List connected integrations such as Slack.", {}),
    ("how_to", "reconnect_slack", "Send the admin a link to reconnect the Slack integration.", {}),
    ("feature_request", "log_feature_request", "Record a feature request with the customer's use case.", {"summary": "string"}),
    ("feature_request", "find_similar_requests", "Search existing feature requests for duplicates.", {"query": "string"}),
    ("how_to", "delete_workspace_data", "Start a GDPR data deletion request for the workspace.", {"confirm": "boolean"}),
]
GROUP = {"search_kb": "how_to", "get_invoice": "billing", "issue_refund": "billing"}


def schema(name: str, description: str, params: dict[str, str]) -> dict:
    return {"type": "function", "function": {"name": name, "description": description, "parameters": {
        "type": "object", "properties": {p: {"type": t} for p, t in params.items()},
        "required": list(params), "additionalProperties": False}}}


ALL = [t.schema() for t in TOOLS.values()] + [schema(n, d, p) for _, n, d, p in EXTRA]
GROUP.update({name: group for group, name, _, _ in EXTRA})


def wilson(hits: int, n: int, z: float = 1.96) -> tuple[float, float]:
    """95% Wilson score interval for a proportion; honest even near 0% or 100%."""
    p = hits / n
    centre = (p + z * z / (2 * n)) / (1 + z * z / n)
    half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
    return centre - half, centre + half


def tool_tokens(tools: list[dict]) -> int:
    return count_tokens(json.dumps(tools), "o200k_base")


three, thirty = tool_tokens(ALL[:3]), tool_tokens(ALL)
price = PRICES["openai/gpt-oss-120b"].input
print(f"3 tools: {three} tokens   30 tools: {thirty} tokens   (o200k_base, JSON as sent)")
print(f"extra per call: {thirty - three} tokens; per 10,000 calls at {price} USD/M input: "
      f"{(thirty - three) * 10_000 * price / 1e6:.2f} USD, and every turn of every loop pays it again")

# Strategy 1: offer the tools for the ticket's triaged category, plus search_kb.
by_group = {}
for tool in ALL:
    by_group.setdefault(GROUP[tool["function"]["name"]], []).append(tool)
print("\nStrategy 1, by category (search_kb always included):")
for group, members in sorted(by_group.items()):
    offered = members + ([] if group == "how_to" else [ALL[0]])
    print(f"  {group:<16} {len(offered):>2} tools {tool_tokens(offered):>4} tokens")

# Strategy 2: retrieve tools by their descriptions with the course's BM25 search.
tool_index = KBSearch([Article(t["function"]["name"], t["function"]["name"].replace("_", " "), (), t["function"]["description"])
                       for t in ALL])
QUERIES = {  # request -> the tool a correct agent needs
    "I was charged twice, refund the duplicate": "issue_refund",
    "What was the amount on invoice INV-2026-004512?": "get_invoice",
    "Please email me a copy of my invoice": "resend_invoice",
    "Add our VAT number to invoices": "update_tax_details",
    "Upgrade us to Business": "change_plan",
    "When does our subscription renew?": "get_subscription",
    "My card expired, how do I change it": "update_payment_method",
    "Cancel our plan please": "cancel_subscription",
    "Can we pause billing for the summer": "pause_subscription",
    "I forgot my password": "send_password_reset",
    "My account is locked after too many attempts": "unlock_account",
    "Lost my phone, need to reset two-factor": "reset_2fa",
    "Make Ana an admin": "change_member_role",
    "Is Brightlane down right now?": "get_status",
    "Boards crash when I drag a card, here are the steps": "create_bug_report",
    "Export my board to CSV": "start_export",
    "Is my export ready yet?": "get_export_status",
    "Slack notifications stopped, reconnect it": "reconnect_slack",
    "Please add Gantt charts": "log_feature_request",
    "Delete all our data under GDPR": "delete_workspace_data",
}
print(f"\nStrategy 2, BM25 over tool descriptions (n={len(QUERIES)} hand-written requests):")
for k in (1, 3, 5):
    hits = sum(gold in [h.article_id for h in tool_index.search(q, k=k)] for q, gold in QUERIES.items())
    low, high = wilson(hits, len(QUERIES))
    print(f"  gold tool in top {k}: {hits}/{len(QUERIES)} = {hits / len(QUERIES):.0%} (95% Wilson CI {low:.0%} to {high:.0%})")
misses = [(q, gold, [h.article_id for h in tool_index.search(q, k=3)]) for q, gold in QUERIES.items()
          if gold not in [h.article_id for h in tool_index.search(q, k=3)]]
for q, gold, got in misses:
    print(f"  miss: {q!r} needs {gold}, got {got}")

Code explained

  • In simple words: define 27 more plausible Brightlane tools, count what 30 definitions cost, then try two ways of offering only the relevant few.
  • What happens: EXTRA lists 27 tool definitions (group, name, description, parameters) without implementations. ALL is the 3 real tools plus those 27. tool_tokens counts the JSON as sent. Strategy 1 groups tools by triage category, so the router's output from Part A picks the tool set, with search_kb always included. Strategy 2 indexes tool descriptions with the course's BM25 KBSearch (each tool becomes an Article) and retrieves the top k tools for a request; 20 hand-written requests, each with the tool it needs, measure whether the right tool is retrieved. wilson gives a 95% interval for a proportion that stays sensible near 0% and 100%.
  • Comes out:

The accuracy cost of many tools needs a real model. Gan and Sun's RAG-MCP (2025, arXiv:2505.03275) retrieved tool descriptions before each call in a stress test with many tools and report tool-selection accuracy of 43.13% versus 13.62% for offering everything, with prompt tokens cut by more than half. Treat that as evidence of the direction, not a number for your model.

SituationUse thisWhy
Under about 10 well-separated toolsOffer them allSimple; the token cost is modest
Tools cluster by task and you already triageOffer the group for the triaged categoryDeterministic, testable, reuses Part A
Dozens of tools, open-ended requestsRetrieve the top k tool descriptions per request, with a fallback setScales; measure recall on real requests
Several overlapping toolsMerge them into one tool with an enum parameterFewer, clearer choices
A large catalog behind a few entry pointsTwo levels: a choose_toolset step, then that setKeeps each prompt small; costs an extra call