CourseLarge Language Models · Module 14: Applications, Patterns, and Frontiers · part 74 of 80
Part 74 · Module 14: Applications, Patterns, and Frontiers

Part D: Frontier

30 min read·22 Sept 2026

The models you build on will change several times during the life of the Brightlane assistant. This part is about reading that change clearly: which trends are measured, what they mean for a support desk, and how to tell capability from marketing. Every trend number below comes from a cited source; follow the links, because these figures are revised often.

D1. Longer contexts and cheaper inference, and what they unlock

Two trends are well documented by Epoch AI, an independent research group that tracks AI progress:

  • Context windows. In "LLMs now accept longer inputs, and the best models can use them more effectively" (Burnham and Adamczewski, 25 June 2025, epoch.ai/data-insights/context-windows), the longest context windows grew about 30x per year since mid-2023, and on two long-context benchmarks (Fiction.liveBench and MRCR) the input length at which top models reach 80% accuracy rose by over 250x in nine months. Advertised length and usable length both grew, but they are not the same number (Module 2's effective context). For a current example, Google's developer documentation lists a 1 million token input limit for gemini-3.5-flash (ai.google.dev).
  • Price. In "LLM inference prices have fallen rapidly but unequally across tasks" (Cottier, Snodin, Owen, and Adamczewski, 12 March 2025, epoch.ai/data-insights/llm-inference-price-trends), the price of reaching GPT-4's level on PhD-level science questions fell about 40x per year, with the rate ranging from 9x to 900x per year depending on the task and threshold. The authors note that the fastest drops happened in the most recent year, so it is less clear they will persist. Read the unit carefully: this is the price of a fixed level of capability, not the price of the newest frontier model, which has not fallen at that rate.

What does that mean for Brightlane? The next script turns the trends into Brightlane arithmetic and also prices one thing longer contexts unlock: skipping retrieval entirely.

python
# examples/m14_frontier_math.py
"""Frontier: turning published trend numbers and a release announcement into Brightlane arithmetic.

Inputs are cited in the module text: Epoch AI price trends (9x to 900x per year
for a FIXED capability level), METR time-horizon doubling times, and the Gemini
3.8 Flash announcement (2 September 2026) with its introductory price.
"""
from __future__ import annotations

import math

from supportdesk.data import load_articles
from supportdesk.llm import Usage
from supportdesk.pricing import PRICES, Price, cost_usd
from supportdesk.tokens import count_tokens


def cost(price: Price, usage: Usage) -> float:
    """Same formula as supportdesk.pricing.cost_usd, for a price not (yet) in the PRICES table."""
    return (usage.input_tokens * price.input + usage.output_tokens * price.output) / 1_000_000


if __name__ == "__main__":
    kb_tokens = sum(count_tokens(a.title + "\n" + a.body) for a in load_articles())
    print(f"whole help center: {kb_tokens} tokens; per 100k tickets with the whole help center in every prompt:")
    for model in ("openai/gpt-oss-120b", "gemini-3.5-flash"):
        plain = cost_usd(Usage(input_tokens=kb_tokens + 150, output_tokens=350), model) * 1e5
        cached = cost_usd(Usage(input_tokens=kb_tokens + 150, cached_tokens=kb_tokens, output_tokens=350), model) * 1e5
        print(f"  {model:22} {plain:8.2f} USD, with the help center cached {cached:8.2f} USD")

    monthly_now = 194.85  # gemini-3.5-flash triage bill for 100k tickets, from m14_triage_at_scale.py
    print("same-capability price decline applied to a 194.85 USD/month bill")
    for per_year in (9, 40, 900):
        print(f"  {per_year:>3}x per year: after 12 months {monthly_now / per_year:8.2f} USD, "
              f"after 6 months {monthly_now / math.sqrt(per_year):8.2f} USD")

    print("\nMETR 50% time horizon: months until a horizon grows from 14.5 h to a 40 h work week")
    for label, days in (("196-day doubling (2019 to 2025 trend)", 196), ("89-day doubling (since 2024, TH1.1)", 89)):
        doublings = math.log2(40 / 14.5)
        print(f"  {label}: {doublings:.2f} doublings = {doublings * days / 30.4:.1f} months "
              f"(a 2x measurement error moves this by {days / 30.4:.1f} months)")

    print("\nper-task cost of one triage call: 141 input tokens, 193 output tokens (incl. 150 assumed reasoning)")
    old = PRICES["gemini-3.5-flash"]
    intro, later = Price(0.75, 3.75, 0.0), Price(1.50, 7.50, 0.0)   # Gemini 3.8 Flash, from the announcement
    base = Usage(input_tokens=141, output_tokens=193)
    print(f"  gemini-3.5-flash         {cost(old, base) * 1e5:7.2f} USD per 100k")
    for name, price in (("3.8 Flash intro price", intro), ("3.8 Flash from Jan 2027", later)):
        # How many times more output tokens can the new model spend before it costs more per task?
        ratio = (cost(old, base) - 141 * price.input / 1e6) / (193 * price.output / 1e6)
        print(f"  {name:24} {cost(price, base) * 1e5:7.2f} USD per 100k at equal tokens; "
              f"breaks even at {ratio:.2f}x the output tokens")

Code explained

  • In simple words: a calculator that converts headline trend numbers and a price sheet into this project's monthly bill and planning dates.
  • What happens:
    1. It counts the whole help center in tokens and prices putting all of it in every prompt, with and without prompt caching of that stable block (Module 7's cache-aware ordering).
    2. It applies the 9x, 40x, and 900x per year declines to A1's gemini-3.5-flash triage bill. Six months is the square root of a yearly factor.
    3. It converts METR doubling times (D3) into months, and shows how much a factor-of-2 measurement error shifts the answer (exactly one doubling time).
    4. It prices the A1 triage call on Gemini 3.8 Flash at its introductory and later prices (D5) and solves for the output-token multiple at which the new model stops being cheaper per task.
  • Comes out:

    text
    whole help center: 1205 tokens; per 100k tickets with the whole help center in every prompt:
      openai/gpt-oss-120b       41.32 USD, with the help center cached    41.32 USD
      gemini-3.5-flash         518.25 USD, with the help center cached   355.57 USD
    same-capability price decline applied to a 194.85 USD/month bill
        9x per year: after 12 months    21.65 USD, after 6 months    64.95 USD
       40x per year: after 12 months     4.87 USD, after 6 months    30.81 USD
      900x per year: after 12 months     0.22 USD, after 6 months     6.50 USD
    
    METR 50% time horizon: months until a horizon grows from 14.5 h to a 40 h work week
      196-day doubling (2019 to 2025 trend): 1.46 doublings = 9.4 months (a 2x measurement error moves this by 6.4 months)
      89-day doubling (since 2024, TH1.1): 1.46 doublings = 4.3 months (a 2x measurement error moves this by 2.9 months)
    
    per-task cost of one triage call: 141 input tokens, 193 output tokens (incl. 150 assumed reasoning)
      gemini-3.5-flash          194.85 USD per 100k
      3.8 Flash intro price      82.95 USD per 100k at equal tokens; breaks even at 2.55x the output tokens
      3.8 Flash from Jan 2027   165.90 USD per 100k at equal tokens; breaks even at 1.20x the output tokens
    

    The first lines are the practical unlock. Brightlane's entire help center is 1,205 tokens. Putting all of it in every prompt costs about 41 USD per 100,000 tickets on gpt-oss-120b, and it removes the retrieval misses you measured in A3 (and Module 7's 52 of 62 top-1 hit rate) at a stroke. For a 12-article help center, long context beats retrieval; for 12,000 articles it does not, and retrieval comes back. The Gemini line shows why prompt caching matters for this layout: the cached help center cuts the bill by about a third. The trend lines say the same triage capability that costs 195 USD a month today would cost somewhere between 22 USD and 22 cents a year from now if those trends hold for this task, a range so wide that the only safe plan is to re-measure prices quarterly, not to forecast them.

SituationUse thisWhy
Knowledge base of a few thousand tokensPut it all in the prompt, cachedNo retrieval misses; cost is small and falling
Knowledge base far beyond the effective contextRetrieval (Module 7)Cost and effective-context limits still bind
Medium size, frequent updatesRetrieval with a generous k, measured against the whole-corpus baselineKeep whichever wins on your eval, not on the trend chart

D2. Test-time compute scaling

Test-time compute is computation spent while answering, not while training: longer hidden reasoning, several samples with a vote, or a search guided by a verifier (Modules 3 and 5). OpenAI's o1 announcement ("Learning to reason with LLMs", 12 September 2024, openai.com) reported that performance "consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute)." Snell, Lee, Xu, and Kumar ("Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters", August 2024, arXiv:2408.03314) found that allocating test-time compute adaptively per prompt improved efficiency by more than 4x over a best-of-N baseline, and that on problems where a smaller model already had some success, extra test-time compute could outperform a 14x larger model in a compute-matched comparison.

For an application builder, three consequences follow:

  • Price per task, not per token. Reasoning tokens bill as output (Module 2). A model with a lower price per token can cost more per ticket if it thinks longer. The A1 cost table only held because we wrote the reasoning-token assumption down.
  • Effort is a dial you tune per route. Triage of a clear ticket needs little thought; a multi-step refund investigation may need more. Measure accuracy and cost at each reasoning_effort level on your eval set and pick per route (Module 3).
  • Verifiers multiply the value of compute. Extra samples help most when something can check them, which is why the SQL pattern in A4 pairs so well with more attempts.

D3. Agentic capability trajectories

The best-known measurement of agent progress comes from METR, a nonprofit that evaluates frontier models. Their 50% time horizon is the length of task (measured by how long it takes skilled humans) that a model completes successfully half the time, on a suite of mostly software and research tasks.

  • The original paper ("Measuring AI Ability to Complete Long Software Tasks", 19 March 2025, metr.org) found the horizon doubling about every 7 months over six years, with Claude 3.7 Sonnet at about one hour.
  • Time Horizon 1.1 (29 January 2026, metr.org) expanded the suite from 170 to 228 tasks and reported a doubling time of 196 days over 2019 to 2025, 131 days since 2023, and 89 days since 2024. Its top estimates were Claude Opus 4.5 at 320 minutes (95% interval 170 to 729) and GPT-5 at 214 minutes.
  • On 20 February 2026 METR estimated Claude Opus 4.6 at about 14.5 hours (95% interval 6 to 98 hours), adding that the measurement is "extremely noisy because our current task suite is nearly saturated" (METR on X). METR's live page (metr.org/time-horizons, last updated 8 May 2026 when this was written) has since added more models; check it for current figures rather than trusting any number printed here.

METR is also explicit about what the metric is not ("Clarifying limitations of time horizon", 22 January 2026, metr.org): it is not how long a model can work unattended, the error bars are roughly a factor of 2 in each direction, horizons differ between domains by orders of magnitude (visual computer-use tasks are 40x to 100x lower), and the task suite is mostly software.

So what should Brightlane do with this? The frontier math output above shows that a factor-of-2 measurement error moves any extrapolation by a full doubling time, 3 to 6 months. The durable lessons are about design, not dates. A 50% success rate is a research metric; Part C showed a customer-facing action needs 84% to 99.6% depending on the cost of an error, and tasks at the edge of a model's horizon are exactly where it fails most. Agents will keep taking on longer tasks, so design the approval gates from B1 around the consequence of each action, and they will stay correct however capable the model becomes.

.

D4. Small models and on-device inference

At the other end of the scale, small models run on phones and laptops with no network, no per-token price, and no data leaving the device. Apple's documentation for its on-device foundation model describes "a compact, approximately 3-billion-parameter model" compressed to 2 bits per weight, and states it "is not designed to be a chatbot for general world knowledge", while it does well at summarization, extraction, text understanding, and short dialog (Apple Machine Learning Research). That list is the pattern catalogue from Part A with the knowledge-heavy patterns removed.

TinyLM, at 1.07 million parameters, is the extreme small case: about 3,000 times smaller than that on-device model. Here is its real CPU latency.

python
# examples/m14_tinylm_latency.py
"""Frontier: small models and on-device inference, with TinyLM as the extreme small case.

Measures real CPU latency for a 1.07M-parameter model: prefill time, decode time
per token and tokens per second (wall clock and CPU time), plus memory footprint.
"""
from __future__ import annotations

import statistics
import time

import torch

from supportdesk import tinylm

PROMPT = "Customer: I was charged twice for the Team plan this month.\nAgent:"


def measure(model, tokenizer, threads: int, runs: int = 9) -> tuple[float, float, float]:
    torch.set_num_threads(threads)
    params = tinylm.SamplingParams(max_new_tokens=40, temperature=0)
    walls, cpus, prefills = [], [], []
    for i in range(runs):
        cpu0, wall0 = time.process_time(), time.perf_counter()
        gen = tinylm.generate(model, tokenizer, PROMPT, params)
        cpu, wall = time.process_time() - cpu0, time.perf_counter() - wall0
        if i >= 2:  # the first two runs warm up; keep the rest
            walls.append(wall * 1000 / len(gen.token_ids))
            cpus.append(cpu * 1000 / len(gen.token_ids))
            prefills.append(gen.prefill_ms)
    return statistics.median(prefills), statistics.median(walls), statistics.median(cpus)


if __name__ == "__main__":
    model, tokenizer = tinylm.load()
    n = model.num_parameters()
    print(f"parameters {n:,}; fp32 weights {n * 4 / 1e6:.1f} MB; at 2 bits per weight {n * 2 / 8 / 1e6:.2f} MB")
    print(f"prompt: {len(tokenizer.encode(PROMPT).ids)} tokens, generating 40")
    with open("/proc/loadavg") as f:
        print("machine load average (1 min):", f.read().split()[0])
    for threads in (1, 2):
        prefill, wall, cpu = measure(model, tokenizer, threads)
        print(f"threads={threads}: prefill {prefill:.1f} ms, whole generation {wall:.1f} ms/token wall "
              f"({1000 / wall:.0f} tokens/s), {cpu:.1f} ms/token CPU")
    torch.set_num_threads(1)
    gen = tinylm.generate(model, tokenizer, PROMPT, tinylm.SamplingParams(max_new_tokens=40, temperature=0))
    print("output:", repr(gen.text[:110]))
    print("capital of France:", repr(tinylm.generate(model, tokenizer, "The capital of France is",
                                                     tinylm.SamplingParams(max_new_tokens=16, temperature=0)).text))

Code explained

  • In simple words: time a very small model on this machine's CPU and look at what it produces.
  • What happens:
    1. It computes the weight size at 32-bit floats and at 2 bits per weight, the compression level Apple reports.
    2. measure runs greedy generation of 40 tokens nine times, discards two warm-up runs, and records the median prefill time plus wall-clock and CPU time per generated token. CPU time (time.process_time) counts only time this process actually computed, so it is less sensitive to other programs on the machine.
    3. It measures once with 1 thread and once with 2, and prints the machine's load average so the numbers can be read in context.
    4. It prints a real continuation of a support prompt and of "The capital of France is".
  • Comes out: this run shared a 2-core machine with other heavy jobs (load average about 17), so treat the wall-clock numbers as a snapshot.

    text
    parameters 1,071,872; fp32 weights 4.3 MB; at 2 bits per weight 0.27 MB
    prompt: 16 tokens, generating 40
    machine load average (1 min): 17.07
    threads=1: prefill 1.6 ms, whole generation 4.5 ms/token wall (220 tokens/s), 1.0 ms/token CPU
    threads=2: prefill 896.1 ms, whole generation 131.0 ms/token wall (8 tokens/s), 28.6 ms/token CPU
    output: ' Team plan this month. Can you refund the duplicate?\nAgent (Eli): Sorry about the duplicate charge. Duplicate '
    capital of France: ' not loading. Is there an outage?\nAgent (Nia): Please check status'
    

    With one thread, TinyLM generates at about 220 tokens a second of wall time and needs about 1 ms of CPU per token. The weights are 4.3 MB, or about a quarter of a megabyte at 2 bits per weight. With two threads on the busy machine it collapsed to 8 tokens a second, because the two threads kept waiting for each other and for other processes. That is a real on-device lesson: a phone or laptop runs your model alongside everything else, so measure under realistic load and prefer fewer threads when the CPU is contended. The outputs are the other lesson. The support continuation is fluent because TinyLM memorized its templated corpus; the France prompt produces confident support-desk text, which is Module 1's hallucination demo in one line. A small model is fast and private, and it knows only what its training data taught it.

SituationUse thisWhy
Privacy-sensitive or offline, narrow task (classify, extract, short rewrite)A small on-device model, evaluated on your taskNo data leaves the device; latency is local
Tasks that need world knowledge or long reasoningA hosted frontier or mid-tier modelSmall models do not hold the knowledge
High volume, narrow task, a big model works but costs too muchDistill or fine-tune a small model (Module 9)Keeps quality on the narrow task at a fraction of the cost

D5. Reading releases critically: separating capability from marketing

Model announcements are written to sell. That does not make them false, but it means the questions you care about are often not the ones the post answers. Apply the same checklist to every release. Here it is applied to a real one: Google's "Introducing Gemini 3.8 Flash and 3.8 Flash Cyber" (Tulsee Doshi and Raluca Ada Popa, 2 September 2026, blog.google), as read on 21 September 2026.

QuestionWhat the announcement saysWhat it means for Brightlane
Which benchmarks, and are they like my task?54.9% on HLE-Verified; claims on DeepSWE v1.1 (long-horizon software engineering), Vals Finance Agent V2, and Harvey's Legal Agent BenchmarkNone is support triage or grounded answering. Expect nothing until our own eval runs
Compared with what, under which settings?"Outperforms most larger frontier models" on DeepSWE; "outperforms 3.7 Flash and other frontier models" on two agent benchmarks"Most" and unnamed competitors; the post links no methodology page, effort level, or number of attempts
What does it cost per task, not per token?0.75 USD input and 3.75 USD output per million tokens, an introductory rate through 31 December 2026, then 1.50 and 7.50The frontier math shows the intro price breaks even with 3.5 Flash at 2.55x the output tokens, but the later price at only 1.20x
Does it change how many tokens it uses?It "might use more tokens to maximize performance, especially at higher effort levels"Per-token price is not per-task price. Measure tokens per ticket on our eval
Context, latency, availability?No context window or latency figures in the post; available in the Gemini API via Google AI Studio, Antigravity, Android Studio, Gemini Enterprise, the Gemini app, and moreLook up the model card and measure latency ourselves
Who measured it?The vendorWait for independent measurements (for example Artificial Analysis, METR, Epoch AI) and, above all, run our own

Only the last question has a universal answer, and it is the one you control. Rerun the Module 10 eval suite with the candidate model (D6 automates the decision), measure tokens and latency per ticket, and let those numbers decide. A benchmark gain on legal agents says nothing about Maya's refund tickets.

D6. Keeping a system adaptable as models change under it

Models are deprecated, repriced, and replaced (Module 13's forced migrations). A system stays adaptable when three things are true, all of which this course has already built:

  • Prompts are versioned data with a content hash, not strings scattered through code (Modules 4 and 5).
  • Each feature names its provider, model, and prompt version in one place, behind the provider-neutral supportdesk.llm.chat (Modules 1 and 13).
  • No change ships without the same eval gate: golden set, invalid-output count, cost and latency budgets (Module 10).

python
# examples/m14_adaptable.py
"""Frontier: keeping a system adaptable as models change underneath it.

Three habits from earlier modules, in one small file: prompts are versioned
data with a content hash (Module 4 and 5), each feature names its model in one
config (Module 13), and no model or prompt change ships without passing the
same golden-set gate (Module 10).
"""
from __future__ import annotations

import hashlib
import json
from dataclasses import dataclass

from pydantic import ValidationError

from examples.m14_common import fmt_rate
from supportdesk.data import load_tickets
from supportdesk.llm import ChatResult, Usage
from supportdesk.pricing import PRICES, cost_usd
from supportdesk.schemas import Triage
from supportdesk.stand_in import ScriptedLLM

PROMPTS = {
    ("triage", "v3"): "Classify the ticket. Reply with only JSON: category, priority, language, summary, needs_human.",
    ("triage", "v4"): "Classify the Brightlane support ticket. Reply with only a JSON object with keys category, "
                      "priority, language, summary, needs_human. No prose, no code fences.",
}


def prompt_hash(feature: str, version: str) -> str:
    return hashlib.sha256(PROMPTS[(feature, version)].encode()).hexdigest()[:8]


@dataclass(frozen=True)
class FeatureConfig:
    feature: str
    provider: str
    model: str
    prompt_version: str
    max_cost_per_1k_usd: float
    max_p95_latency_ms: float


PRODUCTION = FeatureConfig("triage", "groq", "openai/gpt-oss-120b", "v3", 0.50, 3000)


def run_gate(config: FeatureConfig, llm, tickets) -> dict:
    correct = invalid = 0
    priced_as = config.model if config.model in PRICES else "openai/gpt-oss-120b"  # add new models to PRICES first
    latencies, cost = [], 0.0
    for t in tickets:
        messages = [{"role": "system", "content": PROMPTS[(config.feature, config.prompt_version)]},
                    {"role": "user", "content": t.text}]
        result = llm(messages, provider=config.provider, model=config.model, temperature=0.0)
        latencies.append(result.latency_ms)
        cost += cost_usd(result.usage, priced_as)
        try:
            correct += Triage.model_validate_json(result.text).category == t.gold["category"]
        except ValidationError:
            invalid += 1
    p95 = sorted(latencies)[int(0.95 * (len(latencies) - 1))]
    return {"correct": correct, "n": len(tickets), "invalid": invalid, "p95_ms": p95,
            "cost_per_1k": cost / len(tickets) * 1000}


def decide(baseline: dict, candidate: dict, config: FeatureConfig, noise: int = 2) -> str:
    """Promote only if not worse beyond noise, no new invalid outputs, and within cost and latency budgets."""
    if candidate["invalid"] > baseline["invalid"]:
        return f"HOLD: {candidate['invalid']} invalid outputs vs {baseline['invalid']}"
    if candidate["correct"] < baseline["correct"] - noise:
        return f"HOLD: accuracy dropped beyond the noise allowance of {noise} tickets"
    if candidate["cost_per_1k"] > config.max_cost_per_1k_usd or candidate["p95_ms"] > config.max_p95_latency_ms:
        return "HOLD: over cost or latency budget"
    return "PROMOTE"


def fake_model(fence_every: int, wrong_every: int, latency_ms: float):
    """STAND-IN 'models' with known behaviour (not real models): some wrap JSON in fences, some mislabel."""
    counter = {"i": 0}

    def responder(messages, kwargs):
        counter["i"] += 1
        t = next(t for t in TICKETS if t.text == messages[-1]["content"])
        category = "how_to" if counter["i"] % wrong_every == 0 else t.gold["category"]
        body = json.dumps({"category": category, "priority": t.gold["priority"], "language": t.language,
                           "summary": f"About {t.subject}.", "needs_human": False})
        text = f"```json\n{body}\n```" if counter["i"] % fence_every == 0 else body
        return ChatResult(text=text, usage=Usage(input_tokens=120, output_tokens=45), latency_ms=latency_ms)
    return ScriptedLLM(responder=responder)


TICKETS = load_tickets("test")

if __name__ == "__main__":
    for key in PROMPTS:
        print(f"prompt {key[0]}/{key[1]} sha256:{prompt_hash(*key)}")
    baseline = run_gate(PRODUCTION, fake_model(fence_every=10**6, wrong_every=8, latency_ms=900), TICKETS)
    print("baseline  ", PRODUCTION.model, PRODUCTION.prompt_version, fmt_rate(baseline["correct"], baseline["n"]))
    candidates = {
        "new model, old prompt": (FeatureConfig("triage", "groq", "new-model-x", "v3", 0.50, 3000), 6, 12, 700),
        "new model, prompt v4": (FeatureConfig("triage", "groq", "new-model-x", "v4", 0.50, 3000), 10**6, 12, 700),
    }
    for name, (config, fence_every, wrong_every, latency) in candidates.items():
        result = run_gate(config, fake_model(fence_every, wrong_every, latency), TICKETS)
        print(f"{name:22} {fmt_rate(result['correct'], result['n'])}, invalid {result['invalid']}, "
              f"p95 {result['p95_ms']:.0f} ms -> {decide(baseline, result, config)}")

Code explained

  • In simple words: every model or prompt upgrade has to pass the same driving test as the one it replaces, and "drives about as well but crashes sometimes" is a fail.
  • What happens:
    1. PROMPTS stores prompt text by feature and version; prompt_hash gives a short SHA-256 that goes into logs and feedback events (B6), so any output can be traced to the exact prompt text.
    2. FeatureConfig pins provider, model, prompt version, and the cost and latency budgets for one feature.
    3. run_gate sends the 24 test tickets through a configuration, counting correct categories, invalid outputs, p95 latency, and cost per 1,000 tickets via cost_usd. A model missing from PRICES is priced at gpt-oss-120b rates as a placeholder, with a comment telling you to add it first.
    4. decide promotes only if there are no new invalid outputs, accuracy has not dropped by more than a stated noise allowance (2 tickets of 24), and budgets hold.
    5. fake_model builds stand-in "models" with known behaviour, not real models: the new one wraps every sixth reply in a code fence (a common real change between model versions) and mislabels every twelfth.
  • Comes out:

    text
    prompt triage/v3 sha256:031fc023
    prompt triage/v4 sha256:e7835c15
    baseline   openai/gpt-oss-120b v3 21/24 = 87.5% (95% CI 69.0% to 95.7%)
    new model, old prompt  20/24 = 83.3% (95% CI 64.1% to 93.3%), invalid 4, p95 700 ms -> HOLD: 4 invalid outputs vs 0
    new model, prompt v4   22/24 = 91.7% (95% CI 74.2% to 97.7%), invalid 0, p95 700 ms -> PROMOTE
    

    The new model with the old prompt is held because 4 replies came back fenced and failed validation, a regression an accuracy-only comparison would have hidden. With prompt v4, which says "no code fences" explicitly, the same model passes. Read PROMOTE carefully: 22 of 24 versus 21 of 24 is well within noise, so the gate's claim is "not worse and within budget", not "better". That is the right bar for a routine migration.

Module Lab

The lab joins the module's pieces into one intake pipeline for held-out tickets. Its rule is the main lesson of Part C: run the cheapest reliable step first, and call the model only for what the cheap steps cannot do.

.

python
# examples/m14_lab.py
"""Module 14 Lab: the Brightlane intake pipeline, cheapest reliable step first.

For each held-out ticket: regex extracts invoice ids; keyword rules classify
when they match and the LLM classifies only when they do not; the grounded
drafter answers or abstains; the autonomy policy decides who acts; every step
writes a feedback event; the run ends with measured counts and projected cost.
"""
from __future__ import annotations

import json
import re
from collections import Counter
from pathlib import Path

from pydantic import ValidationError

from examples.m14_autonomy import ACTIONS, autonomy_for
from examples.m14_classical_baselines import ROBUST, RULES, extract
from examples.m14_common import fmt_rate, pick_llm
from examples.m14_feedback_events import FeedbackEvent, log
from examples.m14_grounded_answer import answer, choose_threshold
from examples.m14_triage_at_scale import build_messages as triage_messages
from supportdesk.data import load_tickets
from supportdesk.llm import Usage
from supportdesk.pricing import cost_usd
from supportdesk.schemas import Triage
from supportdesk.stand_in import ScriptedLLM

MODEL = "openai/gpt-oss-120b"
POLICY = {a.name: autonomy_for(a) for a in ACTIONS}


def rule_match(text: str) -> str | None:
    low = text.lower()
    return next((category for category, pattern in RULES if re.search(pattern, low)), None)


def responder(messages, kwargs):
    """STAND-IN (not a model): triage always says how_to; drafts cite the first source they were shown."""
    if "You triage support tickets" in messages[0]["content"]:
        return json.dumps({"category": "how_to", "priority": "normal", "language": "en",
                           "summary": "Customer needs help with a feature.", "needs_human": False})
    source = re.search(r'<source id="([^"]+)">', messages[-1]["content"]).group(1)
    return json.dumps({"reply": f"Here is what our help center says ({source}).", "cited_articles": [source],
                       "confidence": "medium"})


class Metered:
    """Wraps any chat function and adds up calls and token usage, whatever the backend."""

    def __init__(self, llm) -> None:
        self.llm, self.calls, self.usage = llm, 0, Usage()

    def __call__(self, messages, **kwargs):
        result = self.llm(messages, **kwargs)
        self.calls += 1
        self.usage.input_tokens += result.usage.input_tokens
        self.usage.output_tokens += result.usage.output_tokens
        return result


if __name__ == "__main__":
    llm = Metered(pick_llm(ScriptedLLM(responder=responder)))
    threshold = choose_threshold(load_tickets("dev"))
    events = Path("m14_out/lab_events.jsonl")
    events.parent.mkdir(exist_ok=True)
    events.unlink(missing_ok=True)
    routes, rule_hits, invoice_count = Counter(), [], 0
    for t in load_tickets("test"):
        invoices = extract(ROBUST, t.text)
        category = rule_match(t.text)
        if category:
            rule_hits.append(category == t.gold["category"])
            routes["category by rule"] += 1
        else:
            result = llm(triage_messages(t), temperature=0.0)
            try:
                category = Triage.model_validate_json(result.text).category
                routes["category by LLM"] += 1
            except ValidationError:
                category, routes["category fallback"] = "how_to", routes["category fallback"] + 1
        draft, why = answer(t, llm, threshold)
        action = "send_reply_with_kb_link" if why == "answered" else "route_to_queue"
        routes[f"{action} ({POLICY[action]})"] += 1
        log(FeedbackEvent(event="draft_shown" if why == "answered" else "escalated", ticket_id=t.id,
                          draft_id=f"lab-{t.id}", feature="draft_reply", prompt_version="lab-v1", model=MODEL,
                          agent="agent_lab", cited_articles=draft.cited_articles), events)
        invoice_count += len(invoices)
    print(json.dumps(dict(routes), indent=1))
    print("rule accuracy where rules fired:", fmt_rate(sum(rule_hits), len(rule_hits)))
    print(f"invoice ids extracted: {invoice_count}; events logged: {sum(1 for _ in events.open())}")
    per_ticket = cost_usd(llm.usage, MODEL) / 24
    print(f"LLM calls: {llm.calls} for 24 tickets ({llm.usage.input_tokens} in, {llm.usage.output_tokens} out tokens)")
    print(f"projected {MODEL} cost: {per_ticket * 100_000:.2f} USD per 100k tickets "
          f"(stand-in token counts; reasoning tokens not included)")

Code explained

  • In simple words: one conveyor belt through every station built in this module, with a meter on the expensive machine.
  • What happens:
    1. It imports the pieces it needs from this module's own examples: the autonomy policy (B1), the robust regex and rules (C2), the feedback schema (B6), the grounded answerer and threshold (A3), and the triage prompt (A1).
    2. rule_match returns a category only when a rule actually fires, instead of falling back to how_to, so the pipeline knows when to escalate to the model.
    3. Metered wraps whichever chat function is in use and counts calls and tokens, so the cost report works the same for the stand-in and for a live model.
    4. For each test ticket: extract invoice ids, classify by rule or by LLM (with a fallback on invalid output), draft or abstain, look up the autonomy level for the resulting action, and log a feedback event.
    5. The stand-in always answers how_to for triage and cites the first source it was shown for drafts. It exists to drive the plumbing; its category answers are not scored.
  • Comes out:

    text
    [stand-in] ScriptedLLM: tests the plumbing only, not model quality
    {
     "category by rule": 11,
     "send_reply_with_kb_link (suggest_and_approve)": 15,
     "category by LLM": 13,
     "route_to_queue (auto)": 9
    }
    rule accuracy where rules fired: 9/11 = 81.8% (95% CI 52.3% to 94.9%)
    invoice ids extracted: 0; events logged: 24
    LLM calls: 28 for 24 tickets (7584 in, 968 out tokens)
    projected openai/gpt-oss-120b cost: 7.16 USD per 100k tickets (stand-in token counts; reasoning tokens not included)
    

    Rules fire on 11 of 24 tickets and are right on 9 of them (82%, interval 52% to 95%). Compare that with 58% for the rules on all tickets in C2: letting rules abstain when they are unsure makes them much more accurate on the tickets they keep, which is the whole idea behind a cascade. The other 13 go to the model. Fifteen tickets get a draft for agent approval and nine are routed straight to the queue, matching A3. The test split contains no invoice ids, so the regex finds none, which is the right answer. The model was called 28 times for 24 tickets (13 triage calls plus 15 drafts), and the projected cost is about 7 USD per 100,000 tickets on gpt-oss-120b. That projection uses the stand-in's short replies and no reasoning tokens, so rerun with M14_LIVE=1 to replace it with a measured figure before quoting it.

Extend the lab in this order, measuring after each step: replace the stand-in with a real model and record its category accuracy on the 13 escalated tickets; train the C2 classifier on a larger labelled set and put it between the rules and the model; and apply the D6 gate before switching the model behind any stage.

Project Milestone

The Brightlane repository (work/m14) now contains:

FileWhat it adds
examples/m14_common.pyLive or stand-in switch, Wilson intervals
examples/m14_triage_at_scale.pyConcurrent triage with dead-lettering and a 100k-per-month cost table
examples/m14_weekly_digest.pyMap-reduce digest with a number checker
examples/m14_grounded_answer.pyRetrieval gate chosen on dev, citation gate, abstention
examples/m14_sql_codegen.pyText-to-SQL with a read-only database and tests as the verifier
examples/m14_review_queue.pyReview queue state machine for article drafts
examples/m14_enrich_pipeline.pyIdempotent enrichment with a content-addressed cache
examples/m14_ticket_search.pyPermission-filtered search over past tickets
examples/m14_autonomy.pyAutonomy policy, undo log, fallbacks
examples/m14_reviewer_fatigue.pyReview placement model, exact and simulated
examples/m14_draft_card.pyAgent card with measured confidence bands, customer disclosure lines
examples/m14_feedback_events.pyFeedback event schema and weekly summary
examples/m14_classical_baselines.pyRegex, rules, and TF-IDF classifier measured against each other
examples/m14_reliability_bar.pyExpected-cost comparison and break-even accuracy
examples/m14_limitations_onepager.pyStakeholder brief generated from measurements
examples/m14_frontier_math.pyTrend and release arithmetic for Brightlane
examples/m14_tinylm_latency.pyTinyLM CPU latency
examples/m14_adaptable.pyPrompt registry and model-upgrade gate
examples/m14_lab.pyThe cascade pipeline
tests/test_m14_patterns.py14 tests pinning the behaviour above

Run the tests with PYTHONPATH=. python -m pytest -q tests/test_m14_patterns.py; in this build all 14 pass in about 9 seconds. The canonical supportdesk/ package is unchanged: everything here imports it.

The assistant can now do more than answer tickets. It can say which of its jobs it should not be doing, what each of its actions is allowed to do alone, what its limits are in numbers a support lead can read, and how it will be checked the next time its model changes.

Interview Questions

1. A product manager wants an LLM to classify 100,000 support tickets a month. What do you do first? Measure the cheap baselines on a labelled sample: majority class, keyword rules, and a classical classifier trained on historical labels, each with a confidence interval on a held-out split. Then price the LLM per ticket from real token counts, including reasoning tokens, and compare batch pricing and caching. In the Brightlane data, rules and TF-IDF are within noise of each other at 48 training tickets, and the right design is usually a cascade: rules or classifier when confident, LLM for the rest, all validated against a schema with a dead-letter path.

2. How do you decide what an AI system may do without human approval? Derive autonomy from the action's consequences, not from how good the model seems: reversibility, customer visibility, whether it moves money, data, or access, and blast radius. Reversible internal actions can be automatic with undo; customer-visible or access-changing actions need approval; irreversible money or data actions stay human-only. Then relax a level only with evidence from approval and edit rates.

3. Your reviewers approve 99% of AI drafts. Is that good news? Not necessarily. A very high approval rate is consistent with good drafts and with fatigued reviewers. Plant known-wrong drafts, measure the catch rate with an interval (you need on the order of 100 planted errors for a usable interval), shorten sessions if catch rates decay, and target review at items a measured signal flags, while auditing a random sample of the rest.

4. When is a regex better than an LLM for extraction? When the field has a fixed format: invoice ids, VAT numbers, order numbers. A tested regex is deterministic, explainable, and runs in hundredths of a millisecond. The work is in the hard cases: case, spacing, neighbours like SINV-, and Unicode word boundaries (a \b after Japanese text fails in Python). Use an LLM only when the format genuinely varies, and even then validate its output with a regex.

5. How do you know whether 90% accuracy is good enough to automate? Price the outcomes. Compute the expected cost per item of each option (manual, draft and review, automatic) from agent time, review time, fix time, catch rate, and the cost of an error that reaches the customer, then find the break-even accuracy. With Brightlane's assumptions, auto-send wins for how-to answers above 84%, but for errors costing 300 USD the bar is 99.6%. Then plan with the lower end of your measured accuracy interval, not the point estimate.

6. How should an assistant decide to say "I don't know"? Use evidence available before generation and after it. Before: a retrieval score threshold chosen on dev data for a target wrong-answer rate, so weak matches never reach the model. After: validation of the output and a citation check that every cited source was actually retrieved. Measure both the unsafe-answer rate and the rate of safe answers you gave up, on held-out data, and report both.

7. What is wrong with filtering search results by permission after ranking? Two things. Any bug in the post-filter leaks restricted content directly. And even a correct post-filter lets restricted documents influence collection-wide statistics such as IDF, so everyone's rankings depend on documents they cannot see; in the Brightlane example a shared ticket's score changed from 3.649 to 3.495 depending on the index. Filter before ranking: per-permission indexes or filtered queries.

8. How do you make an LLM enrichment pipeline safe to rerun? Extract deterministic fields with code, cache model replies under a hash of the exact request including the prompt version, validate every reply, write failures to a dead-letter file, and rewrite outputs rather than appending. Cache only validated replies, or failures replay forever. Changing the prompt version should invalidate the cache on purpose.

9. A vendor announces a model that "outperforms most frontier models" and is half the price. How do you evaluate it? Check whether the benchmarks resemble your task, what the comparisons were against and under which settings, whether the price is introductory, and whether the model uses more tokens per task. Then run your own eval suite through the same gate as any change: accuracy within noise, no new invalid outputs, cost per task and p95 latency within budget. Per-token price is not per-task price.

10. What does METR's time horizon tell you, and what does it not? It estimates the length of task, in skilled-human time, that a model completes with 50% success on a mostly software suite, and it has grown with a doubling time of roughly 3 to 7 months depending on the period. It does not say how long a model can work unattended, its error bars are about a factor of 2 either way, and it varies by orders of magnitude across domains. For product design it means agents will attempt longer tasks, while customer-facing actions still need reliability far above 50%.

11. What feedback should an LLM feature log from day one? Structured events with a fixed schema: draft shown, sent with a measured edit ratio, discarded, escalated, wrong source flagged, undo, each carrying the ticket, prompt version, model, cost, latency, and a pseudonymous user id. That makes "did the new prompt help?" a query with confidence intervals, and it produces labels for future evaluation and fine-tuning.

12. How do you keep an LLM system adaptable as models change? Keep prompts as versioned, hashed data; route every call through a provider-neutral interface with the model named in one configuration per feature; and gate every model or prompt change on the same golden set, invalid-output count, cost, and latency budgets. Expect format regressions (such as new code fences) as often as accuracy changes, and write the promotion rule as "not worse beyond noise", since small test sets cannot show improvement.

Other Tools and Providers

Tool or providerWhat it doesWhen to consider it instead
fastText, SetFitFast text classification; SetFit fine-tunes a small sentence-transformer from few labelsWhen the C2 classifier needs more accuracy but an LLM is too slow or costly
spaCy rule matchers, DucklingRule-based extraction with linguistic features; Duckling parses dates, amounts, and durationsWhen regexes grow unmaintainable but the fields are still well defined
Label Studio, ArgillaHuman review and labelling queuesInstead of the hand-built A5 queue when many reviewers, guidelines, and agreement metrics are involved
Langfuse, Arize Phoenix, LangSmith, BraintrustTracing, feedback capture, and eval dashboardsInstead of the B6 JSONL file once several features and teams log events
Elasticsearch or OpenSearch, Vespa, pgvector, QdrantSearch engines and vector stores with filtered queriesFor A7 when per-group indexes stop scaling
Provider batch APIs (OpenAI, Anthropic, Google, Groq)Asynchronous processing at a discountFor backlog classification and enrichment jobs that can wait
Apple Foundation Models framework, Gemini Nano, llama.cpp, Ollama, MLC LLMOn-device and local inferenceFor D4-style private, offline, narrow tasks
Epoch AI, METR, Artificial Analysis, LMArenaIndependent trend data, agent evaluations, speed and price benchmarks, human preference rankingsFor D5: context before you trust a launch post (still no substitute for your own eval)

Coming Up: Capstone Project

The course ends with a capstone in module15-capstone.md. You pick one of four tracks and build it on the Brightlane assistant you now have:

  • Product track: ship one LLM feature end to end, with prompt versioning, structured outputs, an evaluation suite with CI gates, cost and latency budgets, a safety review, and a documented failure-mode analysis.
  • Agent track: build an agent with real tools, sandboxing, termination guards, trajectory evaluation, human approval gates, and full tracing, and show it behaving correctly under injected adversarial content.
  • Adaptation track: take a task where prompting plateaus, build a fine-tuning dataset, adapt a model, and prove lift over the prompted baseline while measuring general-capability regression.
  • Evaluation track: build a rigorous eval harness for an existing system (assertion tier, calibrated judge tier, noise-floor measurement, CI gates) and use it to find and fix three real quality regressions.

Every track shares four requirements that this module has been practicing: a baseline the final system must beat (Part C), error analysis on at least 50 real failures clustered by cause, cost and latency measured rather than estimated, and a written account of what did not work. The capstone file has the full brief, the milestones, and how to present your results.