CourseLarge Language Models · Capstone Project: Four Tracks · part 75 of 80
Part 75 · Capstone Project: Four Tracks

Part A: Choosing a Track

8 min read·22 Sept 2026

Dark

Capstone Project: Four Tracks

By the end of this capstone, you'll have:

  • Chosen one of four tracks (Product, Agent, Adaptation, Evaluation) for a reason you can defend, and scoped it to about six weeks of part-time work on the Brightlane support assistant.
  • A measured baseline your final system must beat, with confidence intervals, built from keyword triage and BM25 retrieval on the test split.
  • An error-analysis report on at least 50 real failures of your own system, tagged with a cause taxonomy and clustered by cause and by simple text features, with a synthetic set used only to widen the search.
  • Cost and latency measured from real call records (p50, p95, and cost per task), a budget your system is held to, and a written account of what you tried that did not work.
  • A runnable starter for your track, with real output here: a prompt-versioning CI gate, a guarded agent loop that survives an injected help-center article, a TinyLM fine-tune checked against a prompted baseline and a regression budget, or a two-tier evaluation harness with a calibrated judge and a noise floor.
  • An evidence pack and a capstone defense: the numbers, the failures, and the answers to the questions a reviewer will ask.

Prerequisites: Modules 1 to 14. Every track uses Module 2 (tokens and cost), Module 10 (evaluation), and Module 13 (operations and budgets). The Product track leans hardest on Module 4 (prompt engineering), Module 6 (structured outputs), Module 11 (safety review), and Module 14 (application patterns). The Agent track leans on Module 6 (tool calling), Module 7 (retrieval with KBSearch), Module 8 (agents, guards, approval gates), and Module 11 (prompt injection). The Adaptation track leans on Module 1 and Module 3 (TinyLM and scoring by log-probability) and Module 9 (fine-tuning, forgetting, replay). The Evaluation track leans on Module 10 above all, plus Module 5 (judges and self-checks) and Module 13 (CI gates). Working Python.

Where we are: Module 14 ended with a pipeline that runs the cheapest reliable step first and a gate for every model change. The course taught one skill per module. The capstone asks you to put them together on one problem, with the discipline that makes the result believable: a baseline, real failures read one by one, measured cost, and an honest record of dead ends.

How this file is organized

PartWhat it covers
Part A: Choosing a TrackWhat each track proves, how to choose, the six-week shape, and the rules for honest evidence
Part B: Shared RequirementsTooling you build and run for every track: B1 a baseline to beat, B2 error analysis on 50+ failures, B3 cost and latency measured, B4 a written account of what did not work
Part C: Product TrackPrompt versioning, schema-validated triage, an eval and budget CI gate, safety review; starter, milestones, evidence, rubric
Part D: Agent TrackA guarded tool loop with approval gates, tracing, trajectory evaluation, and an injected adversarial article
Part E: Adaptation TrackA TinyLM fine-tune under two minutes, recipe selection without touching test, a regression budget, and a paired comparison with the prompted baseline
Part F: Evaluation TrackAn assertion tier, a judge tier calibrated with Cohen's kappa, a noise floor, and a CI gate that catches regressions
Module LabOne script that runs all four shared requirements and writes the evidence file
Project Milestone, Interview Questions, Other Tools, Course Wrap-UpDeliverables per track, capstone-defense questions, alternatives, and where to go next

Setup for the capstone

The examples run in the order shown, from the root of your own copy of the repository. Every script imports only canonical supportdesk.* modules and other cap_* files in the same examples/ folder, because the capstone is one project.

bash
cp -r supportdesk work/cap && cd work/cap
source /home/claude/venv/bin/activate          # or your own Python 3.11 venv
pip install scikit-learn==1.9.1                # used only by the tests, to cross-check kappa
export PYTHONPATH=.
python -c "import supportdesk.data as d; print(len(d.load_tickets()), 'tickets,', len(d.load_articles()), 'articles')"

Code explained

  • In simple words: make a private workbench, add the one test dependency, and check that the data loads.
  • What happens: the copy keeps capstone experiments away from the canonical code (which you never edit); scikit-learn 1.9.1 is only needed by tests/test_cap_tooling.py, which checks our Cohen's kappa against sklearn.metrics.cohen_kappa_score; PYTHONPATH=. lets every script import supportdesk.
  • Comes out:

    text
    72 tickets, 12 articles
    

A word on what is real in this file. The build environment for this course has no LLM API key. Everything that does not need a hosted model is run for real and its output pasted here: the baselines, the error analysis, TinyLM latency and fine-tuning, the agent's guards, the evaluation harness, and all cost arithmetic. Where a script needs a model decision, it uses ScriptedLLM from Module 1. ScriptedLLM output is not model output: it tests loops, parsers, gates, and budgets, never quality, and every section says so again. Each starter runs against a real provider with one flag (--real) or by passing supportdesk.llm.chat in place of the stand-in, and your capstone numbers must come from that real run.

A few quirks of the canonical code matter here, and the starters work around them without editing it: TinyLM's tokenizer has 2,048 ids but only 1,503 real tokens; generate() crashes with an IndexError on a long prompt when max_new_tokens is 128 or more, because its truncation slice ids[-(context - max_new_tokens):] then stops shortening the prompt, so the starters use 24; greedy generate() with an allowed mask falls back to the lowest allowed id whenever the model's top token is disallowed, so the Adaptation and Evaluation starters score labels by log-probability instead of constrained generation; pricing.py bills Groq cached input at the full input price, although Groq's prompt-caching page states "a 50% discount for cached input tokens" (Groq docs); and every torch script calls torch.set_num_threads(1) so timings are stable on a small machine.

Part A: Choosing a Track

A capstone is not a bigger homework exercise. It is a claim ("this system is better than that one, at this cost, for these reasons") plus the evidence that makes a skeptical reviewer believe it. The four tracks make four different kinds of claim about the same Brightlane assistant.

TrackThe claim you defendWhat you build
Product"This LLM feature is ready to ship, and we will know if it breaks."One feature end to end: versioned prompts, schema-validated output, an eval suite with CI gates, cost and latency budgets, a safety review, a failure-mode analysis
Agent"This agent does useful work and cannot be talked into harm."An agent with real tools, sandboxing, termination guards, approval gates, full tracing, trajectory evaluation, and correct behavior under injected adversarial content
Adaptation"Fine-tuning beats prompting on this task, and it did not break anything else."A task where prompting plateaus, a fine-tuning dataset, an adapted model, lift over the prompted baseline, and a measured general-capability regression
Evaluation"This harness finds real quality problems, and its numbers can be trusted."An assertion tier, a calibrated judge tier, a noise-floor measurement, CI gates, and three real regressions found and fixed with it

[IMAGE: Four tracks drawn as four lanes leaving one shared box labelled "Brightlane assistant from Modules 1 to 14". Each lane shows its main artifact (CI gate, agent trace, fine-tuned model, eval report), and all four lanes pass through one shared band labelled "baseline, 50+ failures, measured cost, what did not work" before reaching "capstone defense". | Alt text: Four capstone tracks sharing four required pieces of evidence | File: cap-four-tracks.png]

How to choose

Pick the track where your evidence will be strongest, not the one that sounds most impressive. An Adaptation project that honestly shows no lift is a pass; an Agent project with no attack suite is not.

SituationUse thisWhy
You have an API key and want something a team would ship next monthProduct trackThe deliverable is a gate and a budget, which any real feature needs; the model can change under it
You care most about actions: refunds, account changes, tool useAgent trackThe risk sits in the tools, and the track makes you prove the guards hold under attack
Prompting gets you stuck at a ceiling on a narrow, well-labelled task, and you can get a few hundred labelsAdaptation trackFine-tuning pays off only with data and a plateau; the track makes you measure both lift and forgetting
You inherit a system and nobody can say whether a change made it betterEvaluation trackThe harness is the product; three real regressions found and fixed is the proof it works
No API key and no GPUEvaluation track (on the extractive drafter and TinyLM), or Adaptation with TinyLMBoth starters run fully offline with real numbers; the conclusions stay about the harness, not about LLM quality
You have a GPU or a hosted fine-tuning accountAdaptation track with a real open-weight modelTinyLM shows the method; a real model can show real lift

The six-week shape

Every track follows the same rhythm, and the milestone tables in Parts C to F fill it in.

WeekAll tracks
1Choose the task and metric; freeze the test split; run the Part B baseline and write down the number to beat
2Build the first end-to-end version, however crude; log every call through Meter
3Collect at least 50 real failures from your own system; tag and cluster them; pick the two biggest causes
4Fix what the clusters point at; measure each fix against the baseline with a paired test; keep a log of what did not work
5Track-specific hardening (gates, attack suite, regression budget, judge calibration); final measured cost and latency
6Evidence pack, written account, defense rehearsal

Rules for honest evidence

These rules come from the course, gathered in one place because a reviewer will check each one.

  1. Freeze the test split in week 1. You may read dev tickets, tune on dev, and select recipes on a slice of dev. You score test once per final system. Module 14 showed why: keyword rules scored 93.8% on the dev tickets they were written from and 58.3% on test.
  2. Always state n and an interval. With 24 test tickets, a Wilson 95% interval is about 35 points wide. Differences inside it are "within noise", not wins.
  3. Compare systems on the same items with a paired test. Count the tickets the new system fixes and the ones it breaks; the sign test on those two numbers is the honest comparison.
  4. Label stand-ins and illustrations. ScriptedLLM output is plumbing. Real-model outputs you did not capture are illustrative. Say so every time.
  5. Real failures drive decisions. Synthetic tickets can widen the search, but a cause that appears only in synthetic data is a hypothesis, not a priority.