Part A: The discipline
Module 7: Context Engineering
By the end of this module, you'll have:
- A measured picture of the help-center retriever (
supportdesk/kb_search.py): top-1 and top-3 hit rates per language, the tickets it misses and why, and one improvement with before-and-after numbers and an honest noise estimate. - A
ContextBuilder(examples/m07_context.py) that assembles the system instruction, customer profile and memory, conversation history, retrieved articles, tool definitions, few-shot examples, and a state scaffold, each under its own token budget, and prints a per-section token report. - A long-term memory store (
examples/m07_memory.py) that extracts facts from conversations, stores them with timestamps and sources, retrieves the relevant ones, resolves conflicts by source trust and recency, redacts secrets, and forgets on request. - An ablation diagnostic (
examples/m07_diagnose.py) that takes a bad answer and its context and finds the section, and then the item, that caused it: poisoning, clash, or distraction. - Measured optimizations (
examples/m07_optimize.py): shared-prefix tokens for prompt caching, just-in-time loading vs preloading, compaction triggers, and an A/B harness that tells you whether added context actually helped.
Prerequisites: Modules 1 to 6. In particular: token counting and the context budget equation (Module 2), prompt anatomy and separating instructions from data (Module 4), and structured outputs with DraftReply and tool calling (Module 6). Working Python; no machine learning background.
Where we are: Module 6 made the assistant's output reliable: typed drafts, validated tool calls, and a loop that returns errors to the model. This module works on the other side of the call. Everything the model knows about a ticket arrives through its context window, so what you put there, in what order, and at what size decides the quality of every draft. Module 8 then hands that context to an agent that runs for many steps.
How this module is organized
| Part | What it covers |
|---|---|
| Setup | One shared setup block; every example runs in order from the repo root |
| Part A: The discipline | Context as a scarce resource, the attention budget, context rot in published research and in a real TinyLM measurement |
| Part B: Retrieved knowledge | kb_search.py in full, hit rates by language, a multilingual improvement, recall@k vs tokens, ordering |
| Part C: Assembling a context | The ContextBuilder, per-section budgets and reports, system instruction stability, history, tool cost, few-shot, scaffolds |
| Part D: Memory | Session vs long-term memory, extraction, retrieval and injection, conflicts and staleness, privacy and user control |
| Part E: Failure modes | Poisoning, distraction, confusion, clash, and an ablation diagnostic that finds the culprit |
| Part F: Optimization | Cache-aware ordering, just-in-time vs preloading, compaction triggers, A/B testing added context |
| Module Lab | The whole pipeline over the dev split with a report and an automatic diagnosis |
Setup
Examples run in order from the repository root. Make a working copy as in earlier modules, activate your virtual environment, and set PYTHONPATH so Python finds both the canonical supportdesk package and this module's example files:
cd supportdesk
source .venv/bin/activate # or your own venv
export PYTHONPATH=.:examples
python examples/m07_setup.py
Code explained
- In simple words: go to the project, switch on the Python environment, and tell Python where to look for code.
- What happens:
PYTHONPATH=.:examplesputs the repo root (forsupportdesk.*) andexamples/(form07_*.py) on the import path. Every script in this module lives inexamples/asm07_<name>.py. Nothing here needs an API key; parts that can use a real model say so and take a--realflag. - Comes out: the setup script below prints the dataset shape.
The shared setup block loads the tickets, the retriever, and the token counter that the rest of the module uses:
# Shared setup for Module 7. Run from the repo root: PYTHONPATH=.:examples python examples/m07_setup.py
from collections import Counter
from supportdesk.data import load_tickets
from supportdesk.kb_search import KBSearch
from supportdesk.tokens import count_tokens
kb = KBSearch()
tickets = {t.id: t for t in load_tickets()}
dev = load_tickets("dev")
answerable = [t for t in load_tickets() if t.gold["answerable"]]
print(f"{len(tickets)} tickets, {len(dev)} dev, {len(answerable)} answerable")
print("languages:", dict(Counter(t.language for t in load_tickets())))
print("answerable by language:", dict(Counter(t.language for t in answerable)))
print("tokens in one ticket (T-1062, o200k_base):", count_tokens(tickets["T-1062"].text))
Code explained
- In simple words: load the 72 Brightlane tickets and the 12 help-center articles, and count how they split.
- What happens:
load_tickets()readsdata/tickets.jsonl; each ticket carries gold labels, includingkb_article(the article that answers it) andanswerable.KBSearch()builds the keyword index you will meet in Part B.count_tokensuses theo200k_basetokenizer from Module 2 by default. - Comes out:
72 tickets, 48 dev, 62 answerable
languages: {'en': 64, 'es': 2, 'de': 3, 'ja': 2, 'hi': 1}
answerable by language: {'en': 54, 'es': 2, 'de': 3, 'ja': 2, 'hi': 1}
tokens in one ticket (T-1062, o200k_base): 30
Sixty-two tickets are answerable from the help center. Eight of those are not in English: small numbers, which will matter every time we compare languages.
Part A: The discipline
A.1 Context is a curated resource, not a dumping ground
The context is everything the model reads on one call: the system instruction, the conversation so far, retrieved documents, tool definitions, examples, and the new message. The model has no other memory of your customer, your policies, or what happened five minutes ago. Context engineering is the work of deciding what goes into that window on each call, in what form, in what order, and at what size, so the model has what it needs and little else.
Prompt engineering (Module 4) was mostly about the wording of an instruction. Context engineering is about the whole request, most of which your code assembles automatically from data. For the Brightlane assistant a single draft request is built from seven sources:
Each box competes for the same window and the same share of the model's attention. The rest of this module is about making each box earn its place, and measuring when it does not.
A.2 The attention budget
Anthropic's engineering post "Effective context engineering for AI agents" (29 September 2025) puts the problem this way: "LLMs have an 'attention budget' that they draw on when parsing large volumes of context," because the transformer architecture "enables every token to attend to every other token across the entire context. This results in n^2 pairwise relationships for n tokens" (the post writes the exponent as a superscript). The same post defines the goal as finding "the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome."
In plain words: a model does not read a long context the way a database scans rows. Each new token's prediction is a weighted mix over everything before it. Every irrelevant paragraph gets some of that weight. Doubling the context does not double what the model notices; it spreads the same attention thinner. That is the attention budget: a metaphor, not a number you can read off a model card, but a useful one. It is why "just paste everything in, the window is a million tokens" usually produces worse drafts than a tight, well-chosen context.
A.3 Context rot: what the research shows
Context rot is the name Chroma's research team gave to quality that degrades as the input grows, even on tasks that stay trivially easy. Their report "Context Rot: How Increasing Input Tokens Impacts LLM Performance" (Kelly Hong, Anton Troynikov, Jeff Huber, 14 July 2025, trychroma.com/research/context-rot) tested 18 models, including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3 variants. Their findings, in their words and numbers:
- "Models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows."
- The less a question's wording resembles the text that answers it, the faster accuracy falls as the input grows.
- Distractors (passages that look relevant but are not) hurt, and some distractors hurt far more than others.
- On LongMemEval conversational questions, models did markedly better with a focused prompt of about 300 tokens than with the full prompt of about 113,000 tokens that contained the same answer.
- Counterintuitively, models did better on haystacks whose sentences were shuffled than on coherent ones.
An older, related result is "Lost in the Middle" (Liu et al., 2023, published in TACL 2024): on multi-document question answering, accuracy was highest when the relevant document sat at the start or end of the context and dropped when it sat in the middle. Module 2 introduced this position effect; here it becomes a design rule for where retrieved articles go (Part B.5).
For Brightlane, the practical reading is: a draft request should carry the articles that answer the ticket and not the ten that do not. Keeping the context small is not only cheaper; the research says it is usually more accurate too.
A.4 Measuring the mechanism in TinyLM
We cannot run GPT-4.1 here, but we can watch a real mechanism in TinyLM, the 1.07M-parameter model from Module 1 that was pretrained on support conversations. The experiment: for each of the 12 question types TinyLM learned, measure how likely it finds the correct agent answer. Then put irrelevant text in front of the question and measure again. Two kinds of irrelevant text:
- same-domain: other support conversations in the exact training format, about different topics.
- off-domain: plain prose about a lighthouse keeper that never appears in the corpus.
Lengths go up to 64 tokens, because TinyLM's whole window is 128 tokens and the question plus answer need the rest.
examples/m07_rot.py
"""Module 7: a real context-rot mechanism in TinyLM.
For each of the 12 question types TinyLM was trained on, measure how likely the
model finds the correct agent answer, then put irrelevant text in front of the
question and measure again. Two kinds of distractor:
same-domain: other support conversations (right format, wrong topic)
off-domain: plain prose that never appears in the corpus
Run: PYTHONPATH=. python examples/m07_rot.py
"""
from __future__ import annotations
import math
import random
import torch
from supportdesk.tinylm import load
torch.set_num_threads(2)
torch.manual_seed(0)
QA = [
("I was charged twice for the Team plan this month. Can you refund the duplicate?",
"Sorry about the duplicate charge. Duplicate charges are always refunded in full"),
("how do I cancel my Team subscription?",
"You can cancel at any time from Settings > Billing > Cancel plan."),
("I forgot my password and the reset link expired.",
"Reset links expire after 30 minutes and can be used once."),
("my account is locked after too many attempts.",
"After 5 failed attempts an account is locked for 15 minutes."),
("does the Team plan include SSO?",
"SSO is available on Business and Enterprise plans."),
("our automations stopped and there is a banner about a limit.",
"Automations pause when the monthly run limit is reached."),
("how do I export my board data?",
"Export a board to CSV or JSON from the board menu > Export."),
("slack notifications stopped posting.",
"The Slack token was usually revoked."),
("can I get a refund on my annual Team plan?",
"Annual plans cancelled within 14 days of purchase or renewal get a full refund."),
("where can I find my invoices?",
"Invoices are emailed to the billing contact"),
("how much does the Team plan cost?",
"Team costs 12 USD per user per month"),
("the board is not loading. Is there an outage?",
"Please check status.brightlane.example for live updates."),
]
OFF_DOMAIN = (
"The old lighthouse keeper walked along the cliff every evening, counting the gulls and "
"listening to the waves. In spring the village held a festival with music, bread, and "
"small boats painted blue and yellow. Nobody remembered who had built the first harbor wall, "
"but everyone agreed it had survived more storms than any of them. "
) * 4
def same_domain(skip: int, rng: random.Random) -> str:
"""Other support conversations from the training format, never about question `skip`."""
parts = []
for i in rng.sample([j for j in range(len(QA)) if j != skip], 6):
q, a = QA[i]
parts.append(f"Customer (Priya): Hi, {q}\nAgent (Jo): {a}\nCustomer (Priya): Thanks!\n")
return "\n".join(parts)
@torch.no_grad()
def answer_stats(model, tok, prefix_ids: list[int], question: str, answer: str) -> tuple[float, float, bool]:
"""(P(first answer token), geometric-mean P per answer token, greedy first 3 tokens right?)."""
prompt_ids = tok.encode(f"Customer (Ben): Hello, {question}\nAgent (Omar):").ids
answer_ids = tok.encode(" " + answer).ids
ids = prefix_ids + prompt_ids + answer_ids
assert len(ids) <= model.cfg.context, len(ids)
logits = model(torch.tensor([ids]))[0]
logp = torch.log_softmax(logits, dim=-1)
start = len(prefix_ids) + len(prompt_ids)
token_lp = [logp[start - 1 + i, t].item() for i, t in enumerate(answer_ids)]
greedy = [int(logits[start - 1 + i].argmax()) for i in range(3)]
return math.exp(token_lp[0]), math.exp(sum(token_lp) / len(token_lp)), greedy == answer_ids[:3]
def run(lengths=(0, 8, 16, 32, 48, 64), seeds=(0, 1, 2)) -> None:
model, tok = load()
sep = tok.encode("\n\n").ids
print(f"{'distractor':12} {'tokens':>6} {'P(first tok)':>13} {'P per tok':>10} {'greedy ok':>10}")
for kind in ("same-domain", "off-domain"):
for length in lengths:
firsts, means, oks, n = 0.0, 0.0, 0, 0
for seed in seeds:
rng = random.Random(seed)
for i, (q, a) in enumerate(QA):
text = same_domain(i, rng) if kind == "same-domain" else OFF_DOMAIN
ids = tok.encode(text).ids
offset = rng.randrange(0, len(ids) - length) if kind == "off-domain" and length else 0
prefix = (ids[offset: offset + length] + sep) if length else []
f, m, ok = answer_stats(model, tok, prefix, q, a)
firsts, means, oks, n = firsts + f, means + m, oks + ok, n + 1
print(f"{kind:12} {length:>6} {firsts / n:>13.3f} {means / n:>10.3f} {oks:>6}/{n}")
POISON = [ # (question index, a wrong answer that looks like a real agent reply)
(3, "Support can unlock your account right away."),
(8, "Annual plans can be refunded at any time for the remaining months."),
(10, "Team costs 8 USD per user per month"),
(4, "SSO is available on all plans, including Free."),
]
def poison(repeats=(0, 1)) -> None:
"""Context poisoning: an earlier turn in the window gave a wrong answer to the same question."""
model, tok = load()
print(f"\n{'question':44} {'wrong turns':>11} {'P/tok right':>12} {'P/tok wrong':>12}")
for idx, wrong in POISON:
q, right = QA[idx]
for r in repeats:
earlier = "".join(f"Customer (Priya): Hi, {q}\nAgent (Jo): {wrong}\nCustomer (Priya): Thanks!\n\n"
for _ in range(r))
prefix = tok.encode(earlier).ids if r else []
_, p_right, _ = answer_stats(model, tok, prefix, q, right)
_, p_wrong, _ = answer_stats(model, tok, prefix, q, wrong)
print(f"{q[:44]:44} {r:>11} {p_right:>12.3f} {p_wrong:>12.3f}")
REAL_PROBES = [ # (ticket-style question, a phrase the right answer must contain)
("How long is an account locked after too many failed sign-ins?", "15 minutes"),
("How many automation runs per month does the Business plan include?", "5,000"),
("Within how many days can an annual plan be cancelled for a full refund?", "14"),
("Above which file size do Android 15 uploads fail?", "25 MB"),
]
def real_model_rot(chat, lengths=(0, 1000, 4000, 16000), repeats: int = 3) -> None:
"""The same experiment on a hosted model: help-center facts, then a question, with filler between.
Filler is support conversations from data/corpus.txt (same domain, so it competes).
Needs an API key; prints pass counts per filler length. We do not show numbers for
this in the module because none were captured in this build.
"""
from supportdesk.data import DATA_DIR, load_articles
from supportdesk.tokens import count_tokens
facts = "\n\n".join(a.body for a in load_articles())
corpus = (DATA_DIR / "corpus.txt").read_text(encoding="utf-8")
for length in lengths:
filler, passed, total = "", 0, 0
for line in corpus.splitlines(keepends=True):
if count_tokens(filler) >= length:
break
filler += line
for question, must in REAL_PROBES:
for _ in range(repeats):
messages = [{"role": "system", "content": "Answer from the help center text only, in one sentence."},
{"role": "user", "content": f"<help_center>\n{facts}\n</help_center>\n\n"
f"<old_chats>\n{filler}</old_chats>\n\nQuestion: {question}"}]
passed += must in chat(messages, temperature=0.0, max_tokens=300).text
total += 1
print(f"filler {length:>6} tokens: {passed}/{total} correct")
if __name__ == "__main__":
import sys
if "--real" in sys.argv:
from supportdesk.llm import chat as real_chat
real_model_rot(real_chat)
else:
run()
poison()
Code explained
- In simple words: ask TinyLM the same 12 questions with growing amounts of junk in front, and watch how sure it stays of the right answer.
- What happens:
QAholds the 12 question and answer templates fromscripts/build_corpus.py(answers shortened to their first clause).OFF_DOMAINis prose the model has never seen.same_domainbuilds distractor conversations from the other 11 templates, so the distractor never contains the right answer.answer_statsruns one forward pass over distractor, question, and answer together (teacher forcing). It reads the probability of the first answer token, the geometric mean probability per answer token, and whether greedy decoding would produce the first 3 answer tokens. Theassertguarantees nothing was cut off by the 128-token window.runrepeats this for lengths 0 to 64 and 3 random seeds (different distractor samples and offsets), so each row averages 36 measurements.poisonis used in Part E: it plants an earlier turn in which an agent gave a wrong answer to the same question.real_model_rotis the same experiment for a hosted model (help-center text, then up to 16,000 tokens of old chats, then a question). It needs an API key: runpython examples/m07_rot.py --real. No real-model numbers were captured for this module, so none are shown.
- Comes out: (the run takes 30 seconds to 3 minutes on a laptop CPU; probabilities are deterministic on the same machine and may differ in the last digit elsewhere)
| Situation | Use this | Why |
|---|---|---|
| You want to know whether your model suffers from long inputs on your task | Run a filler-length sweep like real_model_rot on your own questions | Published curves use other tasks; the rate of decay is task and model specific |
| Context is growing because "it might help" | Default to leaving it out and add it back only when an A/B test (Part F.4) shows a gain | Research and the TinyLM run both show added irrelevant text lowers the right answer's probability |
| The relevant fact is phrased differently from the question | Retrieve and quote the specific passage rather than a whole long document | Chroma found low question-answer similarity decays fastest with length |