CourseLarge Language Models · Module 7: Context Engineering · part 30 of 80
Part 30 · Module 7: Context Engineering

Part A: The discipline

15 min read·22 Sept 2026

Module 7: Context Engineering

By the end of this module, you'll have:

  • A measured picture of the help-center retriever (supportdesk/kb_search.py): top-1 and top-3 hit rates per language, the tickets it misses and why, and one improvement with before-and-after numbers and an honest noise estimate.
  • A ContextBuilder (examples/m07_context.py) that assembles the system instruction, customer profile and memory, conversation history, retrieved articles, tool definitions, few-shot examples, and a state scaffold, each under its own token budget, and prints a per-section token report.
  • A long-term memory store (examples/m07_memory.py) that extracts facts from conversations, stores them with timestamps and sources, retrieves the relevant ones, resolves conflicts by source trust and recency, redacts secrets, and forgets on request.
  • An ablation diagnostic (examples/m07_diagnose.py) that takes a bad answer and its context and finds the section, and then the item, that caused it: poisoning, clash, or distraction.
  • Measured optimizations (examples/m07_optimize.py): shared-prefix tokens for prompt caching, just-in-time loading vs preloading, compaction triggers, and an A/B harness that tells you whether added context actually helped.

Prerequisites: Modules 1 to 6. In particular: token counting and the context budget equation (Module 2), prompt anatomy and separating instructions from data (Module 4), and structured outputs with DraftReply and tool calling (Module 6). Working Python; no machine learning background.

Where we are: Module 6 made the assistant's output reliable: typed drafts, validated tool calls, and a loop that returns errors to the model. This module works on the other side of the call. Everything the model knows about a ticket arrives through its context window, so what you put there, in what order, and at what size decides the quality of every draft. Module 8 then hands that context to an agent that runs for many steps.

How this module is organized

PartWhat it covers
SetupOne shared setup block; every example runs in order from the repo root
Part A: The disciplineContext as a scarce resource, the attention budget, context rot in published research and in a real TinyLM measurement
Part B: Retrieved knowledgekb_search.py in full, hit rates by language, a multilingual improvement, recall@k vs tokens, ordering
Part C: Assembling a contextThe ContextBuilder, per-section budgets and reports, system instruction stability, history, tool cost, few-shot, scaffolds
Part D: MemorySession vs long-term memory, extraction, retrieval and injection, conflicts and staleness, privacy and user control
Part E: Failure modesPoisoning, distraction, confusion, clash, and an ablation diagnostic that finds the culprit
Part F: OptimizationCache-aware ordering, just-in-time vs preloading, compaction triggers, A/B testing added context
Module LabThe whole pipeline over the dev split with a report and an automatic diagnosis

Setup

Examples run in order from the repository root. Make a working copy as in earlier modules, activate your virtual environment, and set PYTHONPATH so Python finds both the canonical supportdesk package and this module's example files:

bash
cd supportdesk
source .venv/bin/activate          # or your own venv
export PYTHONPATH=.:examples
python examples/m07_setup.py

Code explained

  • In simple words: go to the project, switch on the Python environment, and tell Python where to look for code.
  • What happens: PYTHONPATH=.:examples puts the repo root (for supportdesk.*) and examples/ (for m07_*.py) on the import path. Every script in this module lives in examples/ as m07_<name>.py. Nothing here needs an API key; parts that can use a real model say so and take a --real flag.
  • Comes out: the setup script below prints the dataset shape.

The shared setup block loads the tickets, the retriever, and the token counter that the rest of the module uses:

python
# Shared setup for Module 7. Run from the repo root: PYTHONPATH=.:examples python examples/m07_setup.py
from collections import Counter

from supportdesk.data import load_tickets
from supportdesk.kb_search import KBSearch
from supportdesk.tokens import count_tokens

kb = KBSearch()
tickets = {t.id: t for t in load_tickets()}
dev = load_tickets("dev")
answerable = [t for t in load_tickets() if t.gold["answerable"]]

print(f"{len(tickets)} tickets, {len(dev)} dev, {len(answerable)} answerable")
print("languages:", dict(Counter(t.language for t in load_tickets())))
print("answerable by language:", dict(Counter(t.language for t in answerable)))
print("tokens in one ticket (T-1062, o200k_base):", count_tokens(tickets["T-1062"].text))

Code explained

  • In simple words: load the 72 Brightlane tickets and the 12 help-center articles, and count how they split.
  • What happens: load_tickets() reads data/tickets.jsonl; each ticket carries gold labels, including kb_article (the article that answers it) and answerable. KBSearch() builds the keyword index you will meet in Part B. count_tokens uses the o200k_base tokenizer from Module 2 by default.
  • Comes out:
text
72 tickets, 48 dev, 62 answerable
languages: {'en': 64, 'es': 2, 'de': 3, 'ja': 2, 'hi': 1}
answerable by language: {'en': 54, 'es': 2, 'de': 3, 'ja': 2, 'hi': 1}
tokens in one ticket (T-1062, o200k_base): 30

Sixty-two tickets are answerable from the help center. Eight of those are not in English: small numbers, which will matter every time we compare languages.

Part A: The discipline

A.1 Context is a curated resource, not a dumping ground

The context is everything the model reads on one call: the system instruction, the conversation so far, retrieved documents, tool definitions, examples, and the new message. The model has no other memory of your customer, your policies, or what happened five minutes ago. Context engineering is the work of deciding what goes into that window on each call, in what form, in what order, and at what size, so the model has what it needs and little else.

Prompt engineering (Module 4) was mostly about the wording of an instruction. Context engineering is about the whole request, most of which your code assembles automatically from data. For the Brightlane assistant a single draft request is built from seven sources:

.

Each box competes for the same window and the same share of the model's attention. The rest of this module is about making each box earn its place, and measuring when it does not.

A.2 The attention budget

Anthropic's engineering post "Effective context engineering for AI agents" (29 September 2025) puts the problem this way: "LLMs have an 'attention budget' that they draw on when parsing large volumes of context," because the transformer architecture "enables every token to attend to every other token across the entire context. This results in n^2 pairwise relationships for n tokens" (the post writes the exponent as a superscript). The same post defines the goal as finding "the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome."

In plain words: a model does not read a long context the way a database scans rows. Each new token's prediction is a weighted mix over everything before it. Every irrelevant paragraph gets some of that weight. Doubling the context does not double what the model notices; it spreads the same attention thinner. That is the attention budget: a metaphor, not a number you can read off a model card, but a useful one. It is why "just paste everything in, the window is a million tokens" usually produces worse drafts than a tight, well-chosen context.

A.3 Context rot: what the research shows

Context rot is the name Chroma's research team gave to quality that degrades as the input grows, even on tasks that stay trivially easy. Their report "Context Rot: How Increasing Input Tokens Impacts LLM Performance" (Kelly Hong, Anton Troynikov, Jeff Huber, 14 July 2025, trychroma.com/research/context-rot) tested 18 models, including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3 variants. Their findings, in their words and numbers:

  • "Models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows."
  • The less a question's wording resembles the text that answers it, the faster accuracy falls as the input grows.
  • Distractors (passages that look relevant but are not) hurt, and some distractors hurt far more than others.
  • On LongMemEval conversational questions, models did markedly better with a focused prompt of about 300 tokens than with the full prompt of about 113,000 tokens that contained the same answer.
  • Counterintuitively, models did better on haystacks whose sentences were shuffled than on coherent ones.

An older, related result is "Lost in the Middle" (Liu et al., 2023, published in TACL 2024): on multi-document question answering, accuracy was highest when the relevant document sat at the start or end of the context and dropped when it sat in the middle. Module 2 introduced this position effect; here it becomes a design rule for where retrieved articles go (Part B.5).

For Brightlane, the practical reading is: a draft request should carry the articles that answer the ticket and not the ten that do not. Keeping the context small is not only cheaper; the research says it is usually more accurate too.

A.4 Measuring the mechanism in TinyLM

We cannot run GPT-4.1 here, but we can watch a real mechanism in TinyLM, the 1.07M-parameter model from Module 1 that was pretrained on support conversations. The experiment: for each of the 12 question types TinyLM learned, measure how likely it finds the correct agent answer. Then put irrelevant text in front of the question and measure again. Two kinds of irrelevant text:

  • same-domain: other support conversations in the exact training format, about different topics.
  • off-domain: plain prose about a lighthouse keeper that never appears in the corpus.

Lengths go up to 64 tokens, because TinyLM's whole window is 128 tokens and the question plus answer need the rest.

examples/m07_rot.py

python
"""Module 7: a real context-rot mechanism in TinyLM.

For each of the 12 question types TinyLM was trained on, measure how likely the
model finds the correct agent answer, then put irrelevant text in front of the
question and measure again. Two kinds of distractor:
  same-domain: other support conversations (right format, wrong topic)
  off-domain:  plain prose that never appears in the corpus
Run:  PYTHONPATH=. python examples/m07_rot.py
"""
from __future__ import annotations

import math
import random

import torch

from supportdesk.tinylm import load

torch.set_num_threads(2)
torch.manual_seed(0)

QA = [
    ("I was charged twice for the Team plan this month. Can you refund the duplicate?",
     "Sorry about the duplicate charge. Duplicate charges are always refunded in full"),
    ("how do I cancel my Team subscription?",
     "You can cancel at any time from Settings > Billing > Cancel plan."),
    ("I forgot my password and the reset link expired.",
     "Reset links expire after 30 minutes and can be used once."),
    ("my account is locked after too many attempts.",
     "After 5 failed attempts an account is locked for 15 minutes."),
    ("does the Team plan include SSO?",
     "SSO is available on Business and Enterprise plans."),
    ("our automations stopped and there is a banner about a limit.",
     "Automations pause when the monthly run limit is reached."),
    ("how do I export my board data?",
     "Export a board to CSV or JSON from the board menu > Export."),
    ("slack notifications stopped posting.",
     "The Slack token was usually revoked."),
    ("can I get a refund on my annual Team plan?",
     "Annual plans cancelled within 14 days of purchase or renewal get a full refund."),
    ("where can I find my invoices?",
     "Invoices are emailed to the billing contact"),
    ("how much does the Team plan cost?",
     "Team costs 12 USD per user per month"),
    ("the board is not loading. Is there an outage?",
     "Please check status.brightlane.example for live updates."),
]

OFF_DOMAIN = (
    "The old lighthouse keeper walked along the cliff every evening, counting the gulls and "
    "listening to the waves. In spring the village held a festival with music, bread, and "
    "small boats painted blue and yellow. Nobody remembered who had built the first harbor wall, "
    "but everyone agreed it had survived more storms than any of them. "
) * 4


def same_domain(skip: int, rng: random.Random) -> str:
    """Other support conversations from the training format, never about question `skip`."""
    parts = []
    for i in rng.sample([j for j in range(len(QA)) if j != skip], 6):
        q, a = QA[i]
        parts.append(f"Customer (Priya): Hi, {q}\nAgent (Jo): {a}\nCustomer (Priya): Thanks!\n")
    return "\n".join(parts)


@torch.no_grad()
def answer_stats(model, tok, prefix_ids: list[int], question: str, answer: str) -> tuple[float, float, bool]:
    """(P(first answer token), geometric-mean P per answer token, greedy first 3 tokens right?)."""
    prompt_ids = tok.encode(f"Customer (Ben): Hello, {question}\nAgent (Omar):").ids
    answer_ids = tok.encode(" " + answer).ids
    ids = prefix_ids + prompt_ids + answer_ids
    assert len(ids) <= model.cfg.context, len(ids)
    logits = model(torch.tensor([ids]))[0]
    logp = torch.log_softmax(logits, dim=-1)
    start = len(prefix_ids) + len(prompt_ids)
    token_lp = [logp[start - 1 + i, t].item() for i, t in enumerate(answer_ids)]
    greedy = [int(logits[start - 1 + i].argmax()) for i in range(3)]
    return math.exp(token_lp[0]), math.exp(sum(token_lp) / len(token_lp)), greedy == answer_ids[:3]


def run(lengths=(0, 8, 16, 32, 48, 64), seeds=(0, 1, 2)) -> None:
    model, tok = load()
    sep = tok.encode("\n\n").ids
    print(f"{'distractor':12} {'tokens':>6} {'P(first tok)':>13} {'P per tok':>10} {'greedy ok':>10}")
    for kind in ("same-domain", "off-domain"):
        for length in lengths:
            firsts, means, oks, n = 0.0, 0.0, 0, 0
            for seed in seeds:
                rng = random.Random(seed)
                for i, (q, a) in enumerate(QA):
                    text = same_domain(i, rng) if kind == "same-domain" else OFF_DOMAIN
                    ids = tok.encode(text).ids
                    offset = rng.randrange(0, len(ids) - length) if kind == "off-domain" and length else 0
                    prefix = (ids[offset: offset + length] + sep) if length else []
                    f, m, ok = answer_stats(model, tok, prefix, q, a)
                    firsts, means, oks, n = firsts + f, means + m, oks + ok, n + 1
            print(f"{kind:12} {length:>6} {firsts / n:>13.3f} {means / n:>10.3f} {oks:>6}/{n}")



POISON = [  # (question index, a wrong answer that looks like a real agent reply)
    (3, "Support can unlock your account right away."),
    (8, "Annual plans can be refunded at any time for the remaining months."),
    (10, "Team costs 8 USD per user per month"),
    (4, "SSO is available on all plans, including Free."),
]


def poison(repeats=(0, 1)) -> None:
    """Context poisoning: an earlier turn in the window gave a wrong answer to the same question."""
    model, tok = load()
    print(f"\n{'question':44} {'wrong turns':>11} {'P/tok right':>12} {'P/tok wrong':>12}")
    for idx, wrong in POISON:
        q, right = QA[idx]
        for r in repeats:
            earlier = "".join(f"Customer (Priya): Hi, {q}\nAgent (Jo): {wrong}\nCustomer (Priya): Thanks!\n\n"
                              for _ in range(r))
            prefix = tok.encode(earlier).ids if r else []
            _, p_right, _ = answer_stats(model, tok, prefix, q, right)
            _, p_wrong, _ = answer_stats(model, tok, prefix, q, wrong)
            print(f"{q[:44]:44} {r:>11} {p_right:>12.3f} {p_wrong:>12.3f}")


REAL_PROBES = [  # (ticket-style question, a phrase the right answer must contain)
    ("How long is an account locked after too many failed sign-ins?", "15 minutes"),
    ("How many automation runs per month does the Business plan include?", "5,000"),
    ("Within how many days can an annual plan be cancelled for a full refund?", "14"),
    ("Above which file size do Android 15 uploads fail?", "25 MB"),
]


def real_model_rot(chat, lengths=(0, 1000, 4000, 16000), repeats: int = 3) -> None:
    """The same experiment on a hosted model: help-center facts, then a question, with filler between.

    Filler is support conversations from data/corpus.txt (same domain, so it competes).
    Needs an API key; prints pass counts per filler length. We do not show numbers for
    this in the module because none were captured in this build.
    """
    from supportdesk.data import DATA_DIR, load_articles
    from supportdesk.tokens import count_tokens
    facts = "\n\n".join(a.body for a in load_articles())
    corpus = (DATA_DIR / "corpus.txt").read_text(encoding="utf-8")
    for length in lengths:
        filler, passed, total = "", 0, 0
        for line in corpus.splitlines(keepends=True):
            if count_tokens(filler) >= length:
                break
            filler += line
        for question, must in REAL_PROBES:
            for _ in range(repeats):
                messages = [{"role": "system", "content": "Answer from the help center text only, in one sentence."},
                            {"role": "user", "content": f"<help_center>\n{facts}\n</help_center>\n\n"
                                                        f"<old_chats>\n{filler}</old_chats>\n\nQuestion: {question}"}]
                passed += must in chat(messages, temperature=0.0, max_tokens=300).text
                total += 1
        print(f"filler {length:>6} tokens: {passed}/{total} correct")


if __name__ == "__main__":
    import sys
    if "--real" in sys.argv:
        from supportdesk.llm import chat as real_chat
        real_model_rot(real_chat)
    else:
        run()
        poison()

Code explained

  • In simple words: ask TinyLM the same 12 questions with growing amounts of junk in front, and watch how sure it stays of the right answer.
  • What happens:
    • QA holds the 12 question and answer templates from scripts/build_corpus.py (answers shortened to their first clause). OFF_DOMAIN is prose the model has never seen.
    • same_domain builds distractor conversations from the other 11 templates, so the distractor never contains the right answer.
    • answer_stats runs one forward pass over distractor, question, and answer together (teacher forcing). It reads the probability of the first answer token, the geometric mean probability per answer token, and whether greedy decoding would produce the first 3 answer tokens. The assert guarantees nothing was cut off by the 128-token window.
    • run repeats this for lengths 0 to 64 and 3 random seeds (different distractor samples and offsets), so each row averages 36 measurements.
    • poison is used in Part E: it plants an earlier turn in which an agent gave a wrong answer to the same question.
    • real_model_rot is the same experiment for a hosted model (help-center text, then up to 16,000 tokens of old chats, then a question). It needs an API key: run python examples/m07_rot.py --real. No real-model numbers were captured for this module, so none are shown.
  • Comes out: (the run takes 30 seconds to 3 minutes on a laptop CPU; probabilities are deterministic on the same machine and may differ in the last digit elsewhere)

.
SituationUse thisWhy
You want to know whether your model suffers from long inputs on your taskRun a filler-length sweep like real_model_rot on your own questionsPublished curves use other tasks; the rate of decay is task and model specific
Context is growing because "it might help"Default to leaving it out and add it back only when an A/B test (Part F.4) shows a gainResearch and the TinyLM run both show added irrelevant text lowers the right answer's probability
The relevant fact is phrased differently from the questionRetrieve and quote the specific passage rather than a whole long documentChroma found low question-answer similarity decays fastest with length