CourseLarge Language Models · Module 7: Context Engineering · part 34 of 80
Part 34 · Module 7: Context Engineering

Part E: Failure modes

12 min read·22 Sept 2026

E.1 Four ways a context goes wrong

Drew Breunig's essay "How Long Contexts Fail" (22 June 2025, dbreunig.com) named four failure modes that have become common vocabulary. Each with a Brightlane example:

  • Context poisoning: a wrong fact enters the context and gets referenced again and again. Breunig's example is a Gemini agent playing Pokemon whose hallucinated game state "poisoned" its goals. Brightlane version: an earlier assistant turn wrongly said annual plans are refunded pro rata; every later draft in the thread repeats it.
  • Context distraction: the context grows so large that the model leans on it instead of on what it learned in training. Breunig cites the Gemini 2.5 agent drifting toward repeating past actions beyond about 100,000 tokens, and a Databricks study in which Llama 3.1 405b's correctness began to fall around 32,000 tokens. Brightlane version: five articles, four irrelevant, and the one that answers the question is lost among them.
  • Context confusion: superfluous content, often too many tools, changes the answer. Breunig cites a quantized Llama 3.1 8b that failed with 46 tools available but succeeded with 19. Brightlane version: sending every tool to every ticket.
  • Context clash: parts of the context contradict each other. Breunig cites Microsoft and Salesforce research in which spreading a task's information across several turns ("sharding") lowered scores by an average of 39 percent, with o3 falling from 98.1 to 64.1. Brightlane version: a remembered quote of "Team at 15 USD" alongside the pricing article's 12 USD.

These numbers come from the cited sources, measured on their tasks and models, as summarized in Breunig's essay; follow the links there before quoting them further.

E.2 Trying to poison TinyLM

Can we show poisoning with a real model here? The poison function in examples/m07_rot.py (shown in Part A.4) puts an earlier exchange in the window in which an agent gave a wrong answer to the same question, then measures how likely TinyLM finds the right and the wrong answer. It runs after the distractor table:

bash
python examples/m07_rot.py

Code explained

  • In simple words: plant one wrong agent answer in TinyLM's window and see whether it copies it.
  • What happens: for four questions, poison() compares the geometric-mean probability per token of the right answer and of a plausible wrong answer, with 0 and with 1 wrong exchange in front of the question. Two wrong exchanges do not fit in TinyLM's 128-token window with the question and answer, so the run stops at 1.
  • Comes out: the second table printed by the script:
text
question                                     wrong turns  P/tok right  P/tok wrong
my account is locked after too many attempts           0        0.999        0.000
my account is locked after too many attempts           1        0.998        0.000
can I get a refund on my annual Team plan?             0        0.999        0.001
can I get a refund on my annual Team plan?             1        0.999        0.001
how much does the Team plan cost?                      0        0.998        0.062
how much does the Team plan cost?                      1        0.997        0.049
does the Team plan include SSO?                        0        0.998        0.001
does the Team plan include SSO?                        1        0.994        0.001

TinyLM is not poisoned. With one wrong turn in the window, it still assigns the right answer about 0.99 per token and the wrong one almost nothing (0.000 to 0.062). This is a real and useful negative result: TinyLM memorized its corpus so thoroughly that it barely reads its context at all. Large instruction-tuned models are the opposite; they are trained to use their context, which is exactly what makes them easy to poison. So the mechanism we need for this part is one TinyLM cannot show, and we switch to a clearly labelled stand-in to build and test the diagnostic. Run the diagnostic against a real model with the same harness to see real behavior.

E.3 A diagnostic that finds the culprit

When a draft is wrong, the question is which part of the context made it wrong. Ablation answers that empirically: remove one section, rerun, and see whether the answer becomes acceptable. Then do the same for each item within the sections. The pattern of which removals fix the answer tells you the failure mode:

  • exactly one item's removal fixes it, and it contradicts another section on a checkable fact: clash;
  • exactly one item's removal fixes it, with no contradiction found: poisoning by that item;
  • removing any of several items fixes it: distraction (volume, not one bad item);
  • no single removal fixes it: the needed fact is missing (a retrieval problem) or the model cannot do the task.

examples/m07_diagnose.py

python
"""Module 7: find which part of the context caused a bad answer, by ablation.

Given a context split into sections (each a list of items), a way to call the
model, and a check that says whether an answer is acceptable, the diagnostic
removes one section at a time, then one item at a time, reruns, and reports
which removals turn the bad answer into a good one. A small numeric-claim
check flags clashes between sections.

The responder used here is a ScriptedLLM with rules that react to specific
context contents (a planted wrong turn, a stale price, too many articles).
It is NOT a model: it exists so the diagnostic's plumbing can be run and
tested offline. Run the same harness with supportdesk.llm.chat for real results.
Run:  PYTHONPATH=. python examples/m07_diagnose.py
"""
from __future__ import annotations

import re
from collections.abc import Callable
from dataclasses import dataclass, field
from typing import Any

from supportdesk.data import get_article
from supportdesk.kb_search import tokenize
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_tokens

ORDER = ("system", "profile", "history", "knowledge", "ticket")
SYSTEM = "You are the Brightlane support assistant. Answer only from the knowledge and profile. Keep it short."


def render(sections: dict[str, list[str]]) -> list[dict[str, str]]:
    """Sections to chat messages. Each section becomes a tagged block; empty sections vanish."""
    blocks = [f"<{name}>\n" + "\n".join(sections[name]) + f"\n</{name}>"
              for name in ORDER[1:] if sections.get(name)]
    return [{"role": "system", "content": "\n".join(sections.get("system", [SYSTEM]))},
            {"role": "user", "content": "\n\n".join(blocks)}]


# --- the scripted responder (rules, not intelligence) ---------------------------------------------
def scripted_responder(messages: list[dict[str, Any]], kwargs: dict[str, Any]) -> str:
    user = messages[-1]["content"]
    ticket = re.search(r"<ticket>\n(.*?)\n</ticket>", user, re.S)
    question = ticket.group(1) if ticket else ""
    articles = re.findall(r"<article[^>]*>\n(.*?)\n</article>", user, re.S)
    # Rule 1 (poisoning): an earlier claim of pro rata refunds is repeated as fact.
    if re.search(r"refunded pro rata for the remaining", user):
        return "Good news: as we said earlier, annual plans are refunded pro rata for the remaining months."
    # Rule 2 (clash): the first Team price that appears anywhere wins.
    if "Team" in question and ("save" in question or "cost" in question):
        price = re.search(r"Team (?:costs |at )?(\d+) USD per user per month", user)
        if price:
            return f"Team is {price.group(1)} USD per user per month."
    # Rule 3 (distraction): with more than 4 articles, only the first article is read.
    if len(articles) > 4:
        articles = articles[:1]
    # Otherwise: return the knowledge sentence that shares the most words with the question.
    q = set(tokenize(question))
    sentences = [s for a in articles for s in re.split(r"(?<=[.])\s", a)]
    best = max(sentences, key=lambda s: len(q & set(tokenize(s))), default="")
    return best if best and q & set(tokenize(best)) else "I could not find this in the help center."


# --- the diagnostic ---------------------------------------------------------------------------------
@dataclass
class Diagnosis:
    baseline: str
    baseline_ok: bool
    section_fixes: list[str] = field(default_factory=list)
    item_fixes: list[tuple[str, int, str]] = field(default_factory=list)
    clashes: list[str] = field(default_factory=list)
    calls: int = 0
    verdict: str = ""


PRICE = re.compile(r"(Team|Business)\D{0,20}?(\d+) USD per user per month")


def numeric_clashes(sections: dict[str, list[str]]) -> list[str]:
    """Flag the same plan quoted at different monthly prices in different sections."""
    seen: dict[str, set[tuple[str, str]]] = {}
    for name, items in sections.items():
        for item in items:
            for plan, amount in PRICE.findall(item):
                seen.setdefault(plan, set()).add((amount, name))
    return [f"{plan}: " + ", ".join(f"{a} USD in {s}" for a, s in sorted(v))
            for plan, v in seen.items() if len({a for a, _ in v}) > 1]


def diagnose(sections: dict[str, list[str]], call: Callable[[list[dict]], str],
             ok: Callable[[str], bool], protected: tuple[str, ...] = ("system", "ticket")) -> Diagnosis:
    """Ablate each section, then each item of every section, and record which removals fix the answer."""
    d = Diagnosis(baseline=call(render(sections)), baseline_ok=False)
    d.baseline_ok, d.calls = ok(d.baseline), 1
    if d.baseline_ok:
        d.verdict = "answer passes; nothing to diagnose"
        return d
    for name in sections:
        if name in protected:
            continue
        d.calls += 1
        if ok(call(render({k: v for k, v in sections.items() if k != name}))):
            d.section_fixes.append(name)
        for i, item in enumerate(sections[name]):
            trimmed = {**sections, name: sections[name][:i] + sections[name][i + 1:]}
            d.calls += 1
            if ok(call(render(trimmed))):
                d.item_fixes.append((name, i, item[:70]))
    d.clashes = numeric_clashes(sections)
    by_section: dict[str, int] = {}
    for name, _, _ in d.item_fixes:
        by_section[name] = by_section.get(name, 0) + 1
    if any(n >= 2 for n in by_section.values()):
        name = max(by_section, key=by_section.get)
        d.verdict = (f"distraction: removing any of {by_section[name]} items in '{name}' fixes it, "
                     "so volume, not one bad item, is the problem")
    elif len(d.item_fixes) == 1 and d.clashes:
        d.verdict = f"clash: one item in '{d.item_fixes[0][0]}' contradicts another section ({d.clashes[0]})"
    elif len(d.item_fixes) == 1:
        d.verdict = f"poisoning: one item in '{d.item_fixes[0][0]}' carries a wrong claim the answer repeats"
    elif d.section_fixes:
        d.verdict = f"section-level: removing {d.section_fixes} fixes it but no single item does"
    else:
        d.verdict = "no single removal fixes it: check retrieval (missing knowledge) or the model itself"
    return d


def article(article_id: str) -> str:
    return f'<article id="{article_id}">\n{get_article(article_id).body}\n</article>'


CASES = {
    "T-1005 refund after 2 months": (
        {"system": [SYSTEM],
         "profile": ["customer_tier: team", "remembered: pays annually (customer_said, 2026-07-02)"],
         "history": ["user: We paid annually in July. Can we stop now?",
                     "assistant: Annual plans can be refunded pro rata for the remaining months, I will check.",
                     "user: OK, how does that work?"],
         "knowledge": [article("billing-refunds"), article("billing-invoices")],
         "ticket": ["We paid annually in July and want to stop now. Can I get a refund for the remaining months?"]},
        lambda a: "pro rata" not in a and ("14 days" in a or "not refunded" in a)),
    "T-1067 annual discount on Team": (
        {"system": [SYSTEM],
         "profile": ["customer_tier: team", "remembered: was quoted Team at 15 USD per user per month (agent_note, 2025-02-01)"],
         "history": [],
         "knowledge": [article("billing-plans")],
         "ticket": ["How much do we save per user per month going annual on Team? Team cost today?"]},
        lambda a: "12" in a),
    "T-1016 webhook retries": (
        {"system": [SYSTEM],
         "profile": ["customer_tier: business"],
         "history": [],
         "knowledge": [article(a) for a in ("integrations-slack", "billing-plans", "boards-automations",
                                              "status-incidents", "mobile-app")],
         "ticket": ["Our webhook endpoint was down for 2 minutes. Will the automation retry the calls?"]},
        lambda a: "retr" in a.lower()),
}


if __name__ == "__main__":
    for name, (sections, ok) in CASES.items():
        llm = ScriptedLLM(responder=scripted_responder)  # rules that react to planted content, not a model
        d = diagnose(sections, lambda msgs: llm(msgs).text, ok)
        tokens = count_tokens(render(sections)[1]["content"])
        print(f"== {name} (user message {tokens} tokens)")
        print(f"  bad answer: {d.baseline}")
        print(f"  sections whose removal fixes it: {d.section_fixes or 'none'}")
        for section, i, text in d.item_fixes:
            print(f"  item fix: {section}[{i}] {text!r}")
        if d.clashes:
            print(f"  numeric clash: {d.clashes}")
        print(f"  verdict: {d.verdict}")
        print(f"  model calls spent on the diagnosis: {d.calls}\n")

Code explained

  • In simple words: take the bad answer, pull out one piece of the context at a time, and see which missing piece makes the answer good again.
  • What happens:
    • render(sections): turns a dictionary of sections (each a list of items) into chat messages with tagged blocks. Keeping the context as data, not a finished string, is what makes ablation possible.
    • scripted_responder: the ScriptedLLM rules. It is not a model. It repeats a planted "refunded pro rata" claim (poisoning), takes the first Team price it sees (clash), reads only the first article when given more than four (distraction), and otherwise returns the knowledge sentence with the most word overlap. Each rule mimics a failure that real models show, so the diagnostic has something to find.
    • Diagnosis: the baseline answer, which section and item removals fixed it, detected clashes, the number of model calls spent, and a verdict.
    • numeric_clashes(sections): a deterministic check that finds the same plan quoted at different monthly prices in different sections. Real systems extend this to dates, limits, and plan features.
    • diagnose(sections, call, ok): runs the baseline; if it fails, removes each unprotected section and each item, rerunning after every removal; then classifies the pattern as described above. System and ticket are protected: removing the question to "fix" an answer diagnoses nothing.
    • CASES: three realistic tickets, each with a planted problem and an ok check (for example, the refund answer must not say "pro rata" and must mention the 14-day rule).
  • Comes out: (deterministic; every answer below comes from the scripted rules, not from a model)

    text
    == T-1005 refund after 2 months (user message 335 tokens)
      bad answer: Good news: as we said earlier, annual plans are refunded pro rata for the remaining months.
      sections whose removal fixes it: ['history']
      item fix: history[1] 'assistant: Annual plans can be refunded pro rata for the remaining mon'
      verdict: poisoning: one item in 'history' carries a wrong claim the answer repeats
      model calls spent on the diagnosis: 11
    
    == T-1067 annual discount on Team (user message 185 tokens)
      bad answer: Team is 15 USD per user per month.
      sections whose removal fixes it: ['profile']
      item fix: profile[1] 'remembered: was quoted Team at 15 USD per user per month (agent_note, '
      numeric clash: ['Team: 12 USD in knowledge, 15 USD in profile']
      verdict: clash: one item in 'profile' contradicts another section (Team: 12 USD in knowledge, 15 USD in profile)
      model calls spent on the diagnosis: 7
    
    == T-1016 webhook retries (user message 546 tokens)
      bad answer: I could not find this in the help center.
      sections whose removal fixes it: none
      item fix: knowledge[0] '<article id="integrations-slack">\nConnect Slack from Settings > Integr'
      item fix: knowledge[1] '<article id="billing-plans">\nBrightlane has four plans. Free supports '
      item fix: knowledge[3] '<article id="status-incidents">\nLive status is at status.brightlane.ex'
      item fix: knowledge[4] '<article id="mobile-app">\nBrightlane has apps for iOS 17 or later and '
      verdict: distraction: removing any of 4 items in 'knowledge' fixes it, so volume, not one bad item, is the problem
      model calls spent on the diagnosis: 10
    

    All three verdicts are right, and each arrives with evidence you can act on:

  • T-1005 points at history[1], the earlier assistant turn with the false pro-rata claim. The fix is to correct or remove that turn and, more importantly, find out how it got there (an earlier draft that a human approved without catching it).
  • T-1067 points at the remembered "15 USD" quote and names the clash with the knowledge section's 12 USD. The fix belongs in the memory policy: a price is a fact the CRM or help center owns, so it should never be stored as a memory (D.1).
  • T-1016 shows the distraction signature: removing any one of the four irrelevant articles fixes the answer, and no whole-section removal does. The fix is retrieval size (k = 3 instead of 5), not deleting a particular article.
  • The diagnosis cost 7 to 11 calls. With a real model, run each ablation 3 to 5 times at your production temperature and count fixes, because one rerun of a stochastic model can "fix" an answer by chance. That multiplies the cost, so run the diagnostic on failures you have already collected, not on live traffic.
Symptom in the draftLikely failureFirst fix to try
Repeats a specific wrong claim that appears earlier in the thread or in memoryPoisoningRemove or correct the item; add a trust or TTL rule so it cannot return
States one of two conflicting valuesClashDecide which source owns that fact; drop the other from context
Misses an answer that is in the context, gets better with fewer articlesDistractionRetrieve fewer, better items; put the best one first or last
Calls the wrong tool or invents a tool argumentConfusionSend fewer tools; select tools by category
No removal helps, gold article missing from contextRetrieval missFix retrieval (Part B), not the prompt