Part C: Managing context
A support chat grows every turn, and Part B showed that simply sending everything eventually fails: too long, too expensive, or too diluted. Context management is deciding what to keep. We compare four strategies on the 23-turn chat with a deliberately small 300-token budget, so the effects show on a short example; the same logic applies at 30,000 tokens.
Four strategies
# examples/m02_context.py
"""Module 2: four ways to fit a long conversation into a token budget.
Every strategy keeps the system message and the newest user message, and
returns messages in their original order. Budgets are in estimated tokens
(count_messages with o200k_base), so leave a safety margin for the real model.
"""
import re
from collections.abc import Callable
from supportdesk.tokens import count_messages
ID = re.compile(r"\b[A-Z]{2,4}-\d{4}-\d{3,6}\b") # invoice and ticket ids
MONEY_OR_COUNT = re.compile(r"\b\d+(\.\d+)?\s*(USD|EUR|seats?|users?)\b", re.I)
PLAN = re.compile(r"\b(Free|Team|Business|Enterprise) plan\b")
ASK = re.compile(r"\b(please|can you|could you|i need|we need|refund|cancel)\b", re.I)
def _split(messages: list[dict]) -> tuple[dict, list[dict], dict]:
return messages[0], messages[1:-1], messages[-1]
def drop_oldest(messages: list[dict], budget: int) -> list[dict]:
"""Remove the oldest history messages one at a time until the request fits."""
system, history, last = _split(messages)
while history and count_messages([system, *history, last]) > budget:
history = history[1:]
return [system, *history, last]
def sliding_window(messages: list[dict], keep_last: int = 8, pin_first: bool = True) -> list[dict]:
"""Keep the last `keep_last` history messages, plus (optionally) the first user message as an anchor."""
system, history, last = _split(messages)
recent = history[-keep_last:] if keep_last > 0 else []
anchor = [history[0]] if pin_first and history and history[0] not in recent else []
return [system, *anchor, *recent, last]
def importance(message: dict, position: int, total: int) -> float:
"""Score a message: hard facts and requests matter most, then recency."""
text = message["content"]
score = 3.0 * bool(ID.search(text)) + 2.0 * bool(MONEY_OR_COUNT.search(text)) + 2.0 * bool(PLAN.search(text))
if message["role"] == "user" and ASK.search(text):
score += 1.5
return score + position / total # tie-break toward newer messages
def importance_ranked(messages: list[dict], budget: int) -> list[dict]:
"""Keep the highest-scoring history messages that fit, in their original order."""
system, history, last = _split(messages)
ranked = sorted(range(len(history)), key=lambda i: importance(history[i], i, len(history)), reverse=True)
chosen: set[int] = set()
for i in ranked:
trial = [system, *(history[j] for j in sorted(chosen | {i})), last]
if count_messages(trial) <= budget:
chosen.add(i)
return [system, *(history[j] for j in sorted(chosen)), last]
SUMMARY_PROMPT = (
"Summarize the earlier part of this support conversation for the agent who takes over. "
"Keep exactly: every id (invoice, ticket), plan and seat count, amounts, what the customer "
"asked for and whether it is resolved, and anything the agent promised. Drop small talk. "
"At most 80 words."
)
def compact_with_summary(messages: list[dict], summarize: Callable, keep_last: int = 6) -> list[dict]:
"""Replace all but the last few history messages with one summary message.
`summarize` has the signature of supportdesk.llm.chat (or ScriptedLLM).
"""
system, history, last = _split(messages)
old, recent = history[:-keep_last], history[-keep_last:]
if not old:
return messages
transcript = "\n".join(f"{m['role']}: {m['content']}" for m in old)
result = summarize([{"role": "system", "content": SUMMARY_PROMPT},
{"role": "user", "content": transcript}], max_tokens=200)
note = {"role": "system", "content": f"Summary of the earlier conversation: {result.text.strip()}"}
return [system, note, *recent, last]Code explained
- In simple words: four different rules for deciding which messages to keep when the conversation no longer fits.
- What happens: every strategy keeps the system message and the newest user message and preserves order.
drop_oldest(messages, budget): removes the oldest history message untilcount_messagesfits the budget. Simple and predictable; it always loses the beginning first.sliding_window(messages, keep_last, pin_first): keeps a fixed number of recent messages regardless of their size, and optionally pins the first user message, which in support chats often states the problem. Its token cost varies with message length, so pair it with a budget check.importance(message, position, total)andimportance_ranked(messages, budget): score each message (3 for an id, 2 for an amount or seat count, 2 for a plan name, 1.5 for a user request, plus a small recency bonus), then greedily keep the highest scores that fit and restore the original order. The patterns (ID,MONEY_OR_COUNT,PLAN,ASK) encode what matters in Brightlane chats.SUMMARY_PROMPTandcompact_with_summary(messages, summarize, keep_last): send everything except the last few messages to a summarizer with explicit instructions about what to preserve, then replace them with one system note.summarizehas the signature ofllm.chat, so you can pass a real model or aScriptedLLM.- Comes out: nothing when run alone; the next script compares them.
Which facts survive
# examples/m02_compaction.py
"""Module 2: compare four context strategies on a long support chat and check which facts survive."""
import re
from m02_context import ASK, ID, MONEY_OR_COUNT, PLAN, compact_with_summary, drop_oldest, importance_ranked, sliding_window
from m02_conversation import conversation, facts_kept
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages
BUDGET = 300 # deliberately small so the effects show on a 23-turn chat
def extractive_summary(messages, kwargs):
"""A rule, not a model: keep transcript sentences that carry ids, amounts, plans, or requests."""
transcript = messages[-1]["content"]
keep = []
for line in transcript.splitlines():
role, _, text = line.partition(": ")
for sentence in re.split(r"(?<=[.?!])\s+", text):
if ID.search(sentence) or MONEY_OR_COUNT.search(sentence) or PLAN.search(sentence) or (role == "user" and ASK.search(sentence)):
keep.append(f"{role}: {sentence}")
return " ".join(keep)
extractive = ScriptedLLM(responder=extractive_summary) # plumbing stand-in, not an LLM
vague = ScriptedLLM(replies=["The customer raised billing, Slack, export, and automation questions; most were handled."])
full = conversation()
results = {
"no management": full,
"drop-oldest": drop_oldest(full, BUDGET),
"sliding window (8, pinned)": sliding_window(full, keep_last=8),
"sliding window (8, no pin)": sliding_window(full, keep_last=8, pin_first=False),
"importance-ranked": importance_ranked(full, BUDGET),
"summary (extractive rule)": compact_with_summary(full, extractive, keep_last=6),
"summary (vague scripted)": compact_with_summary(full, vague, keep_last=6),
}
print(f"Budget: {BUDGET} tokens. Full conversation: {count_messages(full)} tokens, {len(full)} messages.\n")
print(f"{'strategy':28}{'tokens':>7}{'msgs':>6} invoice id plan+seats customer's ask")
for name, kept in results.items():
facts = facts_kept(kept)
marks = "".join(f"{'kept' if ok else 'LOST':>13}" for ok in facts.values())
print(f"{name:28}{count_messages(kept):>7}{len(kept):>6}{marks}")
print("\nSummary message (extractive rule):", results["summary (extractive rule)"][1]["content"])
kept_idx = [full.index(m) for m in results["importance-ranked"]]
print("Importance-ranked kept message indexes:", kept_idx)Code explained
- In simple words: run each strategy on the same chat and check automatically whether the invoice id, the plan and seats, and the customer's request are still there.
- What happens: the budget is 300 tokens against a 558-token chat. Two summarizers exercise the summary path, and neither is a language model:
extractiveis aScriptedLLMwhose responder is a rule that keeps sentences containing ids, amounts, plans, or user requests;vagueis aScriptedLLMthat returns one generic sentence, the kind of summary a careless prompt or a small model produces.facts_kept()checks each result. - Comes out:
Budget: 300 tokens. Full conversation: 558 tokens, 24 messages.
strategy tokens msgs invoice id plan+seats customer's ask
no management 558 24 kept kept kept
drop-oldest 294 13 LOST LOST LOST
sliding window (8, pinned) 250 11 kept LOST kept
sliding window (8, no pin) 222 10 LOST LOST LOST
importance-ranked 295 13 kept kept kept
summary (extractive rule) 234 9 kept kept kept
summary (vague scripted) 207 9 LOST LOST LOST
Summary message (extractive rule): Summary of the earlier conversation: user: Hi, we were charged twice this month, invoice INV-2026-004512. user: Please refund the duplicate charge. user: We are on the Business plan with 40 seats, paid monthly.
Importance-ranked kept message indexes: [0, 1, 3, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23]Drop-oldest fits the budget and loses all three facts, because they were at the beginning. A plain sliding window loses them too. Pinning the first user message saves the invoice id and the request, but the plan and seat count came in turn 3 and are lost. Importance ranking keeps everything within budget by dropping the chatty middle (Slack, exports) and keeping messages 1 and 3 plus the recent turns. The extractive summary keeps all three facts in 234 tokens; the vague summary is shorter but loses everything, and the automatic check catches it.
Be honest about what this table shows. The importance scorer and the extractive rule were written by someone who knew which facts matter, so they are close to a best case; on real traffic you must decide and encode what matters, and you will miss things. The summary rows measure the plumbing and the checker, not summarization quality. To measure a real summarizer, pass supportdesk.llm.chat as summarize and rerun: compact_with_summary(conversation(), chat).
When to compact and what to preserve
Compaction costs a model call and can lose information, so do it deliberately:
- When: at a threshold, not every turn. A common pattern is to compact when the request passes about 70 to 80% of the usable input budget, which leaves room for the next few turns without compacting again immediately. Also compact at natural boundaries, such as when a sub-issue is resolved.
- What to preserve verbatim: identifiers (invoice, ticket, workspace ids), amounts and counts, the plan, the customer's open request and whether it is resolved, commitments the agent made ("a billing agent will review"), and constraints ("paid monthly", "only owners can cancel").
- What to drop: greetings, thanks, resolved side issues, and step-by-step instructions the customer has already confirmed worked.
- How to verify: run a fact check like
facts_kept()on every compaction, and fall back to a non-lossy strategy when it fails. The lab does exactly this.
| ituation | Use this | Why |
|---|---|---|
| Short chats, facts restated often | Drop-oldest or sliding window | Cheapest, predictable, nothing to tune |
| The opening message states the problem | Sliding window with the first user message pinned | Keeps the request; loses details given later |
| Known fact types (ids, amounts, plans) | Importance-ranked truncation | No extra call; kept all three facts here |
| Long chats with many relevant details | Summary compaction plus a fact check | Compresses the most; must be verified |
| Legal, billing, or audit needs | Keep the full transcript in storage, send a compacted view | The model's view can be lossy; your records must not be |
Long context versus retrieval
Compaction decides what to keep from the conversation. The same question arises for reference material: put the whole help center in every prompt (long context), or search it and include only the top matches (retrieval, built properly in Module 7)?
# examples/m02_long_vs_retrieval.py
"""Module 2: stuff the whole help center into the prompt, or retrieve the top 3 articles?"""
import math
from supportdesk.data import get_article, load_articles, load_tickets
from supportdesk.kb_search import KBSearch
from supportdesk.llm import Usage
from supportdesk.pricing import cost_usd
from supportdesk.tokens import count_tokens
def wilson(hits: int, n: int, z: float = 1.96) -> tuple[float, float]:
"""95% confidence interval for a proportion (works for small n)."""
p = hits / n
centre = (p + z * z / (2 * n)) / (1 + z * z / n)
half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
return centre - half, centre + half
articles = load_articles()
whole_kb = sum(count_tokens(f"[{a.id}] {a.title}\n{a.body}") for a in articles)
search = KBSearch()
answerable = [t for t in load_tickets() if t.gold["answerable"]]
retrieved_tokens, hits = [], 0
for t in answerable:
top = search.search(t.text, k=3)
retrieved_tokens.append(sum(count_tokens(get_article(h.article_id).body) for h in top))
hits += t.gold["kb_article"] in [h.article_id for h in top]
low, high = wilson(hits, len(answerable))
mean_retrieved = sum(retrieved_tokens) / len(retrieved_tokens)
print(f"Whole help center ({len(articles)} articles): {whole_kb:,} tokens in every prompt, gold article always present")
print(f"Top-3 retrieval: {mean_retrieved:.0f} tokens on average, gold article in top 3 for {hits}/{len(answerable)} "
f"answerable tickets ({hits / len(answerable):.0%}, 95% CI {low:.0%} to {high:.0%})")
for model in ("openai/gpt-oss-120b", "gemini-3.5-flash"):
stuffed = cost_usd(Usage(input_tokens=whole_kb + 700, output_tokens=250), model) * 1000
rag = cost_usd(Usage(input_tokens=int(mean_retrieved) + 700, output_tokens=250), model) * 1000
print(f" {model:20} per 1,000 replies: stuffed {stuffed:.3f} USD, retrieved {rag:.3f} USD")
per_article = whole_kb / len(articles)
print(f"At {per_article:.0f} tokens per article, a 131,072-token window holds about "
f"{int(131_072 * 0.9 / per_article):,} articles; a 1M window about {int(1_048_576 * 0.9 / per_article):,}.")Code explained
- In simple words: compare the cost and the risk of stuffing all 12 articles into every prompt against retrieving the top 3.
- What happens: we count the whole help center once, then for each of the 62 answerable tickets retrieve the top 3 articles with
KBSearchand record their size and whether the gold article is among them.wilson()gives a 95% confidence interval for the hit rate. Costs assume a 700-token prompt around the articles and a 250-token reply. - Comes out:
Whole help center (12 articles): 1,264 tokens in every prompt, gold article always present
Top-3 retrieval: 274 tokens on average, gold article in top 3 for 57/62 answerable tickets (92%, 95% CI 82% to 97%)
openai/gpt-oss-120b per 1,000 replies: stuffed 0.445 USD, retrieved 0.296 USD
gemini-3.5-flash per 1,000 replies: stuffed 5.196 USD, retrieved 3.711 USD
At 105 tokens per article, a 131,072-token window holds about 1,119 articles; a 1M window about 8,959.Retrieval sends 78% fewer article tokens but misses the right article for 5 of 62 tickets (the interval, 82% to 97%, is wide because n is small; Module 7 shows most misses are non-English tickets). Stuffing always includes the right article and, at 12 articles, costs only 40 to 50% more per reply. At about 105 tokens per article, the whole help center would fit in a 131K window up to roughly 1,100 articles. Fitting is not the only question, though: Part B's evidence says a model uses a 100,000-token context less reliably than a 1,000-token one.
| Situation | Use this | Why |
|---|---|---|
| Small, stable reference set (a few thousand tokens) | Long context, cached (Part D) | Nothing to miss, and caching makes the repeat cheap |
| Large or growing knowledge base | Retrieval | It will not fit, and irrelevant text dilutes attention |
| Answers need facts spread across many documents | Long context or retrieval with a larger k | Top-k can miss one of the needed pieces |
| Latency or cost sensitive at high volume | Retrieval | Fewer input tokens per call |
| Retrieval misses a known slice (for example non-English) | Hybrid: retrieval plus a pinned core set | Covers the slice without stuffing everything |