Part 3: Stopping, budgets, and recovery
Why a loop needs brakes
A model deciding when to stop is a model that can fail to stop. Real agents loop for mundane reasons: a search that never finds anything, a tool that keeps returning the same error, a model that "double checks" the same record forever, a context so long that the goal gets lost (Module 7's context rot). Each of those burns money on every turn. So the loop, not the model, owns termination:
- Max steps: a hard ceiling on iterations. Simple, and always on.
- No-progress detection: stop after N consecutive steps that produced no new observation. We fingerprint each observation (tool name plus result) and count steps that added no new fingerprint.
- Repeated-action detection: stop when the model issues the identical tool call (same name, same arguments) more than
repeat_limittimes. - Budget ceilings: stop before a call that would push the task over its token or dollar limit. The check happens before spending, using
prompt_tokensfor the input andoutput_reservefor the reply.
When a guard fires, the run ends with a status (max_steps, no_progress, repeated_action, budget_tokens, budget_usd) that the caller can route on, typically to a human queue with the trace attached.
Triggering every guard on purpose
This script runs eight scenarios through the same loop. Two small policies misbehave on purpose: StuckSearcher keeps rephrasing a search that finds nothing, and Looper asks for the same account again and again.
"""Module 8: every termination guard, the budget ceiling, replanning, and crash recovery, one scenario each.
All runs are driven by ScriptedLLM policies (not a model), so the numbers are real
counts of the loop's bookkeeping, not measurements of model behavior.
"""
from __future__ import annotations
import json
from pathlib import Path
from m08_agent import (T1001, AgentConfig, BillingStore, ScriptedApprover, SimulatedCrash, SupportPolicy,
history, make_tools, print_trace, run_agent)
from supportdesk.stand_in import ScriptedLLM
WORK = Path("runs/m08/guards")
WORK.mkdir(parents=True, exist_ok=True)
for old in WORK.glob("*"):
old.unlink()
class StuckSearcher(SupportPolicy):
"""Keeps rephrasing the customer's words; every query returns nothing new."""
QUERIES = ["charged twice", "paid twice", "twice charged", "charged 2x", "twice paid"]
def __call__(self, messages, kwargs):
n = len(history(messages))
return self.reply(messages, "Thought: try other words.", "search_kb",
{"query": self.QUERIES[n % len(self.QUERIES)]}, kwargs.get("tools"))
class Looper(SupportPolicy):
"""Calls get_account with the same arguments again and again."""
def __call__(self, messages, kwargs):
return self.reply(messages, "Thought: check the account.", "get_account",
{"email": "dana.k@northwind.example"}, kwargs.get("tools"))
def scenario(name: str, policy, *, config: AgentConfig | None = None, approve=(True,), flaky: int = 2,
crash_at_step: int | None = None, resume: bool = False) -> dict:
store = BillingStore(WORK / f"{name}_billing.json")
tools = make_tools(store, flaky_failures=flaky)
kwargs = dict(config=config, approver=ScriptedApprover(list(approve)),
trace_path=WORK / f"{name}.jsonl", checkpoint_path=WORK / f"{name}_ckpt.json",
sleep=lambda s: None)
try:
state = run_agent(T1001, "T-1001", tools, ScriptedLLM(responder=policy), crash_at_step=crash_at_step, **kwargs)
except SimulatedCrash as exc:
print(f" {name}: {exc}; refunds on disk: {len(json.loads((WORK / f'{name}_billing.json').read_text()))}")
if not resume:
raise
store = BillingStore(WORK / f"{name}_billing.json") # a fresh process reloads everything
tools = make_tools(store, flaky_failures=0)
kwargs["approver"] = ScriptedApprover([True]) # Maya approves the re-sent request
state = run_agent(T1001, "T-1001", tools, ScriptedLLM(responder=SupportPolicy()), **kwargs)
return {"scenario": name, "status": state.status, "steps": state.step, "llm_calls": state.llm_calls,
"tokens": state.input_tokens + state.output_tokens, "cost_usd": round(state.cost_usd, 6),
"refunds": len(store.refunds), "plan_version": state.plan.version}
rows = [
scenario("happy_path", SupportPolicy()),
scenario("max_steps", SupportPolicy(), config=AgentConfig(max_steps=4)),
scenario("no_progress", StuckSearcher()),
scenario("repeated_action", Looper()),
scenario("budget_usd", SupportPolicy(), config=AgentConfig(price_model="gemini-3.5-flash", max_usd=0.02)),
scenario("approval_denied", SupportPolicy(), approve=(False,)),
scenario("tool_down_replan", SupportPolicy(), flaky=10),
scenario("crash_resume", SupportPolicy(), crash_at_step=5, resume=True),
]
print()
print(f"{'scenario':18} {'status':16} {'steps':>5} {'calls':>5} {'tokens':>7} {'cost_usd':>9} {'refunds':>7} {'plan_v':>6}")
for r in rows:
print(f"{r['scenario']:18} {r['status']:16} {r['steps']:>5} {r['llm_calls']:>5} {r['tokens']:>7} "
f"{r['cost_usd']:>9.6f} {r['refunds']:>7} {r['plan_version']:>6}")
print("\nTrace of tool_down_replan (steps 4 onward):")
lines = [ln for ln in (WORK / "tool_down_replan.jsonl").read_text().splitlines() if json.loads(ln)["step"] >= 4]
(WORK / "replan_tail.jsonl").write_text("\n".join(lines) + "\n")
print_trace(WORK / "replan_tail.jsonl")
print("\nTrace of crash_resume (steps 5 onward):")
lines = [ln for ln in (WORK / "crash_resume.jsonl").read_text().splitlines() if json.loads(ln)["step"] >= 5]
(WORK / "crash_tail.jsonl").write_text("\n".join(lines) + "\n")
print_trace(WORK / "crash_tail.jsonl")Code explained
- In simple words: one run per guard, each set up so that exactly that guard (or recovery path) fires, then a summary table and two traces.
- What happens:
scenariobuilds fresh tools and a fresh billing file, runs the agent with a trace and a checkpoint, and returns the totals.sleep=lambda s: Noneskips the real backoff waits so the script is fast; the backoff values still appear in the trace.max_stepssetsmax_steps=4.no_progressusesStuckSearcher, whose queries all return[].repeated_actionusesLooper.budget_usdprices the run asgemini-3.5-flash(1.50 USD input, 9.00 USD output per million tokens inpricing.py) with a 0.02 USD ceiling.approval_deniedhas Maya reject the refund.tool_down_replanmakesget_invoicefail 10 times, more than the 3 attempts allowed.crash_resumecrashes at step 5, after the refund ran but before the checkpoint was written. Theexceptbranch then does what a restarted worker would: reload the billing store from disk, rebuild the tools, and callrun_agentagain with the same checkpoint path, which resumes from the last checkpoint.
- Comes out: (ScriptedLLM-driven; counts are real)
crash_resume: crashed during step 5; refunds on disk: 1
scenario status steps calls tokens cost_usd refunds plan_v
happy_path done 8 8 9739 0.001631 1 3
max_steps max_steps 4 4 3811 0.000641 0 2
no_progress no_progress 4 4 3428 0.000548 0 0
repeated_action repeated_action 2 3 2664 0.000431 0 0
budget_usd budget_usd 7 7 8084 0.014571 1 3
approval_denied done 8 8 9628 0.001597 0 3
tool_down_replan done 7 7 7791 0.001291 0 3
crash_resume done 8 8 9753 0.001633 1 3
Trace of tool_down_replan (steps 4 onward):
[4] THINK Thought: fetch the invoice named in the ticket to see the charges.
[4] ACT get_invoice({"invoice_id": "INV-2026-004512"})
[4] RETRY get_invoice attempt 1: 503 billing API timeout (call 1); backoff 0.061 s
[4] RETRY get_invoice attempt 2: 503 billing API timeout (call 2); backoff 0.119 s
[4] RETRY get_invoice attempt 3: 503 billing API timeout (call 3); giving up
[4] REPLAN_REQUESTED {'old_plan': ['get_account', 'get_invoice', 'issue_refund', 'add_internal_note', 'reply']}
[5] THINK Thought: billing API is down after retries, so I cannot verify the charge. Hand it to a person.
[5] ACT escalate_to_human({"ticket_id": "T-1001", "reason": "get_invoice unavailable; duplicate charge unv)
[5] OBS {"queued_for": "billing-team", "position": 1}
[6] THINK Thanks for flagging this. A member of our billing team is checking the duplicate charge and will
[7] STOP {'status': 'done'}
Trace of crash_resume (steps 5 onward):
[5] THINK Thought: two charges of 288.0 on 2026-09-03; overpaid 288.0. Refund the second one.
[5] ACT issue_refund({"invoice_id": "INV-2026-004512", "charge_id": "ch_7Hq2", "amount_usd": 288.0, ")
[5] GATE issue_refund approved=True by human:maya
[5] OBS {"refund_id": "RF-00001", "invoice_id": "INV-2026-004512", "charge_id"
[5] CRASH {'note': 'simulated crash before checkpoint'}
[5] RESUME {'from_checkpoint': 'crash_resume_ckpt.json', 'llm_calls': 5}
[5] THINK Thought: two charges of 288.0 on 2026-09-03; overpaid 288.0. Refund the second one.
[5] ACT issue_refund({"invoice_id": "INV-2026-004512", "charge_id": "ch_7Hq2", "amount_usd": 288.0, ")
[5] GATE issue_refund approved=True by human:maya
[5] OBS {"already_refunded": true, "refund_id": "RF-00001", "invoice_id": "INV
[6] THINK Thought: record what I did for the team.
[6] ACT add_internal_note({"ticket_id": "T-1001", "note": "Refunded duplicate ch_7Hq2 (288.0 USD) as RF-00)
[6] OBS {"saved": true, "note_count": 1}
[7] THINK Hi, sorry about the double charge. We refunded the duplicate payment of 288.00 USD on invoice IN
[8] STOP {'status': 'done'}How to read the table, row by row:
- happy_path: 8 steps, one refund. The baseline.
- max_steps: stopped after 4 model calls with no refund. A hard cap is crude but it bounds the worst case: the most this task can ever cost is 4 calls.
- no_progress: the first empty search was new information (a new fingerprint); the next three returned the same
[], so after three stale steps the loop stopped. Without this guard,StuckSearcherwould spin untilmax_steps. - repeated_action: the third identical
get_accountcall was refused before it ran. The model made 3 calls but only 2 steps completed. - budget_usd: this one is worth studying. The full run at Gemini prices would cost 0.0174 USD (Part 4), under the 0.02 ceiling, yet the guard stopped it after 7 calls at 0.0146 USD. Why? Before call 8 the loop projected about 1,600 input tokens plus a 400-token output reserve. At 9 USD per million output tokens, the reserve alone is 0.0036 USD, and 0.0146 plus the projection crosses 0.02. The actual final reply was far shorter than 400 tokens. A budget guard with a reserve is conservative by design; size the reserve from the 95th percentile of real reply lengths, and when it trips, hand off to a person rather than leaving the ticket half done. Here the refund went through but the customer reply was never written, which is exactly the state you do not want to leave silently.
- approval_denied: Maya rejected, the policy escalated instead of retrying the refund, zero refunds.
- tool_down_replan: three attempts with backoff, then a replan request, then escalation. The agent did not guess at the charges it could not see.
- crash_resume: the process died right after the refund. On resume, the loop replayed step 5 from the checkpoint, the model asked for the same refund, Maya approved again, and the billing store answered
already_refunded: truewith the original RF-00001. One refund on disk, not two.
Notice one more thing in crash_resume: its token total (9,753) is barely above the happy path, but the model call made just before the crash is missing from the state's totals, because that state was never saved. The trace file still has it. When you reconcile spend, reconcile from traces or provider bills, not from checkpointed counters.
| Situation | Use this | Why |
|---|---|---|
| Every agent, always | max_steps plus a dollar ceiling | Bounds the worst case even when every other guard misses |
| Searches or reads that can come back empty or identical | No-progress detection on observation fingerprints | Catches loops where arguments change but nothing new is learned |
| A model that re-requests the same call | Repeated-action detection | Cheap, exact, and catches the most common loop |
| A cost-sensitive queue | Per-task token and dollar budgets checked before each call, plus a batch ceiling | Stops overspend before it happens, not after the bill |
| A guard fires mid-task | Route to a human with the trace and the checkpoint | A half-done task needs a person more than a retry |