Part D: Self-correction
Self-critique and revision loops
A critique-and-revise loop asks the model to review its own answer and fix it: answer, critique, revise, repeat. It is appealing because it looks like what a careful person does. The question is where the information to fix a mistake comes from. If the critique is the same model reading the same prompt, it has no new information, only a second chance to sample, and a second chance can move a right answer as easily as a wrong one.
Huang et al., "Large Language Models Cannot Self-Correct Reasoning Yet" (ICLR 2024), tested exactly this intrinsic self-correction (no external feedback). With GPT-3.5, accuracy went from 75.9 to 75.1 to 74.7 percent on GSM8K over two rounds of self-correction, and from 75.8 to 38.1 (round 1) and 41.8 (round 2) on CommonSenseQA. On GSM8K the model kept its first answer 74.7 percent of the time, and among the changes it was more likely to turn a correct answer into an incorrect one than the reverse. Earlier reports of self-correction gains often relied on oracle labels telling the model when to stop, which is external information.
Verifier models and external checks, measured
The script builds the loop above and runs it three ways: with no external signal, with the billing-record verifier available for only 15 of the 20 scenarios, and with the verifier for all 20. The model is a ScriptedLLM whose behavior follows stated assumptions, chosen in the direction Huang et al. observed: first answers are right with probability 0.9 on easy cases and 0.45 on traps; when asked to critique, it changes a right answer to a wrong one 30 percent of the time and fixes a wrong answer 20 percent of the time. The loop code is real; the accuracies measure what that loop does to a model with those habits.
examples/m05_self_correct.py
"""Part D: a critique-and-revise loop, with and without an external verifier.
The model is a SCRIPTED stand-in whose behavior follows stated assumptions:
first answers are right with probability 0.9 (easy cases) or 0.45 (trap cases);
when asked to critique, it changes a right answer to a wrong one 30 percent of
the time and fixes a wrong answer 20 percent of the time (Huang et al. 2023 saw
more right-to-wrong than wrong-to-right changes). The loop code is real; the
accuracies measure the loop's arithmetic under those assumptions, not a model.
"""
import random
from examples.m05_helpers import CASES, LABELS, case_from_messages, decide, parse_decision, prompt_scaffold, shortcut_baseline
from supportdesk.stand_in import ScriptedLLM
CRITIQUE = ("Review your answer above against the policy. If it is correct, reply exactly NO_ISSUES. "
"Otherwise explain the mistake and end with a corrected line: DECISION: <decision>")
def make_stand_in(rng: random.Random) -> ScriptedLLM:
def responder(messages, kwargs):
case = case_from_messages(messages[:1])
wrong = [label for label in LABELS if label != case.gold]
if len(messages) == 1: # first answer
trap = shortcut_baseline(case.text) != case.gold
ok = rng.random() < (0.45 if trap else 0.9)
return f"DECISION: {case.gold if ok else rng.choice(wrong)}"
current = parse_decision(messages[-2]["content"])
roll = rng.random()
if current == case.gold:
return f"On reflection the policy points elsewhere.\nDECISION: {rng.choice(wrong)}" if roll < 0.30 else "NO_ISSUES"
if roll < 0.20:
return f"I misapplied a rule.\nDECISION: {case.gold}"
if roll < 0.45:
return f"Another rule applies.\nDECISION: {rng.choice([w for w in wrong if w != current])}"
return "NO_ISSUES"
return ScriptedLLM(responder=responder)
def verifier(case, label: str | None) -> bool | None:
"""External check against the billing record. Returns None when the record is unavailable."""
if case.id in NO_RECORD:
return None
return label == decide(case.facts)
NO_RECORD: set[str] = set()
def critique_and_revise(chat_fn, case, max_rounds: int, use_verifier: bool) -> tuple[list[str | None], list[int]]:
"""Run up to max_rounds of critique. Returns the answer after each round (index 0 = first answer)
and the cumulative number of model calls after each round."""
messages = prompt_scaffold(case)
reply = chat_fn(messages)
calls, current = 1, parse_decision(reply.text)
history, spent = [current], [calls]
messages = messages + [reply.as_message()]
for _ in range(max_rounds):
if use_verifier and verifier(case, current):
history.append(current) # verified: nothing to fix, no call spent
spent.append(calls)
continue
critique = chat_fn(messages + [{"role": "user", "content": CRITIQUE}])
calls += 1
proposed = parse_decision(critique.text) if "NO_ISSUES" not in critique.text else current
if use_verifier and verifier(case, proposed) is False:
proposed = current # the check rejects the revision; keep what we had
if proposed != current:
messages = messages + [{"role": "user", "content": CRITIQUE}, critique.as_message()]
current = proposed
history.append(current)
spent.append(calls)
return history, spent
def table(use_verifier: bool, repeats: int = 300, rounds: int = 4, seed: int = 3) -> list[tuple[float, float]]:
"""Accuracy and average calls per case after 0..rounds rounds, over repeats x 20 cases."""
llm = make_stand_in(random.Random(seed))
right = [0] * (rounds + 1)
calls = [0] * (rounds + 1)
for _ in range(repeats):
for case in CASES:
history, spent = critique_and_revise(llm, case, rounds, use_verifier)
for r in range(rounds + 1):
right[r] += history[r] == case.gold
calls[r] += spent[r]
total = repeats * len(CASES)
return [(right[r] / total, calls[r] / total) for r in range(rounds + 1)]
if __name__ == "__main__":
f, g = 0.30, 0.20
print(f"No-stop Markov estimate: accuracy drifts toward g/(f+g) = {g / (f + g):.2f} whatever it starts at\n")
runs = {"no external signal": table(False)}
NO_RECORD.update({"R03", "R07", "R12", "R15", "R18"}) # 5 of 20 cases with no billing record reachable
runs["verifier, 15/20 records"] = table(True)
NO_RECORD.clear()
runs["verifier, all records"] = table(True)
print(f"{'round':<6}" + "".join(f"{name:>30}" for name in runs))
for r in range(5):
cells = "".join(f"{f'{acc:.3f} ({calls:.2f} calls)':>30}" for acc, calls in (runs[n][r] for n in runs))
print(f"{r:<6}{cells}")
Code explained
- In simple words: let a scripted model second-guess itself for up to four rounds, with and without a check that knows the right answer, and record accuracy and calls after each round.
- What happens:
make_stand_inreturns aScriptedLLMwith the stated flip rates. The first message decides the first answer; later calls read the current answer from the conversation and flip it or not.verifierrecomputes the decision from the billing record, and returnsNonewhen the record is not reachable (NO_RECORD), which is how partial coverage is modelled.critique_and_reviseis the loop. A verified answer costs no more calls. A revision the verifier rejects is thrown away (proposed = current). A revision without a check is accepted, as a pure self-critique loop would.max_roundsis the iteration ceiling: the loop cannot run forever even if the model never says NO_ISSUES. It returns the answer and the cumulative call count after every round.tablerepeats the 20 cases 300 times and averages.- The first printed line is a two-state Markov chain estimate: if each round keeps a right answer with probability 0.7 and fixes a wrong one with probability 0.2, accuracy drifts toward 0.2 / (0.3 + 0.2) = 0.40 regardless of where it starts.
- Comes out:
- No external signal: accuracy falls every round, 0.706, 0.558, 0.482, 0.439, 0.428, heading for the 0.40 the Markov estimate predicts, while calls rise to 5 per case. The loop is actively harmful.
- Verifier for all records: accuracy rises from 0.685 to 0.869 after four rounds and can never fall, because only verified fixes are accepted; it spends only 1.94 calls per case, since verified cases stop.
- Verifier for 15 of 20 records: in between (0.781). On the 5 unverifiable cases the loop behaves like the no-signal loop, eroding their accuracy while the verified cases improve.
Why self-correction without an external signal fails
The table makes the argument concrete. A critique round is a coin flip weighted by the model's habits, not by the truth. If the model is more willing to abandon a right answer than to find a wrong one (which is what Huang et al. saw), every round loses accuracy. Signals that do carry new information include: a verifier (tests, a schema, a database lookup, a rule function), a tool result, retrieved documents the first answer did not see, a different and stronger model as the critic, or a human. Without one of those, spend the call on a second independent sample and a vote instead, or on nothing.
| Situation | Use this | Why |
|---|---|---|
| An exact check exists | Revise only when the check fails, accept only checked fixes | Accuracy can only go up; calls are spent only on failures |
| A partial check exists | Loop on checkable cases; send uncheckable ones to review | The no-signal part of the loop erodes accuracy |
| No check at all | No self-critique loop; one answer, or a vote, plus human review | Measured drift toward the model's flip-rate equilibrium |
| The critique adds new information (retrieval, tool output) | One revision round with that information in the prompt | The revision has something to act on |
Iteration ceilings and diminishing returns
Reading the verifier column round by round shows how returns shrink:
| Rounds (ceiling) | Accuracy, verifier on all records | Gain from this round | Calls per case | Extra calls this round | Gain per extra call |
|---|---|---|---|---|---|
| 0 | 0.685 | 1.00 | |||
| 1 | 0.749 | +0.064 | 1.31 | 0.31 | 0.21 |
| 2 | 0.795 | +0.046 | 1.57 | 0.26 | 0.18 |
| 3 | 0.836 | +0.041 | 1.77 | 0.20 | 0.21 |
| 4 | 0.869 | +0.033 | 1.94 | 0.17 | 0.19 |
Each round gains less accuracy than the one before (+0.064 down to +0.033), because each round works on the cases that survived the previous one, which are the hard ones. The gain per extra call stays near 0.2 only because verified cases stop spending: with a fixed-cost loop (every case critiqued every round), the calls per round would be flat at 1.0 and the gain per call would fall with the gain. Latency does not stop growing either: a case still unverified after round 3 has waited for four sequential calls.
Two practical rules follow. Always set a ceiling (here 2 to 4 rounds) and route what is still unverified after the ceiling to a person, instead of letting the loop run. And pick the ceiling from a table like this one, measured on your dev set with your model, not from intuition: stop where one more round's gain is worth less than a human review of the cases it would touch.