Last lesson ended on one line: stopping a loop is not the same as checking it. So who checks it? Right now, nobody. Every check this loop has runs inside the run that made the change. This lesson spawns one that doesn't, and hands it a list.
The checker you already have
You do have a checker. /goal runs one after every turn: a small model reads the
transcript and says done, or not done. This is the transcript-checker from
Lesson 1.2.
Look at what it sees: what the run said happened. Its own account of its own work. The documentation is blunt about the limit. It can't "run commands or read files independently."
Lesson 1.3 already gave you the rule that applies here: store what you measured, not what you decided. A verdict isn't evidence.
So if the transcript-checker is the only check a loop has, you've built a rubber-stamp verifier. It isn't broken. It just isn't a verifier. It is reading the maker's account, and an account is not evidence.
Two ways to spawn a checker
There are two ways to get a second opinion, and only one of them helps.
| 01: the same session, asked to check itself | 02: a sub-agent invocation | |
|---|---|---|
| Context | the maker's, entire: every step, and every reason it gave itself | its own window, empty when it starts |
| Evidence | the transcript it wrote itself | whatever it runs and reads for itself |
| Verdict | agrees, almost always | whatever the output says |
Ask the same session to check its own work and it has the whole transcript already. It agrees. So would you. That isn't evidence; it's a story that ends well.
A sub-agent is different. Claude Code gives it its own context window. It starts cold and has to go and get the facts.
That's maker-checker: two roles, never the same invocation. Whatever made the change doesn't get to grade it. The verifier is told, in writing, to be adversarial and to trust the tests over anyone's account of them.
Dimension blindness
A maker checks the dimension it was already thinking about. Fix a desktop layout, ship it, and it breaks on mobile, because nobody looked at mobile. If the checker inherits the maker's context, it inherits that blind spot with it.
The checklist
What do you send the checker to do? For this repo, four items, and each one is a command.
1. The failing test, on its own, by name.
pytest tests/test_pricing.py::test_apply_discount_larger_price
Green isn't enough. You need exit 0 and a non-zero collected count. This is
Lesson 1.2's trap: pytest exits 5 when it collects nothing, so a mistyped path
passes a check that never ran.
2. The whole suite.
pytest -v
This is what CI runs, and the only canonical test command this repo has. The fix working was never the question. Regressions live where nobody was looking.
3. Lint, format, type-check.
This repo has no lint, format or type-check config. So the box comes back empty, and that is the finding, not a free pass. A repo with no static gate hands the checker one fewer independent signal, and you should know that before you trust it.
4. The diff.
git diff --name-only master...HEAD
List every file the branch touched, and ask why each one is there. Three dots, not two: two would also list what master did after the branch diverged.
I'd run item 4 first. A fix that quietly edits a test is the one you most want to catch.
The checker, as a file
A sub-agent is defined in a file, .claude/agents/verifier.md. It has the same
shape as the skill from Lesson 2.3: frontmatter on top, instructions underneath.
---
name: verifier
description: Verify a fix on a branch. Run the checklist and report what each command printed.
tools: Read, Grep, Bash
---
You are verifying someone else's work on a branch you did not write.
Be adversarial. Trust the test output over any account of it.
Run every item on the checklist. Do not skip one because an earlier one passed.
Report the command, its exit code, and the last line of its output — for every item. Report what you ran, not what you concluded.
A name, a description, and the tools it may use: read and run, not edit. A checker that can edit can fix what it finds, and the moment it does, it is a maker again, grading its own work.
The last line is the whole module: report what you ran, not what you concluded. "Here's the green run," not "tests pass." A verifier that can't produce the evidence hasn't verified anything.
You call it by name when a fix lands, from the maker's session:
> use the verifier subagent on fix/apply-discount
Hand it the branch name, and nothing else. It reads none of the maker's session.
What that buys, and what it costs
| What you bought | What it costs |
|---|---|
| A second opinion that had to earn it, from output it produced, not from a story it was handed. | Another context window, and the time it takes to actually run the list. |
| A shot at the dimension the maker was structurally unable to check. | A list that rots. The checklist is a project artifact like any other; 2.3's skill goes stale the same way, and nothing checks either of them for you. |
And be honest about the limit: the checker can be wrong too. It just can't be wrong the same way the maker was. That is the only kind of independence a second model can give you.
After this lesson
You will be able to:
- explain why
/goal's transcript-checker, as a loop's only check, is a rubber-stamp verifier; - choose a sub-agent invocation over a self-check, and keep the maker and checker roles separate;
- name dimension blindness, and why inheriting the maker's context inherits it;
- write a verification checklist for this repo where every item is a command;
- define a read-and-run
verifiersub-agent that reports evidence, not conclusions.
What comes next
The checker can be wrong too, and sometimes the loop is stuck anyway. In Lesson 3.2: the failure modes a loop falls into, and the tasks a loop should never own.
Companion article
Every hook, script, and command in this course, written up and executed: How to build an agent loop in Claude Code