Last lesson ended on one line: the stopping rule matters as much as the goal. What it left out is where that rule lives. That turns out to decide whether it's real.
Two goals, the same intent
Here is the goal from Lesson 1.1, and a version of it you can actually run.
/goal fix failing tests until they pass.
/goal Fix the failing tests in tests/test_pricing.py
Done when pytest tests/test_pricing.py exits 0
and collects a non-zero number of tests,
and the full suite has no new failures.
The first one is fine for showing the shape of a loop. But "until they pass": checked by whom? Nothing outside the agent ever checks. Done is whenever the agent says it is.
The second is the same intent, made mechanical. The agent's word is a judgment. Exit 0 is a fact, and it's the same fact every run.
The "collects a non-zero number of tests" clause is not padding. "No failures"
is also true of a run that never ran: pytest exits 5 when it collects
nothing. The exit codes are 0 passed, 1 failed, 2 interrupted,
3 internal error, 4 usage error, 5 nothing collected.
Every way this loop is allowed to end
Success is one way a loop ends, and it's the one you'll get least often. So the goal carries its other stops too, written inside the goal's own text:
/goal Fix the failing tests in tests/test_pricing.py
Done when pytest tests/test_pricing.py exits 0, having collected >0 tests,
and the full suite has no new failures.
Stop after 12 iterations.
Stop at $5.
Stop and ask on no progress.
"No progress" is the one people leave vague. Define it before the loop runs:
No progress = the same failing assertion and an unchanged
git diffhash, for two consecutive iterations.
The shape every /goal in this course follows
Read it as slots:
When [trigger], inspect [fresh input]. Take one bounded action. Check [mechanical condition] the same way every time. Stop on success; stop at [iteration cap] (hook-enforced); stop at [budget] (prose only); stop and ask on [no progress].
Filled in for this course's running example:
| Slot | Value |
|---|---|
| Trigger | the suite goes red |
| Fresh input | this run's pytest output |
| One bounded action | one edit and one test run, not a plan |
| Mechanical condition | pytest exits 0, tests collected > 0 |
| Iteration cap | 12 iterations, hook-enforced |
| Budget | $5, prose only |
| No progress | same assertion + same diff hash, twice |
One stop is marked hook-enforced, and the other prose only. That difference is the point of this lesson: one is a cap, the other is a wish. Module 4 hands you this template again, blank, for a loop of your own.
Two stop checks, and only one is a ceiling
/goal already has a stop check built in. After each turn, a small model reads
the transcript and judges whether the goal is done. Call it the
transcript-checker. It sees what the run said happened. According to the
docs, it doesn't "run commands or read files independently." It's useful, but
it is not a ceiling.
You're probably thinking: "stop after 12 iterations" is right there in the goal. Isn't that a cap? It looks like one. But it's a sentence in a prompt, read by the same model that wants to finish the job.
A hook is different. Claude Code normally asks before it runs a command, and
there are flags that turn that asking off. A PreToolUse hook runs ahead of all
of it, and if it says no, the call doesn't happen. In the docs' own words,
PreToolUse hooks "fire before any permission-mode check, in every permission
mode", and a hook denying a call "blocks the tool even in bypassPermissions
mode or with --dangerously-skip-permissions."
The transcript-checker reads the story. The hook holds the door.
The ceiling, in code
Register a PreToolUse hook in your project's .claude/settings.json:
{
"hooks": {
"PreToolUse": [{
"matcher": "Bash|PowerShell",
"hooks": [{
"type": "command",
"command": "python \"${CLAUDE_PROJECT_DIR}/.loop/guard.py\"",
"timeout": 30
}]
}]
}
}
The outer hooks is the config block. The inner hooks is the list of things to
run for this matcher. The matcher is a regex: on Windows the agent's shell tool
is PowerShell, so a Bash-only cap would never fire there.
Then the script, .loop/guard.py:
import json, sys
from pathlib import Path
CAP = 40
STATE = Path(__file__).with_name("state.json") # .loop/state.json
SEED = {"run": {"session_id": "", "bash_calls": 0}}
payload = json.load(sys.stdin)
try: # missing, corrupt, or half-written by the agent
state = json.loads(STATE.read_text())
run = state["run"]
except Exception: # fail to a known cap, never to none
state, run = SEED, SEED["run"]
# new session -> the cap resets itself
if run.get("session_id") != payload.get("session_id"):
run.update(session_id=payload.get("session_id"), bash_calls=0)
run["bash_calls"] = run.get("bash_calls", 0) + 1
STATE.parent.mkdir(parents=True, exist_ok=True)
STATE.write_text(json.dumps(state, indent=2))
if run["bash_calls"] > CAP:
print(f"call {run['bash_calls']} exceeds cap {CAP}", file=sys.stderr)
raise SystemExit(2) # 2 blocks the call; stderr goes back to the agent
raise SystemExit(0)
It reads the counter, adds one and writes it back. Over forty, it exits 2, which
blocks the call. The hook payload carries a session_id, so a new session
starts a fresh count, with no startup step to forget.
It's in Python rather than shell because you already have Python (you run
pytest with it), and it avoids needing jq on Windows.
What that ceiling does, and doesn't
There is already a native cap, but on a different hook. A Stop hook fires
when the agent tries to finish, and it can push the agent to keep going. Claude
Code overrides one after 8 consecutive blocks without progress
(CLAUDE_CODE_STOP_HOOK_BLOCK_CAP raises that number). A runaway loop is busy,
editing and running pytest, so it is making progress and never trips that cap.
The units differ. The goal says 12 iterations, but the hook counts 40 shell calls. A hook can't see a turn, only a tool call. At roughly three commands per iteration, twelve iterations is about 36 calls, so forty is the backstop just past it.
Shell calls are your proxy for iterations. Set the cap high, as a backstop. A hard ceiling on something measurable beats a soft one on something that isn't.
This is the iteration cap only. A spend cap would need a running-cost
number, and a PreToolUse hook is never handed one. That is why the template
marks the budget prose only.
After this lesson
You will be able to:
- turn a vague goal into a mechanical one, including the
pytestexit-5 trap; - write every stop a loop needs: success, an iteration cap, a budget, and a "no progress" rule defined before the loop runs;
- tell a prose stop condition, which the model interprets, from a hook-enforced ceiling, which it can't argue past;
- wire a working
PreToolUseiteration cap.
What comes next
The hook writes its counter to .loop/state.json. In Lesson 1.3: where that
file lives once the session ends, and what must never go in it.
Companion article
The same hook, with a progress detector and flaky-test triage, built and executed step by step: How to build an agent loop in Claude Code