Learn AI Engineering
Lessons
← Loop Engineering

Write a /goal, and the Hook That Enforces It

A /goal is only a loop if its success condition is mechanical, and the ceiling that halts a runaway has to live outside the transcript, in a hook.

Last lesson ended on one line: the stopping rule matters as much as the goal. What it left out is where that rule lives. That turns out to decide whether it's real.

Two goals, the same intent

Here is the goal from Lesson 1.1, and a version of it you can actually run.

/goal fix failing tests until they pass.
/goal Fix the failing tests in tests/test_pricing.py
Done when pytest tests/test_pricing.py exits 0
and collects a non-zero number of tests,
and the full suite has no new failures.

The first one is fine for showing the shape of a loop. But "until they pass": checked by whom? Nothing outside the agent ever checks. Done is whenever the agent says it is.

The second is the same intent, made mechanical. The agent's word is a judgment. Exit 0 is a fact, and it's the same fact every run.

The "collects a non-zero number of tests" clause is not padding. "No failures" is also true of a run that never ran: pytest exits 5 when it collects nothing. The exit codes are 0 passed, 1 failed, 2 interrupted, 3 internal error, 4 usage error, 5 nothing collected.

Every way this loop is allowed to end

Success is one way a loop ends, and it's the one you'll get least often. So the goal carries its other stops too, written inside the goal's own text:

/goal Fix the failing tests in tests/test_pricing.py
Done when pytest tests/test_pricing.py exits 0, having collected >0 tests,
and the full suite has no new failures.
Stop after 12 iterations.
Stop at $5.
Stop and ask on no progress.

"No progress" is the one people leave vague. Define it before the loop runs:

No progress = the same failing assertion and an unchanged git diff hash, for two consecutive iterations.

The shape every /goal in this course follows

Read it as slots:

When [trigger], inspect [fresh input]. Take one bounded action. Check [mechanical condition] the same way every time. Stop on success; stop at [iteration cap] (hook-enforced); stop at [budget] (prose only); stop and ask on [no progress].

Filled in for this course's running example:

SlotValue
Triggerthe suite goes red
Fresh inputthis run's pytest output
One bounded actionone edit and one test run, not a plan
Mechanical conditionpytest exits 0, tests collected > 0
Iteration cap12 iterations, hook-enforced
Budget$5, prose only
No progresssame assertion + same diff hash, twice

One stop is marked hook-enforced, and the other prose only. That difference is the point of this lesson: one is a cap, the other is a wish. Module 4 hands you this template again, blank, for a loop of your own.

Two stop checks, and only one is a ceiling

/goal already has a stop check built in. After each turn, a small model reads the transcript and judges whether the goal is done. Call it the transcript-checker. It sees what the run said happened. According to the docs, it doesn't "run commands or read files independently." It's useful, but it is not a ceiling.

You're probably thinking: "stop after 12 iterations" is right there in the goal. Isn't that a cap? It looks like one. But it's a sentence in a prompt, read by the same model that wants to finish the job.

A hook is different. Claude Code normally asks before it runs a command, and there are flags that turn that asking off. A PreToolUse hook runs ahead of all of it, and if it says no, the call doesn't happen. In the docs' own words, PreToolUse hooks "fire before any permission-mode check, in every permission mode", and a hook denying a call "blocks the tool even in bypassPermissions mode or with --dangerously-skip-permissions."

The transcript-checker reads the story. The hook holds the door.

The ceiling, in code

Register a PreToolUse hook in your project's .claude/settings.json:

{
  "hooks": {
    "PreToolUse": [{
      "matcher": "Bash|PowerShell",
      "hooks": [{
        "type": "command",
        "command": "python \"${CLAUDE_PROJECT_DIR}/.loop/guard.py\"",
        "timeout": 30
      }]
    }]
  }
}

The outer hooks is the config block. The inner hooks is the list of things to run for this matcher. The matcher is a regex: on Windows the agent's shell tool is PowerShell, so a Bash-only cap would never fire there.

Then the script, .loop/guard.py:

import json, sys
from pathlib import Path

CAP   = 40
STATE = Path(__file__).with_name("state.json")  # .loop/state.json
SEED  = {"run": {"session_id": "", "bash_calls": 0}}

payload = json.load(sys.stdin)
try:  # missing, corrupt, or half-written by the agent
    state = json.loads(STATE.read_text())
    run   = state["run"]
except Exception:  # fail to a known cap, never to none
    state, run = SEED, SEED["run"]

# new session -> the cap resets itself
if run.get("session_id") != payload.get("session_id"):
    run.update(session_id=payload.get("session_id"), bash_calls=0)

run["bash_calls"] = run.get("bash_calls", 0) + 1
STATE.parent.mkdir(parents=True, exist_ok=True)
STATE.write_text(json.dumps(state, indent=2))

if run["bash_calls"] > CAP:
    print(f"call {run['bash_calls']} exceeds cap {CAP}", file=sys.stderr)
    raise SystemExit(2)  # 2 blocks the call; stderr goes back to the agent
raise SystemExit(0)

It reads the counter, adds one and writes it back. Over forty, it exits 2, which blocks the call. The hook payload carries a session_id, so a new session starts a fresh count, with no startup step to forget.

It's in Python rather than shell because you already have Python (you run pytest with it), and it avoids needing jq on Windows.

What that ceiling does, and doesn't

There is already a native cap, but on a different hook. A Stop hook fires when the agent tries to finish, and it can push the agent to keep going. Claude Code overrides one after 8 consecutive blocks without progress (CLAUDE_CODE_STOP_HOOK_BLOCK_CAP raises that number). A runaway loop is busy, editing and running pytest, so it is making progress and never trips that cap.

The units differ. The goal says 12 iterations, but the hook counts 40 shell calls. A hook can't see a turn, only a tool call. At roughly three commands per iteration, twelve iterations is about 36 calls, so forty is the backstop just past it.

Shell calls are your proxy for iterations. Set the cap high, as a backstop. A hard ceiling on something measurable beats a soft one on something that isn't.

This is the iteration cap only. A spend cap would need a running-cost number, and a PreToolUse hook is never handed one. That is why the template marks the budget prose only.

After this lesson

You will be able to:

  • turn a vague goal into a mechanical one, including the pytest exit-5 trap;
  • write every stop a loop needs: success, an iteration cap, a budget, and a "no progress" rule defined before the loop runs;
  • tell a prose stop condition, which the model interprets, from a hook-enforced ceiling, which it can't argue past;
  • wire a working PreToolUse iteration cap.

What comes next

The hook writes its counter to .loop/state.json. In Lesson 1.3: where that file lives once the session ends, and what must never go in it.

Companion article

The same hook, with a progress detector and flaky-test triage, built and executed step by step: How to build an agent loop in Claude Code