Three more loops, and not one of them needs a primitive you haven't already built. Same commands, same worktree, same maker and checker, same five rungs.
What changes is where you point them, and how far up the ladder each one is allowed to climb. For one of them, the answer is not very far at all.
Every pattern here is measured against the rule from Lesson 3.3:
The rung you're on is the one you have evidence for, not the one the tooling can reach.
Pattern 1: the test-stabilizer
You have this problem already; Module 2 made you triage it. Now fix it. Typed in a session:
/goal fix the flaky suite until 5 consecutive green runs. Stop after 15 turns.
Five consecutive green runs is the check. One red run sends the counter back to zero. It's mechanical, and a machine can tell you when it's done.
Read the second sentence again. "Stop after 15 turns" is inside the command, so
it is prose: a request the model honours, not a ceiling. The ceiling is the
PreToolUse hook you wrote in Module 1, reused here as-is. It fires before any
permission check and can't be argued past.
| Test-stabilizer | |
|---|---|
| Loop type | goal-based |
| Target | the flaky suite, not one failing test |
| Primitive | /goal, typed in a session |
| The check | 5 consecutive green runs |
| Isolation | a worktree per attempt, unchanged |
| Checker | a sub-agent that isn't the maker, unchanged |
| Starts on | Manual: you type it and watch every turn |
| Ceiling | open: every rung above still has a mechanical check underneath it |
Pattern 2: the housekeeper
It removes dead code, and dependencies that nothing imports any more. This is the cleanup you keep putting off.
/loop every 7d "Find dead code and unused deps. Remove one, rerun the full suite, and stop before touching the next."
Notice what isn't in there: /goal. A built-in command reaches a scheduled fire
as plain text, so the stop condition goes inside the quotes, as prose. That rule
holds for every other built-in too.
Each fire does three things, then stops:
- Find dead code, and dependencies nothing imports any more.
- Remove one. One removal per fire; the word one is doing all the work.
- Rerun the full suite, the whole thing, not just the tests near the deletion.
"Stop before touching the next" is the whole pattern. One removal, one suite run, one reviewable diff. Nine deletions in one fire is not a reviewable diff.
It is a time-based loop. It starts on Draft: it opens the pull request, and you read the diff. The worktree, the maker and checker split, and the teardown are all reused unchanged.
Pattern 3: UI/UX scoring
This one is different in a way that matters. It is a cloud routine, created
with /schedule (alias /routines), on a GitHub push event:
/schedule "Screenshot the changed page, score it against the design checklist, flag regressions for review."
The trigger is a push, not a tick: the proactive loop type from Lesson 1.1,
still built with /schedule, with an event where the clock was. Each run:
- A push arrives on GitHub.
- The changed page is screenshotted: start the dev server, capture before and after.
- Score it against the design checklist.
- Flag for review: a report a person reads.
Now the hard part. What is the check here? A score is a model's opinion. As Lesson 3.2 put it: a model will answer every time you ask, and differently each time. An answer is not a check.
So this one stops at draft, permanently. Not because the tooling is immature, but because there is nothing mechanical to graduate to. The gate is human, and it stays.
Three patterns, one ladder
| Pattern | Command | Starts on | How far it can climb |
|---|---|---|---|
| Test-stabilizer | /goal | Manual | Open, up to auto-merge |
| Housekeeper | /loop every 7d | Draft | Open, up to auto-merge |
| UI scoring | /schedule | Triage | Draft, and never past it |
Two of these gates are temporary. You hold them until you have evidence, and then you let go. The third gate is permanent. Two of these loops can go and get the evidence; a score is never evidence, so that gate isn't waiting for anything.
Open is not the same as earned. Lesson 3.3's stance stands: I have never given rung five to anything on a repo someone pays for.
All three also tear down the same way: /goal clear, delete the task by id,
set CLAUDE_CODE_DISABLE_CRON=1, then read the listing back, and don't treat
an empty listing as proof.
Where I land (a judgment, not a rule): I'd run the housekeeper on a repo I'm paid for. I wouldn't let a score decide anything on its own. You're allowed to disagree; you should just have to argue for it, out loud, to someone.
After this lesson
You will be able to:
- map the test-stabilizer, housekeeper and UI scoring patterns onto the primitives you already built, without new mechanics;
- write the stop condition as prose inside a
/loopprompt, and know that the hook, not the prose, is the ceiling; - check the
/loopexpiry against a long interval before relying on it; - tell a temporary human gate from a permanent one, by asking whether there is a mechanical check underneath.
What comes next
In Lesson 4.2: the /goal template again, blank, for a loop of your own, and
the rung it honestly starts on. Don't invent a task to practice on. If there's a
check you already run every morning by hand, that's your first loop.
Companion article
Every hook, script, and command in this course, written up and executed: How to build an agent loop in Claude Code