Learn AI Engineering
Lessons
← Loop Engineering

Three Loop Patterns: Test-Stabilizer, Housekeeper, UI Scoring

Three more loops built from the primitives you already have, and the one question that decides how far up the ladder each is allowed to climb: is there a mechanical check underneath it?

Three more loops, and not one of them needs a primitive you haven't already built. Same commands, same worktree, same maker and checker, same five rungs.

What changes is where you point them, and how far up the ladder each one is allowed to climb. For one of them, the answer is not very far at all.

Every pattern here is measured against the rule from Lesson 3.3:

The rung you're on is the one you have evidence for, not the one the tooling can reach.

Pattern 1: the test-stabilizer

You have this problem already; Module 2 made you triage it. Now fix it. Typed in a session:

/goal fix the flaky suite until 5 consecutive green runs. Stop after 15 turns.

Five consecutive green runs is the check. One red run sends the counter back to zero. It's mechanical, and a machine can tell you when it's done.

Read the second sentence again. "Stop after 15 turns" is inside the command, so it is prose: a request the model honours, not a ceiling. The ceiling is the PreToolUse hook you wrote in Module 1, reused here as-is. It fires before any permission check and can't be argued past.

Test-stabilizer
Loop typegoal-based
Targetthe flaky suite, not one failing test
Primitive/goal, typed in a session
The check5 consecutive green runs
Isolationa worktree per attempt, unchanged
Checkera sub-agent that isn't the maker, unchanged
Starts onManual: you type it and watch every turn
Ceilingopen: every rung above still has a mechanical check underneath it

Pattern 2: the housekeeper

It removes dead code, and dependencies that nothing imports any more. This is the cleanup you keep putting off.

/loop every 7d "Find dead code and unused deps. Remove one, rerun the full suite, and stop before touching the next."

Notice what isn't in there: /goal. A built-in command reaches a scheduled fire as plain text, so the stop condition goes inside the quotes, as prose. That rule holds for every other built-in too.

Each fire does three things, then stops:

  1. Find dead code, and dependencies nothing imports any more.
  2. Remove one. One removal per fire; the word one is doing all the work.
  3. Rerun the full suite, the whole thing, not just the tests near the deletion.

"Stop before touching the next" is the whole pattern. One removal, one suite run, one reviewable diff. Nine deletions in one fire is not a reviewable diff.

It is a time-based loop. It starts on Draft: it opens the pull request, and you read the diff. The worktree, the maker and checker split, and the teardown are all reused unchanged.

Pattern 3: UI/UX scoring

This one is different in a way that matters. It is a cloud routine, created with /schedule (alias /routines), on a GitHub push event:

/schedule "Screenshot the changed page, score it against the design checklist, flag regressions for review."

The trigger is a push, not a tick: the proactive loop type from Lesson 1.1, still built with /schedule, with an event where the clock was. Each run:

  1. A push arrives on GitHub.
  2. The changed page is screenshotted: start the dev server, capture before and after.
  3. Score it against the design checklist.
  4. Flag for review: a report a person reads.

Now the hard part. What is the check here? A score is a model's opinion. As Lesson 3.2 put it: a model will answer every time you ask, and differently each time. An answer is not a check.

So this one stops at draft, permanently. Not because the tooling is immature, but because there is nothing mechanical to graduate to. The gate is human, and it stays.

Three patterns, one ladder

PatternCommandStarts onHow far it can climb
Test-stabilizer/goalManualOpen, up to auto-merge
Housekeeper/loop every 7dDraftOpen, up to auto-merge
UI scoring/scheduleTriageDraft, and never past it

Two of these gates are temporary. You hold them until you have evidence, and then you let go. The third gate is permanent. Two of these loops can go and get the evidence; a score is never evidence, so that gate isn't waiting for anything.

Open is not the same as earned. Lesson 3.3's stance stands: I have never given rung five to anything on a repo someone pays for.

All three also tear down the same way: /goal clear, delete the task by id, set CLAUDE_CODE_DISABLE_CRON=1, then read the listing back, and don't treat an empty listing as proof.

Where I land (a judgment, not a rule): I'd run the housekeeper on a repo I'm paid for. I wouldn't let a score decide anything on its own. You're allowed to disagree; you should just have to argue for it, out loud, to someone.

After this lesson

You will be able to:

  • map the test-stabilizer, housekeeper and UI scoring patterns onto the primitives you already built, without new mechanics;
  • write the stop condition as prose inside a /loop prompt, and know that the hook, not the prose, is the ceiling;
  • check the /loop expiry against a long interval before relying on it;
  • tell a temporary human gate from a permanent one, by asking whether there is a mechanical check underneath.

What comes next

In Lesson 4.2: the /goal template again, blank, for a loop of your own, and the rung it honestly starts on. Don't invent a task to practice on. If there's a check you already run every morning by hand, that's your first loop.

Companion article

Every hook, script, and command in this course, written up and executed: How to build an agent loop in Claude Code