Learn AI Engineering
Lessons
← Loop Engineering

Automations: /loop, /schedule, and Triaging Flaky Tests

/goal, /loop and /schedule are one loop with three different triggers, and a three-rerun triage keeps a flaky test out of the fix loop before it disables the no-progress rule.

Last lesson was about where a loop's memory lives. This one asks the question that comes before it: what starts a run?

Three commands, three questions

/goal, /loop and /schedule can all start a run. Ask the same three questions of each: what triggers the run, what stops it, and where the work lives.

/goal/loop/schedule
Loop typegoal-basedtime-basedtime-based
TriggerYou, once.An interval: 1d, or every 2 hours.A cloud schedule, minimum one hour.
Stops whenYour success condition holds, or a cap trips.You close the session, or a cap trips.You delete the routine. It outlives the terminal.
Where the work livesThis session, this one task.This session, your machine.A fresh clone in the cloud that cannot see your repo.

The trigger is the taxonomy from Lesson 1.1. /loop and /schedule are both time-based; /goal is goal-based.

The difference between the two time-based commands is where they run. Every /loop fire is a new session on your machine, which is exactly what Lesson 1.3's .loop/state.json is for. A /schedule routine runs in the cloud on a fresh clone, and that clone never sees the file at all.

A few details from the screen worth keeping:

  • /schedule is also available as /routines; the web view is claude.ai/code/routines.
  • Cloud routines need a claude.ai subscription login. An API key, Bedrock, Vertex or Foundry each fail with a named error.
  • /loop documents two interval forms: a leading token (1d) or a trailing clause (every 2 hours), in units s, m, h, d. An interval that doesn't map to a clean cron step is rounded, and Claude tells you what it picked.
  • At an interval of an hour or more, /loop asks whether the run belongs in this session or in the cloud.

Triage: rerun three times before you fix anything

Two tests are failing. Before the loop touches either one, rerun each on its own, three times. On its own means outside the suite, one test at a time.

CandidateRerun 1Rerun 2Rerun 3VerdictLane
tests/test_retry.py call_with_retries()failpassfailflakyQuarantine and report it. A human reads this one. It never enters the fix loop.
tests/test_pricing.py apply_discount()failfailfailrealSame failure, three for three. This is the only one the loop is allowed to fix.

Inconsistent means flaky. Consistent means real. Only the real failure earns a fix.

This is not hygiene. Recall the no-progress rule from Lesson 1.2: stop when the same failing assertion and the same diff hash repeat across two iterations. A flaky test fails differently every time, so the counter never sees a repeat and the rule never fires. The quarantine list that triage produces is the input the no-progress check reads.

The composition: one command, five decisions

Here is the whole thing as one command:

/loop 1d "Check the repo for newly failing tests. Rerun each 3x to rule out flakiness. If a real failure remains, fix it and keep going until the suite passes, then push the branch (not master)."

Read it in pieces.

  1. 1d: the trigger. Once a day. That is the only thing /loop decides.
  2. "Check the repo for newly failing tests": fresh input. Not yesterday's list, but whatever is red right now.
  3. "Rerun each 3x to rule out flakiness": triage. The habit from the table above, written into the instruction so it happens on every run.
  4. "keep going until the suite passes": the stop condition, in prose. /loop decides when each run starts; this sentence decides when it's done.
  5. "push the branch (not master)": one bounded action, on a branch. Never master. The loop proposes; you still merge.

Notice what isn't in those quotes: the iteration cap. Module 1's template says every goal carries an iteration cap, a spend cap and a no-progress rule, and they belong in this prompt too. Written there, they are prose that the model honours. The PreToolUse hook from Lesson 1.2 is what makes the iteration cap real. No hook covers spend, because no hook is handed a running-cost number.

The trigger slot: a clock is only one option

A clock tick is one kind of trigger. An event somewhere else is another: a pull request opening, a webhook firing. That is the fourth loop type from Lesson 1.1, proactive. It isn't a fourth command. It's /loop or /schedule with an event in the trigger slot instead of a clock tick.

Keep two things apart that are easy to conflate:

  • A trigger is what says "go" and starts the loop.
  • A connector is what the loop acts through once it is already running. Lesson 2.4 covers that half.

Where each trigger is created:

TriggerCreated from
Schedulethe CLI
GitHubthe CLI (since v2.1.225), or the web
APIthe web only; the CLI can't create or revoke tokens

/web-setup does none of these. It grants repo access for cloning.

Three things that change when nobody is watching

1. The bill you can't see. Closing the terminal is supposed to stop a /loop. One reported /loop worker survived Ctrl-C and a closed terminal, and ran roughly 864 iterations over 72 hours, for about $300 nobody approved. All the while, CronList and TaskList both came back empty. The failure wasn't the spend; it was that nothing could see the loop. An empty task list is not evidence that a loop has stopped. Turning one off is its own procedure, in Lesson 2.5.

2. The reading. Somebody still has to check what the loop did. Each primitive removes one manual step; none of them removes the check. Automation moved the work, not the verification. That is Module 3, and it is not optional.

3. The collision. Two fixes running at once, in one working copy. They edit each other's files, and neither fix can be trusted.

After this lesson

You will be able to:

  • tell /goal, /loop and /schedule apart by trigger, stop condition, and where the work lives;
  • triage a failing test by rerunning it on its own three times, and route a flaky one to quarantine instead of the fix loop;
  • explain why an unstable failing set silently disables the no-progress rule;
  • compose a /loop with a prose stop condition, and say why /goal can't go inside its prompt;
  • separate a trigger from a connector, and place a proactive loop as /loop or /schedule with an event trigger.

What comes next

Two fixes, one working copy. In Lesson 2.2, "Worktrees: Isolating Parallel Runs, and What They Don't Isolate", each run gets its own checkout, and you see exactly which half of the problem that solves.

Companion article

Every hook, script, and command in this course, written up and executed: How to build an agent loop in Claude Code