Learn AI Engineering
Lessons
← Loop Engineering

Loop Failure Modes, and the Tasks a Loop Should Never Own

When a loop misbehaves it takes one of four shapes, each with a fix, and before you build one ask whether a machine can tell you it's done.

That's where Lesson 3.1 left it. The checker is in place, the loop is running, and something is still wrong. So what is it?

Two questions, and the order matters:

  1. Is the loop broken? That takes four shapes.
  2. Or was this never a job for a loop at all?

Four failure modes

Each row is what you see, what it actually is, and what fixes it.

What you seeWhat it isThe fix, and where you built it
01The same failing test, run after run.Stuck, no progress: the same failing assertion and the same diff hash, twice.Stop and ask; the hook-enforced cap ends the run (1.2). It only fires once triage has quarantined the flaky tests (2.1).
02It reports success. The bug is still there."Done" but broken: the maker graded its own work.The verification checklist, run by a checker that is not the maker (3.1).
03The check has never once disagreed.The rubber-stamp verifier: it reads the transcript, not the repo.Demand the output, not the verdict: "here's the green run", never "tests pass" (3.1).
04Hour three is worse than hour one, and nothing failed.Quality decay, from context rot: the window is mostly its own old output.No exit code, and no hook. This is the one you have to go looking for.

Look down the fix column. Three of those four you already built. That was the plan.

Quality decay: the one you haven't built

Every iteration feeds its own output back in. Picture one context window at three points in the same run. At iteration 1 it holds the goal, what it read from the repo, and plenty of room left. By iteration 6, its own earlier output has moved in. By iteration 12, most of what the model reads is itself.

There's no exit code for that. Nothing exits non-zero, nothing is flagged; the work just gets vaguer.

The tell is the diff. It starts editing files it has no reason to edit, and the commit messages stop saying anything.

So bound the run, and start the next one clean. Short runs beat long ones. Lesson 1.3 put the progress on disk in .loop/state.json, which is what makes a restart cheap: the progress lives in the file, not in the window.

The tasks a loop should never own

Now the other half. Sometimes nothing in that table applies: the loop was the wrong shape from the start. One question sorts it.

Can a machine tell you it's done?

If a person has to look at the result to know, you don't have a stop condition. You have an opinion, and a loop can't stop on one.

Kind of taskExampleWhy it fails as a loop
01A vague goal"Make the app faster."Faster than what, measured how? A goal that can't fail can't finish. There is no condition to test, so it runs until something else stops it.
02A subjective call"Does this copy read well?"A model will answer all day, and differently each time. An answer is not a check.
03Human taste"Is this the right abstraction here?"A judgment about a codebase you have to live in afterwards. It has no exit code, and it is yours. That is not a gap in the tooling.

None of this is off limits to an agent. It just never graduates past a draft, and never to auto-merge. Point the loop at the drafting, and keep a permanent human gate on the judgment. Module 4's UI-scoring loop is exactly this shape.

The card to keep

When it breaks, name it:

  • Stuck: the cap ends the run, and triage is what makes the rule fire at all.
  • "Done" but broken: the checklist, run by a checker that isn't the maker.
  • Rubber stamp: evidence, not an account of the evidence.
  • Quality decay: bound the run and start clean. Nothing will tell you; go and look.

Before you build it, ask: can a machine tell you it's done? If yes, it can be a loop, and everything above applies. If no, it's a draft, and a gate that stays.

Ask that question before you write the goal, not after the third failed run. A loop that shouldn't exist can't be debugged into one that should.

After this lesson

You will be able to:

  • diagnose a misbehaving loop against four failure modes (stuck, "done" but broken, rubber-stamp verifier, quality decay) and apply the matching fix;
  • spot quality decay from the diff, when no check will flag it, and bound the run instead of letting it drift;
  • decide, before writing a goal, whether a task can be a loop at all, and keep a human gate on the ones that can't.

What comes next

In Lesson 3.3, "Staying the Engineer: The Risk Audit and the Adoption Ladder": three risks your loop creates, and which rung of the adoption ladder it is honestly ready for.

Companion article

Every hook, script, and command in this course, written up and executed: How to build an agent loop in Claude Code