Learn AI Engineering
Lessons
← Loop Engineering

State Management: Surviving a Session

A loop's memory is a file, not a conversation: which fields reset when a new session starts, which must survive it, and three values that never belong there.

Last lesson, the iteration cap reset itself: the hook payload carries a session_id, so a new session gets a fresh count. That is right for the counter. It says nothing about the values that should not reset, and a loop that runs overnight has some.

Two kinds of memory

Your loop has two kinds of memory, and only one of them survives the night.

The transcript holds everything the run tried and everything it concluded. One /clear, one compaction, or one closed terminal, and it's gone. Resuming a session restores the conversation, but it has never restored a counter.

The state file holds a few fields, written to disk before the turn ends. Tomorrow morning's run is a new session, and this file sits on both sides of that boundary. The new session reads what the last one left.

Survives a new session?What it holds
TranscriptNoEverything the run tried and concluded
State fileYesA few fields, written to disk each turn

One file, and where it splits

The state file is the counter from Lesson 1.2, grown up. It is one file with five fields, split in half:

{
  "run": {
    "session_id":        "a3f9c1e2",
    "bash_calls":        31
  },
  "work": {
    "failing_signature": "test_apply_discount_larger_price",
    "diff_hash":         "9f2c1ab",
    "unchanged":         1
  }
}

The top half is the run. It records which session this is and how many shell calls the hook has counted. It is run-scoped: it resets on a new session, so a new run gets a fresh cap of 40 calls.

The bottom half is the work. It holds the two values from Lesson 1.2's definition of no progress (the failing assertion and the diff hash) and how many iterations they have stayed identical. It is work-scoped: it survives the session.

The work half exists because one iteration cannot compare itself with the one before. That comparison only exists if something wrote it down. Your test step writes these values, never the model. A number the agent can edit is not a measurement.

The two halves are nested rather than sitting side by side, so "reset the run half" is one assignment, not a field-by-field decision. That is also why the guard.py from Lesson 1.2 already works with this layout: it seeds the run half and updates it in place, so a work half next to it is left alone when the session changes. Reading a field back is json.loads(...) in the same Python, with no jq needed.

Two ways to get the split backwards

If the run-scoped half never resets, the counter just climbs:

"run.bash_calls": 312

The cap has outlived the run it was capping. By Thursday the loop can't run at all, because every shell call is blocked.

If the work-scoped half resets every session, you get this every morning:

"work.unchanged": 0

The loop retries the same broken fix forever. No progress never trips, because every attempt is attempt one.

So ask one question of every field: what should this be when a new session starts?

The run half resets itself. Nobody runs a reset step: the hook reads the session_id and does it, so it is never the model's job and never a startup step you have to remember. The work half is the one nothing resets for you, so deciding what goes in it is your call.

The guard script and the goal are committed to the repo. The state file is a run artifact and stays out of git. Ignore the file, .loop/state.json, not the whole .loop/ folder.

Three values that never go in it

Everything in this file gets read back, by your hook, by the agent, and by whatever the agent writes next. That rules out three kinds of value.

Never storeExampleWhy
A secret"api_key": "sk-ant-…"Everything here is read back by the hook, the agent, and whatever it writes next.
A conclusion"verified": trueStore what you measured, not what you decided. The next run can't re-check a verdict; it just believes it.
A copy of anything you can recompute"diff": "--- a/src/loopdemo/pricing.py …"Keep the hash instead. A stale copy is one the loop still trusts.

After this lesson

You will be able to:

  • explain why a loop's memory has to live in a file, not the transcript, and what /clear, compaction, and resuming a session do and don't keep;
  • split a state file into a run-scoped half that resets on a new session and a work-scoped half that survives it;
  • spot both ways to get that split backwards: a cap that never resets, and a no-progress counter that always does;
  • keep secrets, conclusions, and recomputable copies out of the state file, and keep the file itself out of git.

What comes next

In Lesson 2.1, "Automations: /loop, /schedule, and Triaging Flaky Tests", the loop starts running without you. /loop re-reads this file on every run. /schedule gets a fresh clone that can't see it at all, which is where the question of where the work lives starts to matter.

Companion article

Every hook, script, and command in this course, written up and executed: How to build an agent loop in Claude Code