08. The Session Log and Recovery
The source of truth for an agent runtime should be the session log, not the in-memory messages array. Memory gets lost, UIs get refreshed, processes crash, and users come back tomorrow to keep going. As long as the log is complete enough, you can rebuild the context, display history, audit tool calls, resume unfinished tasks, and even fork from some intermediate node.
Why saving only the prompt is not enough
Many minimal implementations only save the current messages on exit. This causes several problems:
- Tool execution progress and errors are lost.
- Non-message events like model switches, compaction, and user steering are lost.
- Branching cannot be expressed; you can only overwrite linearly.
- On a crash, the last chunk of state may never have been written at all.
An agent session is not a chat transcript; it is an event log. The chat UI is just one projection of that log.
A JSONL append-only log
A teaching project can use JSONL: one entry per line, appended as you go. On a crash, at most the last line is corrupted, and everything before it can still be parsed. Every entry needs at least an id, parentId, timestamp, and type.
type SessionEntry =
| { type: "session"; id: string; timestamp: string; cwd: string; version: number; parentSession?: string }
| { type: "message"; id: string; parentId: string; timestamp: string; message: Message }
| { type: "model_change"; id: string; parentId: string; timestamp: string; model: string }
| { type: "compaction"; id: string; parentId: string; timestamp: string; summary: string; firstKeptEntryId: string }
| { type: "custom"; id: string; parentId: string; timestamp: string; customType: string; data: unknown }
| { type: "label"; id: string; parentId: string; timestamp: string; targetId: string; label: string };
parentId turns the session into a tree rather than just an array. When the user resumes from some point in history, you can create a new branch; the original follow-up records are still preserved. The currently active branch can be identified by its leaf id.
Using a time-ordered id such as UUIDv7 for each entry id is a practical choice: sorting by id sorts by time, and you can take a prefix for a short display in the UI. A mature system has more entry types than the ones above — switching reasoning effort, branch summaries, custom messages injected by extensions, labeling a node, renaming the session; each is its own entry type. The shared principle is this: any state change that affects "how the session runs afterward" goes into the log; transient state that only affects the current process (such as scroll position) does not.
Storage layout and discovery
A single file is not enough; you also have to think about how "all the sessions on this machine" are organized. One battle-tested layout is: group by project directory, encode the workspace path into the directory name, one JSONL file per session, and include the session id in the file name. This supports three high-frequency operations:
- "Continue this project's most recent session": scan the files under this project's directory and take the newest by modification time.
- "List this project's past sessions": read only the first-line header of each file (limited to the first few hundred bytes), without parsing the whole thing.
- "Recover a session by id prefix": look in the current project first, then globally.
The header must contain cwd and a format version. The cwd is used to verify "which project this session belongs to" (note that symlinks must be normalized, or the same project will show up as two separate histories); the version is used for migration.
The log is the source of truth; the context is a projection
When building the model context, you do not read the whole log and send it to the model. Instead, you walk from the current leaf back to the root to get the active branch, then project the entries into LLM messages. The projection rules can be:
messageentries go into the context.model_changedoes not go into the context, but it determines which model subsequent requests use.customstays out of the context by default, unless an extension declares it to be a message type.compactionenters the context as a single summary message and determines which older entries are kept.label, session renames, and other pure metadata never enter the context.
This way the log records the complete facts, while the context contains only what the model needs to keep working.
The recovery flow
When recovering a session, the runtime should:
- Parse the JSONL, skipping or reporting corrupted lines.
- Validate the session version and migrate old entries if necessary.
- Index by id, find the current leaf, and walk back along parentId to build the active branch.
- Rebuild the messages, model, compaction state, and queue state from the entries.
- Re-create the tools and the provider.
- Let the UI project the history from the log, instead of asking the model to recite it.
Recovery should not automatically re-execute historical tools. Tool results are already facts. Unless the user explicitly asks to rerun the tests, recovery only rebuilds state.
The migration in step 2 is not a hypothetical requirement. A real system's log format will inevitably evolve: upgrading from a linear array to a tree (backfilling id and parentId on all old entries), renaming message roles, changing index references into id references. Every change needs an idempotent migration, executed in version order, with old-format sample files kept in the tests. A log format with no migration mechanism turns all of a user's past sessions into waste paper the first time you make a breaking change.
Branching and fork
Branching is not an advanced feature; it is a natural need for an agent. A user might have the agent try approach A, decide they are not happy with it, and switch to approach B from a midpoint. If the log is a tree, all you have to do is point the leaf at some historical entry and append new messages. The old branch remains viewable.
Cross-session fork can be built on the same mechanism: copy a session's active branch into a new file and record parentSession in the header pointing at the source. That way "starting a new task from this discussion" has a complete lineage you can trace, while the two sessions stay out of each other's way.
Note the relationship between branching and compaction. Compaction should not destroy the old log; it merely tells the context builder "from this point on, represent everything earlier with a summary." If compaction deleted history outright, you could never return to a pre-compaction branch, nor audit early tool calls.
What you should observe
A stretch of session log might look like this:
{"type":"session","id":"s1","timestamp":"2026-01-01T10:00:00.000Z","cwd":"/repo","version":1}
{"type":"message","id":"m1","parentId":"s1","timestamp":"2026-01-01T10:00:01.000Z","message":{"role":"user","content":[{"type":"text","text":"Fix the failing test"}]}}
{"type":"message","id":"m2","parentId":"m1","timestamp":"2026-01-01T10:00:03.000Z","message":{"role":"assistant","stopReason":"toolUse","content":[{"type":"toolCall","id":"t1","name":"bash","input":{"command":"npm test"}}],"model":"example","usage":{"inputTokens":500,"outputTokens":40}}}
{"type":"message","id":"m3","parentId":"m2","timestamp":"2026-01-01T10:00:06.000Z","message":{"role":"toolResult","toolCallId":"t1","toolName":"bash","isError":true,"content":[{"type":"text","text":"1 test failed"}]}}
These few lines are already enough to recover the user request, the model's action, the tool result, and the context for the next step.
Exercise
Implement JSONL session storage.
Acceptance criteria:
- Each message is appended, never overwriting the existing file; a corrupted last line does not affect the entries before it.
- After a process restart, the current active branch can be recovered.
model_changeaffects the default model after recovery, but does not enter the LLM messages.- After forking from a historical entry, the old branch remains readable.
- A compaction entry does not delete old entries; it only changes the context projection.
- Write a v1-to-v2 format migration (for example, backfilling parentId on a linear log), and an old sample file recovers correctly after migration.