Why an AI Agent Session Is More Than a JSONL File

How Infai evolved its AI agent session store from a simple JSONL log into a durable, branchable, lazy-loaded, and compactable timeline.

infai ai system-design ai-agents sustainable agent-harness opencode pi claudecode
13 min read 2,523 words
Why an AI Agent Session Is More Than a JSONL File

You can use an AI coding agent for a few minutes and never think about its session storage. Use one for a full day, branch the work, recover from a crash, and return to it tomorrow, and storage becomes part of the product.

That was the problem I kept running into while using agent tools. Over the past year, I have used OpenCode, Pi, Claude Code, Google ADK, Jules, and others. Many of them are useful, but the experience often becomes awkward when you move beyond the happy path: custom approval screens flicker on small terminals, subagent permissions are difficult to reason about, and multi-agent state can leak into the main session.

The tools often expose an SDK, but the SDK is sometimes another layer of TUI customization. What I wanted was lower-level control over the harness and agent loop: permissions, branching, recovery, persistence, and the rules that decide what context reaches the model.

That is why I started building Infai. Part of the motivation was to learn, but the larger goal was to build an agent harness that treats failure as a normal operating condition instead of designing only for the successful path.

TL;DR

  • An agent session is not just chat history. It is a durable, navigable event timeline.
  • Branching, HEAD, chunked logs, lazy payloads, and compaction reduce the work needed to rebuild active context.
  • The best optimization is often deciding what not to load, not adding another feature to the interface.

The Session Is the Agent’s Working Memory

An agent has several kinds of state, and they do not all have the same lifetime. The model needs a working context for the next request. The harness needs runtime state such as permissions and in-flight tool calls. The user needs a durable record that can be inspected, resumed, or branched.

I call the model-facing part of that working context the ActiveSession. It is short-term memory, but it is not disposable. If it cannot be reconstructed after a restart, a process crash, or a branch switch, the harness cannot be trusted with long-running work.

The design question therefore became more specific:

How do we store enough history to reconstruct the active timeline without loading every byte of every previous tool result on every turn?

The answer emerged through eight iterations.

Iteration 1: Start With Durable Events

The first step was to understand what an LLM provider actually returns. A response is not one plain string. Depending on the provider and the operation, it can contain roles, content, reasoning, tool calls, tool results, usage, and provider-specific metadata.

The first useful abstraction was an event. Each interaction became an append-only record, with session metadata stored before the conversation began. Infai uses UUIDv7 event IDs, which carry creation-time ordering while remaining globally unique.

{"kind":"meta","role":"","id":"xxxx-aaa-bbb-cccc","session_id":"xxxx-aaa-bbb-cccc","record":{"last_used_model":"model-name","session_name":"example","working_dir":"/workspace","created_at":"2026-08-25T17:21:22Z","updated_at":"2026-08-25T17:21:22Z"}}
{"kind":"message","role":"system","record":{"content":"..."}}
{"kind":"message","role":"user","record":{"content":"Build the feature."}}
{"kind":"message","role":"assistant","record":{"content":"I will inspect the repository first."}}

JSONL was a reasonable first choice for local storage. It is easy to append, easy to inspect, and naturally supports replay. The initial layout was intentionally boring:

~/.config/infai/harness/sessions/<session-id>/events.jsonl

The important decision was not JSONL itself. It was treating the session as a sequence of facts rather than repeatedly rewriting one large mutable object.

There is one important distinction in the implementation: streamed deltas are live-only. They fan out through the session event hub for the TUI or SSE client, but the settled message is what becomes durable history. Persisting every token would make the event log noisy without improving session recovery.

Primitive agent session events stored as JSONL
The first session primitive: append events to a durable log.

Iteration 2: Add Branching Without Duplicating History

The moment a user can revisit an earlier point and try a different approach, a session stops being a list. It becomes a tree.

The simplest representation is an event with a pointer to its parent. A new response normally points to the current event. A response created from an older event points to that event instead, creating a branch.

{"id":"01a02f3f-a89d-7d25-b4ac-ac89a183258d","parent_id":"00000000-0000-0000-0000-000000000000","branch_from":null,"kind":"message","role":"user","record":{"content":"Try approach A."}}
{"id":"01a02f3f-cb65-73b7-a632-1d948274475f","parent_id":"01a02f3f-a89d-7d25-b4ac-ac89a183258d","branch_from":null,"kind":"message","role":"assistant","record":{"content":"Approach A has a problem."}}
{"id":"01a02f54-b76b-7808-a895-3069dcca66f6","parent_id":"01a02f3f-a89d-7d25-b4ac-ac89a183258d","branch_from":"01a02f3f-a89d-7d25-b4ac-ac89a183258d","kind":"message","role":"user","record":{"content":"Try approach B instead."}}

parent_id tells us how to walk backward. branch_from tells us that this event begins a deliberate alternate path instead of being an ordinary child. That distinction matters when rendering a timeline and when summarizing work from an abandoned branch.

For the local representation, I chose an append-only array of events and a hash map from event ID to array index. The map avoids scanning the whole log to locate a parent.

  • Event lookup is average-case O(1) and worst-case O(n) if a hash table degenerates.
  • Appending is amortized O(1).
  • Reconstructing a path is O(k), where k is the number of events on that path.

This does not require a B-tree. A B-tree could be useful for a different storage backend or query pattern, but the session’s primary operation is append plus parent traversal. Keeping the representation small was itself an optimization.

Branch selection is intentionally two-phase. Selecting an older event changes a pending in-memory parent, not the persisted HEAD. The old active branch remains current until the user submits a new prompt. Only then does the harness append the new event and move HEAD to the new branch.

Agent session events forming a branchable tree
Parent links let one session contain multiple timelines.

The index stores enough information to locate an event without reading every event first:

{"id":"01a039f0-e016-7dd2-aac4-669467e8d7e4","parent_id":"00000000-0000-0000-0000-000000000000","timestamp":"2026-08-25T17:21:22.454897183Z"}
Event ID to parent and index mapping
The index gives the harness a fast route to event locations.

Iteration 3: Add HEAD

Branching gives us many possible paths, but the harness still needs to know which path is active. Git solves this with HEAD, and the same idea works here.

HEAD stores the ID of the latest event on the active timeline. To rebuild the current context, the harness starts at HEAD and follows parent_id until it reaches the root or a compaction boundary.

The pointer must be updated carefully. A crash between writing an event and updating HEAD should leave us with a valid older session, not a pointer to an event that was never fully persisted. On startup, the harness validates the pointer and can finish an interrupted temporary HEAD update when the referenced event is already indexed.

The current write sequence is deliberately conservative: sync the chunk, write and sync a temporary HEAD, append and sync the index entry, then atomically rename the temporary HEAD and sync the session directory. If the process dies after the chunk is synced but before the index is, startup reconciliation rebuilds missing index entries from the chunks. A torn final JSONL line is truncated; a complete but invalid line is treated as corruption instead of being silently ignored.

HEAD pointing to the active agent session timeline
HEAD identifies the timeline the agent should continue.

Iteration 4: Move Session Metadata Out of the Event Log

The first design put metadata and conversation events in the same file. It worked, but it forced simple operations such as listing sessions or changing the selected model to understand the event log.

The next layout separated stable session metadata from the event stream:

sessions/<session-id>/
  events.jsonl
  index.jsonl
  session.json
  HEAD

session.json can hold the session name, working directory, selected model, and timestamps. The event log can remain focused on what happened.

I also removed the system prompt from the durable conversation events. The system prompt belongs to the agent configuration and can change with the active model or agent. Keeping it in the loop as a dynamic input gives the harness more flexibility without pretending that every generated prompt is a user-visible event.

This separation is useful beyond storage. It distinguishes session facts from context assembly. The session records what happened; the agent loop decides which instructions and events should be sent to the model now.

Session metadata separated from the event stream
Metadata and conversation events have different responsibilities.

Iteration 5: Chunk a Growing Event Log

One events.jsonl file is convenient until a session becomes large. Listing a session should not require opening a multi-gigabyte tool history, and loading the current branch should not require scanning unrelated branches.

The next step was chunk rotation:

sessions/<session-id>/
  chunks/
    000000.jsonl
    000001.jsonl
  blobs/
  index.jsonl
  session.json
  HEAD

The chunk and blob thresholds are configuration choices, not properties of the data model. The threshold can be tuned for different storage characteristics: chunks bound event-log reads, while sufficiently large serialized records move to the blobs/ directory.

The index now records the physical location of each event:

{"id":"01a039f0-e016-7dd2-aac4-669467e8d7e4","parent_id":"00000000-0000-0000-0000-000000000000","timestamp":"2026-08-25T17:21:22.454897183Z","location":{"chunk":0,"offset":0,"length":226}}

The harness can read only the chunks needed for a timeline or a bounded range. The exact rotation threshold is a policy choice. Smaller chunks make targeted reads cheaper but increase file and index overhead; larger chunks reduce metadata overhead but increase the amount of data read during recovery.

This is the same kind of trade-off found in log rotation and write-ahead storage: bounded work matters more than a perfect universal chunk size.

Chunked event files with an indexed event location
Chunking bounds the amount of history a read must touch.

Iteration 6: Load Large Tool Results Lazily

Tool calls and tool results can be much larger than ordinary messages. A compiler output, a test log, or a file read can consume more memory than the conversation around it.

The event store still needs to preserve those payloads, but every operation does not need them. A session list needs a name and timestamp. A timeline view needs event types and short summaries. Only active-session reconstruction may need the full content. In the current implementation, ToolCallRecord and ToolResultRecord define the durable shape, but the tool loop has not yet been wired to emit those records. The storage layer is ready; the integration is still future work.

That led to a lazy-loading rule: keep metadata and a digest in the index or lightweight event record, then resolve the large payload only when the caller asks for it.

{"id":"01a02f3f-cb65-73b7-a632-1d948274475f","branch_from":null,"parent_id":"01a02f3f-a89d-7d25-b4ac-ac89a183258d","kind":"message","role":"assistant","digest":"sha256:2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824","record":null}

The digest gives us an integrity check and a stable reference for a payload that is not loaded yet. This saves memory during browsing and avoids unnecessary CPU and disk work during session listing.

Blob files are content-addressed by the SHA-256 digest of the serialized record. The same payload therefore does not need to be written twice, and loading a blob verifies its digest before decoding it. LoadEvent keeps the blob unresolved; the engine resolves it only when rebuilding chat history.

Large tool result payloads resolved lazily from session storage
Heavy content is available when needed, not loaded by default.

Iteration 7: Make Compaction a First-Class Event

Even a large model context window is finite. More importantly, sending the entire history on every request increases latency and cost long before the model reaches its hard limit.

Compaction is the point where storage design and model behavior meet. Instead of deleting old events, the harness records a compaction event that summarizes a bounded portion of the timeline. The original events remain available for audit, debugging, or a future policy that needs them. Infai currently triggers automatic compaction when reported token usage reaches 80% of the configured model context window, and also exposes manual compaction.

{"id":"01a02f3f-cb65-73b7-a632-1d948274475f","branch_from":null,"parent_id":"01a02f3f-a89d-7d25-b4ac-ac89a183258d","kind":"compaction","role":"assistant","digest":"sha256:2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824","record":{"summary":"The agent inspected the repository, chose the chunked event design, and left the migration test unfinished."}}

The active timeline can now stop at the root or at the relevant compaction boundary:

HEAD -> parent -> parent -> compaction
                         or root

For a path with N events, reconstruction is O(N) when walking the parent links. Each indexed event lookup is average-case O(1), so the combined path walk is approximately O(N) with O(1) lookup per step. The best case is O(1) when HEAD is itself a compaction boundary or the requested context contains one event. The worst case remains O(N) when the active path has no useful boundary and every event must be visited.

Compaction is lossy, so the summary must preserve more than a vague paragraph. At minimum, it should retain goals, constraints, decisions, files changed, unresolved work, and facts that would be expensive or dangerous to rediscover. Important identifiers should live in durable artifacts as well, because no summary is guaranteed to preserve every detail.

This is also why compaction should be a harness-level operation. It needs a trigger policy, a boundary policy, failure handling, and a way to retry without corrupting the session. It is not only a button in the TUI.

Compaction event creating a shorter active session path
Compaction shortens the active path without erasing the original history.

Iteration 8: The Simplified Design

After these iterations, the design became less about adding features and more about assigning each responsibility to the right layer:

sessions/<session-id>/
  chunks/
    000000.jsonl
    000001.jsonl
  blobs/
  index.jsonl
  session.json
  HEAD

The event stream is append-only. Parent links provide the tree. HEAD selects the active branch. The index makes event locations discoverable. Chunks bound disk reads. Digests and lazy payloads keep lightweight operations lightweight. Compaction gives long-running sessions a deliberate context boundary.

The final system is not optimized because it uses a clever data structure. It is optimized because each operation reads only the state it needs.

There is an operational boundary worth calling out. A user message is persisted before model generation begins, so a hard crash does not lose what the user typed. The assistant’s settled response is persisted when the turn completes; an interrupted response is the part that can be lost. That is a reasonable first durability contract for an interactive harness, but it should be explicit.

Final Infai agent session storage design
The final design combines durable events, branches, chunks, lazy payloads, and compaction.

What This Changed About My View of Agent Harnesses

I started with a frustration about agent tools: the visible interface often looked more customizable than the underlying loop really was. I wanted better permissions, better subagent boundaries, and fewer surprises when a session became complicated.

The session store turned out to be a central part of that problem. If the harness cannot identify the active branch, recover after a crash, load a bounded context, or explain where a tool result came from, then higher-level controls will always feel unreliable.

The broader lesson is that agent context needs the same engineering discipline as any other stateful system. Durability, indexing, concurrency, lazy loading, and recovery are not distractions from the intelligence of the agent. They determine whether that intelligence remains useful after the first successful demo.

We began with a JSONL file because it was easy. We ended with a timeline because the agent needs to remember not only what happened, but also which version of what happened it is continuing.

That is the difference between storing a transcript and building a session.

Further Reading

Dipankar Das

Dipankar Das

Building sustainable and efficient platforms without firefighting