Skip to content

feat: stable turn ids, per-turn workspace snapshots, and conversation history that survives restart (#354) - #361

Merged
saucam merged 5 commits into
mainfrom
feat/turn-checkpoints-354
Oct 10, 2026
Merged

saucam merged 5 commits into
mainfrom
feat/turn-checkpoints-354

Conversation

@saucam

@saucam saucam commented Oct 10, 2026

Copy link
Copy Markdown
Collaborator

Closes #354

What this does

Every turn in a session now has a stable identity, codeoid snapshots the files each turn starts from, and the conversation history survives a daemon restart.
This is the foundation for going back a turn (#355), forking from an earlier point (#356), comparing backends (#357) and per-turn diffs (#358).
It works the same on every backend, today's and future ones: nothing depends on a backend's own session format.

  • Turns. Each prompt, and everything the agent does in reply, carries one turn id. A message sent while the agent is still working joins the running turn instead of starting a new one. Clients can list a session's turns with session.turns (order, a short preview of the prompt, when it started, and its snapshot).
  • Snapshots. At the start of each turn, codeoid records the working directory in its own storage, one shadow git repository per session. It never touches your repository: no refs, no objects, nothing in git log --all, no hooks or filters run. It works in folders that aren't git repositories, and leaves out .gitignored files, common secret files (.env, keys, credential files) and dependency folders. If a snapshot is still being taken when the agent starts, it's marked "late".
  • History survives restart. Codeoid's backend-neutral copy of the conversation is now saved to disk. Before this change, after a daemon restart:
    • forking a session, or switching its backend, silently carried no conversation;
    • a stateless backend (OpenAI, Gemini API) answered the next prompt with no memory of the session.
  • Stop works while a message is still being prepared. Pressing Stop (or shutting down) while a send was preparing its turn used to stop nothing: the turn started anyway afterwards. Now it's cancelled, the user is told, and anything waiting on it (a pipeline phase, a background wake) is released.

How it was checked

Live, against a real daemon with the real Claude backend (a scratch git repo with an untracked .env):

Check Result
Two turns that create and then edit hello.txt Each message carries its turn's id. session.turns lists both with snapshots. Turn 1's snapshot holds only README.md (no .env), turn 2's holds hello.txt = one. The user repo has no new refs or log entries.
Restart the daemon, fork the session, ask "what is my codename?" (told in turn 1, no tools) This branch: "Your codename is PELICAN." The fork lists the inherited snapshots.
Same on main "I don't know your codename". The fork had no conversation.

Tests: typecheck, lint and build pass, plus 2,848 daemon tests and 577 web tests. New: checkpoints.test.ts (21), turns.test.ts (20) and canonical-log.test.ts (14).

  • The checkpoint suite covers: the user repo left exactly as it was; a malicious core.fsmonitor, filter driver or hook never running; inherited GIT_* variables ignored; snapshots surviving deletion of the user's .git; planted skip-worktree / assume-unchanged / split-index entries having no effect; a .gitignore that negates .env not pulling it in; nested repositories; stale locks; storage budget; forks.
  • Each guard added during the audits was mutation-checked: removing it makes a test fail.

Audits: two rounds, each with a correctness reviewer and a security reviewer. All findings were fixed in this PR.

  • Round 1: running git in the user's repo let the repo's own config run commands on every turn, outside the approval gate. That is the reason for the shadow repository.
  • Round 2: made the shadow repository fully self-contained, fixed a race with turns the backend starts on its own, and stopped a compaction edge case from wiping the history log.

Behaviour changes and limits

  • The first snapshot of a session stores one compressed copy of the working directory. Later snapshots store only what changed.
  • Snapshots are skipped, with a log line, when:
    • the first snapshot exceeds 1 GB or 100k files;
    • a later snapshot adds more than 100 MB or 20k new files.
  • Each session keeps its newest 200 snapshots, capped at 2 GB. Settings live under session.checkpoints (enabled, maxPerSession, maxUntrackedBytes, waitMs).
  • A turn waits at most 2 s for its snapshot.
  • Nested git repositories are recorded as a pointer to their commit, not their files.
  • .env.example is excluded along with .env, because the exclusions are deliberately not overridable.
  • The history log caps very large fields (attachments, tool output) and keeps its newest 16–32 MB.
  • Sessions created before this change get their history rebuilt from their transcript, once.
  • Context rotation still starts a fresh context, but the turn list is no longer reset.
🤖 Implementation context (for agents / maintainers)

Files

  • src/daemon/checkpoints.ts: shadow repo per session at <transcriptDir>/checkpoints/<id>.git, used with GIT_WORK_TREE=<workdir>.
    • Environment: allowlist (PATH, HOME, LANG, LC_ALL, TMPDIR) plus GIT_CONFIG_NOSYSTEM=1, GIT_CONFIG_GLOBAL=/dev/null, -c core.hooksPath=/dev/null -c core.fsmonitor=false.
    • Storage: self-contained, with no alternates and no copied index. Created atomically.
    • Snapshot: exclusions are :(exclude,glob) pathspecs. Then add -A, write-tree, commit-tree, refs/turns/<turnId>, and a line in the daemon-owned turns.log (order, plus the late flag).
    • Listing: only refs that match turns.log are trusted.
    • Housekeeping: prune by count and by storage budget, with batched repack -a -d + prune. Also copyCheckpoints (fork, via fetch), deleteCheckpoint(s), and sweepCheckpoints (unknown session ids, run at resume).
  • src/daemon/session.ts:
    • Turn ids: #currentTurnId is stamped by #makeMessage. In #sendInner: joinsRunningTurn for mid-turn pushes; hooks, then snapshot, then stop-check, then the adopted-turn check, then runTurn; a try/catch hands attribution back.
    • Stop: SendStoppedError via #stopGen / #preparingSends (and preparingTurn for drain).
    • Turn index: #turnIndex / #indexTurn (capped at 1000).
    • History and forks: restoreTurns, fullCanonicalHistory, inheritTurns.
  • src/daemon/providers/canonical.ts: CanonicalTurn gains turnId / prompt / at / background. The accumulator emits onChange (append / replace), and restore() loads without re-emitting.
  • src/daemon/transcript.ts:
    • <id>.canonical.jsonl: serialized at call time, field caps plus a 1 MiB whole-turn cap, compaction at 32 MiB, tail reads that widen until they hold a prompt, 0600.
    • <id>.turns.jsonl turn index.
    • A shared #chain helper. This also fixes append()'s stored promise chain, which leaked an unhandled rejection on a failed write and never cleared itself.
  • src/daemon/session-manager.ts:
    • Resume: restore from a 4 MiB canonical tail plus the turn index. Legacy sessions are rebuilt from the full transcript.
    • Fork: fullCanonicalHistory + inheritTurns.
    • New request: session.turns (attach/watch scope + ownership).
    • Stopped sends: SendStoppedError handled in #send, in the pipeline phase driver and in fleet event delivery.
    • Shutdown: drain() counts sends that are still preparing.
  • Protocol: packages/protocol/src/turns.ts (session.turns / session.turns.result, TurnSummary) and SessionMessage.turnId.
  • src/daemon/canonical-restore.ts: canonicalFromTranscript, turnIndexFromHistory.

Known edges

  • After a stalled-run recovery, or when a prompt joins a backend-started turn, the already-broadcast user row keeps the id first minted for it.
  • fsmonitor and the untracked cache are deliberately off for snapshots, so a snapshot walks the tree.

🤖 Generated with Claude Code

saucam and others added 5 commits October 10, 2026 10:55
…history that survives restart

Foundation for going back a turn, forking from history, comparing backends
and per-turn diffs (#354). Backend-agnostic: everything comes from codeoid's
own records, never a backend's native session.

- Every turn gets a turnId, stamped on each SessionMessage produced while it
  runs and on the canonical history (user and assistant turns).
- At turn start in a git workdir, the working tree (tracked + untracked,
  non-ignored) is snapshotted under refs/codeoid/checkpoints/<session>/<turn>
  via a throwaway index and commit-tree — branch, index, stash, working tree
  and git log untouched. Bounded (untracked size/count caps, timeouts, a 2s
  max wait before the turn starts), pruned per session, deleted on destroy.
  Config: session.checkpoints.
- The canonical history is persisted per session (<id>.canonical.jsonl) and
  restored on resume; sessions from before the log are rebuilt once from the
  transcript. Before this, after a restart a fork or backend switch carried
  no conversation, and stateless backends answered with no memory of it.
- session.turns lists a session's turns with their snapshots.
- A Stop (or drain) while a send is still preparing its turn now cancels it;
  the turn used to start anyway after the Stop.
- TranscriptStore.append's stored chain no longer leaks an unhandled
  rejection on a failed write, and clears itself.

Closes #354

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ounded history log

Audit round 1 on #354:

Security (High): snapshots ran git inside the user's repository, so its own
config ran code on every turn outside the approval gate — core.fsmonitor,
filter drivers, reference-transaction hooks — with the daemon's environment.
Checkpoints now live in a daemon-owned shadow repository per session
(<transcriptDir>/checkpoints/<session>.git, GIT_WORK_TREE = the workdir):
config comes only from the shadow repo (system/global off), hooks and
fsmonitor are forced off, the environment is an allowlist with no GIT_*.
Nothing lands in the user's repo (no pushable refs, no leftover objects),
destroy is rm -rf, the user repo's objects are an alternate for dedup, the
snapshot is scoped to the session's directory, common secret files and
dependency dirs are excluded, it works in non-git directories, refs are
trusted only when they match the daemon's own turns.log, storage is capped
per session, one snapshot runs at a time, and unknown sessions' repos are
swept at resume.

Correctness:
- A Stop before a send's turn starts throws SendStoppedError: a pipeline
  phase or background wake waiting on that turn no longer hangs or drops
  reports; SessionManager#send doesn't report it as a failure.
- A message sent mid-turn joins the running turn instead of stealing its
  output; no snapshot wait on that path.
- The snapshot is taken at the last moment before the agent starts; one
  still running when it starts is recorded as late.
- A turn index separate from the canonical history: rotation no longer
  empties session.turns; refused/stopped sends aren't listed and hand
  attribution back.
- Forks inherit the parent's turns and snapshots, which outlive the parent;
  forks after a restart use the whole history log.
- The history log caps oversized fields, compacts past 32 MiB, is read with
  a 4 MiB tail on resume, serializes at call time, and is owner-only.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…join backend turns after the snapshot

Audit round 2 on #354:

- Checkpoint repos are fully self-contained: no copied user index (an agent
  could plant skip-worktree entries naming any blob; assume-unchanged and
  split-index broke snapshots) and no alternates (a rewrite + gc in the
  user repo corrupted snapshots). A snapshot holds exactly the work tree's
  bytes, stored in the shadow repo. First snapshot is capped by total size
  and file count, later ones by new files; a missing index (a fork's copied
  history) is a first snapshot, so forks keep snapshotting.
- Exclusions are :(exclude,glob) pathspecs — a .gitignore negation can't
  pull .env or keys back in, and tracked secret files stay out — with a
  longer list; the daemon's data dirs are excluded too.
- Shadow repos are created atomically; a stale index.lock is cleared; a
  nested repository with no commit no longer fails every snapshot; over
  the storage budget the oldest snapshots are dropped (down to none) so the
  latest always fits; space is reclaimed in batches with repack + prune;
  turns.log rewrites are atomic.
- The "did the backend start its own turn?" check runs after the snapshot
  wait: a turn started in that window is joined instead of having its
  queue closed by a second runTurn.
- The history log can't compact or tail-read to empty: tail reads widen
  until they hold a prompt, and a whole turn is capped (1 MiB).
- Sessions from before the log are rebuilt from the whole transcript, once.
- A snapshot taken for a send that never became a turn is deleted; fleet
  event delivery survives a stopped send; a stopped phase reports that; the
  turn index is capped; a fork inherits only turns its history contains.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The env-leak test sets GIT_DIR / GIT_INDEX_FILE / GIT_OBJECT_DIRECTORY on
process.env; its own git show then inherited the decoy object directory,
which newer git (CI) honours over --git-dir.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@saucam
saucam merged commit a0ac66d into main Oct 10, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Sessions have no durable turn history: every turn needs a stable id and a snapshot of the files it started from

2 participants