Skip to content

feat: compare backends side by side — one prompt, 2–4 forks, keep the best - #364

Merged
saucam merged 3 commits into
mainfrom
feat/compare-357
Oct 10, 2026
Merged

saucam merged 3 commits into
mainfrom
feat/compare-357

Conversation

@saucam

@saucam saucam commented Oct 10, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #357. Builds on #363 (fork from a turn), #362 and #361, all merged.

What

Send one prompt to 2–4 backends/models at once and compare what each did, then keep the best. Each branch runs in its own fork and git worktree, starting from the same conversation and files: the latest point, or after an earlier turn.

For each branch you see:

  • status, including needs your approval (branches ask in their own session);
  • its reply to the prompt;
  • files changed (+/−, paths);
  • cost and time.

Keep one, optionally destroying the rest.

  • Web: "⚖ compare" in session controls, "⚖ compare from here" on a message. The dialog shows the branches in columns, with open, keep, and keep + discard others.
  • CLI:
    • codeoid compare run <s> --with claude,codex:gpt-5.5 [--at N] [--no-wait] <prompt…>
    • codeoid compare show <id> [--wait]
    • codeoid compare ls <s>
    • codeoid compare keep <id> <branch> [--discard-others]
  • TUI: /compare <targets> [--at N] <prompt>, /compare, /compare keep N [--discard-others].
  • Protocol: session.compare, compare.get, compare.list, compare.keep → compare.state / compare.list.result.

How (backend-agnostic)

It is pure orchestration over session.fork and send, so it works with every current and future backend. The comparison record lives in the store (compare_runs).

When a branch's compared turn settles, its result is frozen from codeoid's own records:

  • the reply comes from the canonical history for that turn only;
  • the files come from the start→end snapshot numstat;
  • the cost and time are that turn's totals.

Carrying on in a branch later never changes the comparison, and reads stay cheap.

Safety and robustness

  • Refused up front, leaving nothing behind:

    • a model the backend doesn't have, or one that isn't a model id;
    • a parent that's still working;
    • a folder that isn't a git repo, since every branch needs its own worktree.

    Any failure while forking rolls back the branches made, along with their git branches.

  • Request handling: the request answers as soon as the branches exist, and the sends run without the client waiting.

  • Keeping: one keep per comparison, and only a finished branch can be kept. A slow settle can't undo a keep.

  • Cleanup: comparisons are deleted with their session.

  • Forks: they now carry the parent's pack, re-resolved, so a role's tool deny survives. A pack that can't be re-activated refuses the fork.

  • Unrelated fix in codeoid approve: it always sent an empty approval id, which the daemon rejected. It now shows the pending tool call in full and asks before approving (--pick N, --yes).

Testing

  • src/tests/compare.test.ts (12) uses the real SessionManager with mock backends. It covers:
    • replies, files, and cost per branch;
    • comparing from an earlier turn;
    • the same backend with two models;
    • a failing branch, with no inherited reply;
    • frozen results;
    • keep and discard, and the second-keep refusal;
    • the up-front refusals with no leftovers (model, flag-like model, busy parent, non-git);
    • destroy and restart cases;
    • the keep-vs-slow-settle race;
    • no resurrection after the parent is destroyed;
    • gone and never-started branches.
  • src/tests/terminal-compare.test.ts (7) and web CompareModal.test.tsx (4) cover the clients.
  • Full gate passes: typecheck, biome, bun test (2931 pass), web vitest (593 pass), build.
  • Live E2E, on a real daemon with real backends, through the CLI:
    • claude vs codex:
      • Approvals are surfaced, and codeoid approve shows the exact command and asks first.
      • claude's result: its own reply, fib.py +8, $0.18, 14s.
      • codex reported its own failure honestly. Its sandbox can't start inside this test environment, which is unrelated to compare.
      • Keep with discard removed the other branch and its worktree, and a second keep was refused.
    • claude vs claude:sonnet from turn 1: both branches started from that turn's files and each appended its line, while the parent's files stayed untouched.

Audits

Two deep rounds, each covering correctness and security; all fixes are folded into the two fix: commits.

One finding is noted but not changed: a follow-up sent to a branch during its compared turn joins that turn, as any mid-turn message does.

🤖 Generated with Claude Code

@saucam
saucam force-pushed the feat/fork-at-turn-356 branch from e8f9037 to 00e7549 Compare October 10, 2026 17:30
@saucam
saucam force-pushed the feat/fork-at-turn-356 branch from 00e7549 to fe7bd9a Compare October 10, 2026 17:40
saucam and others added 3 commits October 11, 2026 01:46
… best

Closes #357. Pick 2–4 backend/model targets and a prompt: each runs in its
own fork (own worktree) from the same conversation and files — the latest
point or after an earlier turn — then compare status, reply, files changed
(start→end snapshot numstat), cost and time, and keep one (optionally
destroying the rest). Pure orchestration over session.fork + send, so it
works with every backend; runs are recorded in the store (compare_runs).

- Protocol: session.compare, compare.get, compare.list, compare.keep →
  compare.state / compare.list.result.
- Web: "⚖ compare" in session controls and "⚖ compare from here" on a
  message open the compare dialog (columns, keep / keep + discard).
- CLI: codeoid compare run <s> --with claude,codex:model [--at N] [--shared]
  [--no-wait] <prompt>; compare show <id> [--wait]; compare ls <s>;
  compare keep <id> <branch> [--discard-others].
- TUI: /compare <targets> [--at N] [--shared] <prompt>, /compare, /compare keep N.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…proval, and never half-start

Round-1 audit of #357 (and a live run), folded in.

- Each branch's result is frozen when its compared turn settles — that
  turn's reply (never the parent's inherited one), files (after its end
  snapshot lands), cost and time — so carrying on in a branch later changes
  nothing, and list/poll is cheap (no git per read).
- The request answers as soon as the branches exist and starts them all at
  once (a send can wait on setup longer than a client waits); the record is
  saved first.
- Refused up front, leaving nothing behind: a model the backend doesn't have
  or that isn't a model id, a parent that's still working (every branch must
  start from the same point), a folder that isn't a git repo. Any failure
  while forking rolls back the branches made; a branch that can't get its
  own worktree aborts it (shared folders made the file counts lie). The
  shared-folder option is gone.
- A branch waiting on an approval says so — "needs your approval" in the
  web, CLI and TUI, with how to give it — instead of "working…" forever;
  `codeoid approve` now sends the real approval id (it sent an empty one).
- One keep per comparison (a second "keep, discard others" would destroy the
  first); only a finished branch can be kept; failed destroys are reported.
- Forks carry the parent's pack (a role's tool deny included).
- Comparisons go with their session when it's destroyed; model ids are
  charset-checked and errors sanitized for terminals; branch names capped.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…eir session, and approve shows what it approves

Round-2 audit of #357, folded in.

- Settles in flight are tracked: a comparison isn't retired (and re-read as
  a second copy) while a branch is still being frozen, and a save never
  clears a keep — so a second "keep, discard others" can no longer destroy
  the branch kept first.
- After creation a comparison is only ever updated, never re-inserted, and a
  destroyed session's live comparisons leave memory: destroying the parent
  mid-run no longer brings the record (and its prompt) back.
- A destroyed branch stops being waited for (reads "gone"); a branch whose
  prompt never ran because the daemon restarted reads "failed" instead of
  "working…" forever.
- Rolling back a comparison that couldn't start deletes its branches' git
  branches too.
- Forks re-resolve the parent's pack instead of reusing it: a pack since
  untrusted or removed doesn't ride along, and a fork that can't carry it is
  refused rather than losing a role's tool deny.
- Model ids reject ".." too.
- `codeoid approve` shows the tool call in full (its input, as plain text
  the agent can't restyle), says how many are waiting (--pick N), and asks
  before approving (--yes to skip; required without a terminal).
- TUI: /compare reads the comparison it last started/showed (fresh state,
  not another client's newer one), loads the branch sessions, and keep says
  which comparison it kept from.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@saucam
saucam changed the base branch from feat/fork-at-turn-356 to main October 10, 2026 17:46
@saucam
saucam merged commit 35c644f into main Oct 10, 2026
4 checks passed
saucam added a commit that referenced this pull request Oct 11, 2026
* fix: approve shows what the gate is actually waiting on (#349)

#364 made `codeoid approve` work by attaching to the session and
scanning its scrollback for tool calls in `waiting_confirmation`. That
reads the ANNOUNCED tool message, which is not always what is parked:

- A backend that ties a gate to its announcement by tool name alone
  (qwen) can show one call while another is parked, and approve now
  presents what it shows as exactly what will run.
- A leftover `waiting_confirmation` row from a past run (after a daemon
  restart) looked pending; approving it printed "Approved." while
  nothing happened.
- The scan gave up after 3s on a long scrollback, and attaching just to
  read had side effects.

The daemon now answers `session.approvals {sessionId}` with only the
calls its approval gate is really waiting on, each with the input the
gate received (after hook rewrites). It needs session:watch or
session:attach, like the scrollback it replaces. `session.approve`
reports whether the answer took effect (`data.outcome`: settled / early
/ stale), still as response.ok so the web ApprovalBar's dismissal of
leftovers is unchanged.

The CLI:
- uses that list; adds --id (stable, unlike --pick, for scripts)
- exits 1 whenever it does not answer, and reports a stale answer
- refuses to approve on a partial view when the input is longer than
  it shows, instead of silently cutting it
- shows C1 controls, bidi and other invisible characters as visible
  escapes on the lines a decision rests on, dialogs included
- treats Ctrl-D at the prompt as "no" instead of hanging

Closes #349

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: stricter approve lines, refuse an empty --id, state the backend limits

Security review of the session.approvals change:

- An empty `--id` (an unset $ID in a script) fell back to the oldest
  waiting call. It is refused now.
- Escape every Unicode default-ignorable and format character on the
  lines a decision rests on (variation selectors, Hangul fillers, ...),
  not a hand-picked list.
- Single-line fields (tool name, id, dialog title) escape newlines, so
  a crafted tool name can't print a fake input block above the real one.
- The protocol said the listed input is exactly what runs. Not on every
  backend: codex and the Gemini CLI ignore a hook rewrite, and codex
  ties an MCP elicitation to its call by server name. Say so; those are
  backend fixes, tracked separately.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep escaping U+2028/U+2029 on approve lines

The Unicode-property rewrite dropped the line and paragraph separators
(Zl/Zp, not Cf or default-ignorable), which most terminals draw as
nothing. Add them back, with a test.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Run one prompt on several backends side by side and keep the best result

2 participants