Repository navigation
fix(server): queued runs land on the timeline where they started - #18094
juliusmarminge wants to merge 2 commits into
Conversation
A run's turn items were banded by its run ordinal, which is fixed when the run is queued. A message queued before a steer therefore rendered above the steer even though the steer's run started first (#16987). Each run now gets its own timeline band the first time it writes an item, which is when it starts. TurnItemPositionStore owns the new orchestration_v2_run_timeline_bands table. Migration 061 and the projection rebuild both derive a run's band from the positions its items already hold, so existing threads keep their order. The run ordinal itself is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Note GPT 6.1 Sol responding on behalf of voltcrash @juliusmarminge would you consider merging #17773 for the queued-run ordering fix instead? Both address the same transcript bug, while #17764 already handles queued runs disappearing on the client. #17773 moves a queued run's ordinal forward before it starts when newer work has already started. It keeps run IDs unchanged and preserves the execution-order model used by existing ordinal consumers, without adding a table or migration. The separate timeline bands in this PR are a reasonable design, but the conversion is incomplete. As the description notes, the transcript can show A → C → B while a fork through C still includes B. Checkpoint restore needs consideration too: CheckpointRollbackService selects later runs using I'd favor #17773 with focused held-queue → resume → fork/restore regression coverage before merging. The timeline-band approach could then be reconsidered once the remaining ordering consumers are reconciled. Is there a constraint requiring queued run ordinals to remain immutable that makes #17773 unsuitable? |
|
Note 🤖 Claude Fable 5.1 on behalf of Mnigos I hit this bug myself (a steer promoted from the queue into a held PR-watch run that started after an automatic continuation, the #16987 pattern), so I read this PR and #17773 (#17773) against that case. Both fix it for future runs; neither repairs already-recorded history, which is fine. Two things I found here, from source review plus an isolated SQLite check, not a runtime run:
For what it's worth, #17773's dequeue-time ordinal reassignment is the smaller model, since fork cuts and ordinal-based "latest run" lookups follow it without a second source of order; it has the same prewritten-item gap and would need the same test. |
|
Additional real-thread evidence for the ordering defect tracked in #16987 (read-only inspection on macOS, 2026-10-11, Claude provider):
This matches the start-order problem this PR addresses. I checked current main's Separately, a queued run 115 was cancelled before starting: requested 13:12:56.565 UTC, cancelled 13:13:04.299 UTC, no Run 105 was cancelled at 13:12:09.253 UTC and followed by continuation runs 114 and 116. These records alone do not prove whether Claude read Inspected with Codex (gpt-6.1-sol) in T3 Code; no live database writes or application changes. |
Server half of #16987. #17764 already fixed queued runs vanishing on the client.
Problem
Queued runs render above a steer that ran first. Every turn item's position sits in a band of its run's ordinal (
ordinal * 1_000_000 + n), and the ordinal is fixed when the run is queued. Take a queued message B (run 2) and then a steer C (run 3). C starts first, but B's items still sort under run 2, so B shows above C even though it ran after it.Fix
orchestration_v2_run_timeline_bandstable owned byTurnItemPositionStore.EventSinkno longer passes run ordinals in.throughRunOrdinaland provider-thread ranges are all derived from them.ProjectionMaintenancerefills the bands from stored positions with the same query.ordinal * 100 + nplaceholders in turn-item payloads (queued and direct start, handoff, checkpoint, run signals) are left alone. Every write goes throughnormalize, which replaces them with the next position in the run's band. A queued run writes nothing until it starts, so its message lands in its start-order band.Tests
runtimeLayer.test.ts: "places a queued run that starts late after the runs that started before it".expected ['A','B','C'] to deeply equal ['A','C','B'].061_RunTimelineBands.test.ts: the backfill keeps stored bands, and a run started afterwards goes above all of them.apps/servertypecheck.Not covered yet
Fork "through run" cuts still compare run ordinals:
visibleTurnItemsThroughRun, the timeline index,itemCountThroughRun,fork_boundary, and the portable-fork handoff list;The fix compares bands at those sites. Those are the same functions #17763 changes, so it follows as a separate PR once #17763 lands.
🤖 Generated with Claude Code (Opus 5.5)