Skip to content

Cost pipeline (cost-aggregator.ts + UsageAggregator.ts) undercounts long sessions and overcounts short ones #2284

Description

@jacobo-ortiz

Observed: the Performance and Usage tabs show roughly half of what the transcripts price to. Five independent defects, measured on one install over ~9,000 transcripts (June to October 2026, Claude Code 2.1.x):

  1. A session row is frozen the first time it is seen (PULSE/Performance/cost-aggregator.ts: ids already present are skipped, and files are skipped by mtime). Sessions that keep growing after their first sighting stay at their first value. 15.9% of rows had a transcript that later grew; on those rows the recorded cost is 0.38x the deduplicated cost. Long sessions are the expensive ones, so this is where most of the gap comes from.

  2. Usage is summed per transcript line, not per request. One API response is written as N lines (one per content block) with the same usage repeated. Lines per request measured: main thread 2.06, subagents 2.37, workflows 2.20. On rows that were not frozen, recorded / deduplicated = 2.22x. Same pattern in TOOLS/UsageAggregator.ts. Dedup key that works: requestId ?? message.id, last line wins with the max output_tokens (subagent lines grow between blocks).

  3. Workflow transcripts are never read. <sid>/subagents/workflows/<wf>/agent-*.jsonl is one level deeper than both walkers go. Measured: 13.6% of all requests and 13.4% of the priced cost.

  4. Every cache write is priced as the 5-minute write (1.25x input). 99% of main-thread cache writes are 1-hour writes (2x input); usage.cache_creation.ephemeral_1h_input_tokens / ephemeral_5m_input_tokens are in every transcript line but the row does not keep the split, so history cannot be repriced.

  5. Days are cut in UTC, and each tab assigns a whole session to a different day. usage-daily uses the session's first timestamp, Performance its last one; "today" and the ISO week are UTC. For a principal at UTC-5, 16.3% of cost lands on the wrong local day, and the two tabs disagreed by three orders of magnitude for the same "today" in one live check.

Cross-check that made us trust the recomputation: the harness writes its own per-session total (cost-state lines in the transcript). Recomputing from deduplicated requests with the official price table matches it within 6.4% over a 30-day window, and the remaining gap is explained by calls that leave no transcript line (WebSearch, measured; others unmeasured). Every cost-state row with n ≥ 30 per model falls inside the band "all writes 5-minute" to "all writes 1-hour" of the official table.

Suggested direction: regenerate both ledgers from the transcripts on each run (dedup globally, walk workflows, keep the 5m/1h split, price from one table), keep every existing field's meaning so the tabs need no change, and add the day in the principal's timezone as a separate series. We run that locally as a replacement tool and can share the reader if useful.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions