Skip to content

docs(qa): refresh qa/SKILL.md, keep every rule and drop the history - #855

Merged
atomchung merged 1 commit into
mainfrom
claude/qa-skill-refresh
Oct 5, 2026
Merged

atomchung merged 1 commit into
mainfrom
claude/qa-skill-refresh

Conversation

@atomchung

@atomchung atomchung commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

What

qa/SKILL.md is the fomo-qa skill that Claude, Codex, Antigravity and Grok load as instructions (all four discovery links point at this directory). Next to its rules it carried dated decisions, incident and audit write-ups, issue and PR pointers used as provenance, "used to" and "now" notes, a remark about a removed flag, and superseded examples. This rewrites each in place so the rule stands alone. No rule is added, dropped or weakened, and nothing outside qa/SKILL.md changes. The stories move into this description.

Numbers

before after
lines (wc -l) 661 659
bytes 66,022 62,613
frontmatter description 956 chars 450 chars
lines matching grep -nE '20[0-9]{2}[-/][0-9]{2}' 19 4 (listed below)
lines matching grep -nEi 'historical|legacy|previously|used to|deprecated|歷史|舊版|以前|曾經|拍板' 4 2 (the literal label legacy-unattributed, below)
distinct issue/PR numbers 26 1 (#486)
old lines carrying a rule keyword 116 116 kept: 77 unchanged, 39 reworded, 0 lost

The description now says what the skill does, its triggers (the Chinese trigger phrases stay on that one line) and the out-of-scope boundary with the pointer to the product skill fomo-kernel.

The workspace skill lint reports 0 errors and 5 warnings (before: 13 errors and 5 warnings). The warnings are the length, the two as-of dates and the two legacy-unattributed lines, all explained below.

Length. 659 lines against the checklist's 200. 299 of them are fenced commands, and 209 of those sit in seven # qa-trace: fences that qa/tests/test_skill_commands.py replays through ux_receipt.py verify. The order of those commands is a contract (a wrong order voids the run and the trace is append-only), and the prose between them is the per-route rule behind that order, so none of it can go without changing what an operator does. Splitting the file into references is a separate decision, not part of a history refresh.

Repository rules checked first

  • docs/expression-contract.md §3.4 ("Where a local answer-shape phrasing was superseded by this chapter, the surface document says so rather than deleting its own history") does not cover this file. It governs answer-shape phrasing on the contract's own surfaces (§6: review card, consider, freeform, no-book, weekly read, plus output-voice.md's V1), and asks only that a surface that once had a local answer-first phrasing say it was superseded. qa/SKILL.md is a maintainer procedure with no answer-shape phrasing and is not a surface. I found no other repository rule that asks a skill, flow or procedure to keep its history; the history-bearing documents are ruling logs (docs/expression-contract.md §8, docs/output-contract.md), the maintainer guide's receipt rows, and V1 in docs/output-voice.md, none of which is a procedure an agent follows.
  • What does bind it is qa/tests/test_skill_commands.py, docs/language-policy.md (qa/ is English; CJK is allowed only on the description: line and the - The user says bullet), docs/development-guide.md §6 (find the enforcing test before cleaning; read and classify "never" lines instead of counting them) and the maintainer-guide row for the isolation gate, which names this file's Step 0 and guardrail 2 as readers. All of those stay intact: the gate count and its seven-name roster, every # qa-trace: fence and its replay, the placeholder map, the "Continuing in the same campaign" section, guardrails 1 to 5.

Kept on purpose

where why it stays
line 407: As of 2026-07-21 the date of a tool-behaviour fact (generic visualization tools normalize a supplied <style>); it tells an agent when to re-verify
line 464: as of 2026-07-29 the date of the case-id lists in acceptance issue #486 (its checklist has since grown to M0-T11), so the lists read as a dated example and the next sentence sends the operator to the issue
lines 589, 615: 2026-08-14 example values in the --challenge-check-file template; a test replays this JSON
lines 477, 651: legacy-unattributed the literal label report prints (qa/receipts.py), a live rule
#486 the live acceptance issue an operator must read before archiving

Left untouched because they are not history: the M0-U01, M0-U02 example ids in the lineage table (they disagree with the M0-F ids two paragraphs above; example values stay) and the count "three things" before a four-row table.

What was removed

Source commit = the oldest commit whose diff adds the text (git log -S, restricted to qa/SKILL.md). 58d46f5 is the commit that translated the Chinese import 0957af0 into English, so the oldest stories (the 2026-07-19 to 2026-07-27 dates, the #230/#236/#238/#293 notes) were already in that import.

edit kind what was removed source
E01 rationale Design rationale in the frontmatter description: the purpose (kill the rework from every session testing a different environment), the mechanism list (version gate, detached worktree, simulated new user, … 58d46f5
E02 date Dated version stamp "v1, fixed 2026-07-20" 58d46f5
E03a date Start date of the runbook as contract source and the PR that introduced it 58d46f5
E03b incident Dated account of when the seventh gate was added and why 58d46f5
E04 date Start date of registry discovery 58d46f5
E05 incident "Why it exists" block: 2026-07-19 audit that found 17 of 18 worktrees behind main (issue #250) 58d46f5
E06 date "had always lacked" framing, eval's current "pending owner dogfood" marking, "Since 2026-07-27 it produces one more thing", and the quoted "stable enough to keep" gloss 58d46f5
E07a date Version stamp "v1, fixed" in the heading 58d46f5
E07b used-to "This round" release framing 58d46f5
E07c incident Pointer to #230 as the origin of the false-pass trap 58d46f5
E08 superseded "next round" roadmap note and the as-of remark about test_card_html.py 58d46f5
E09 incident Pointer naming #230 as the ceiling 58d46f5
E10 superseded "next round" roadmap note 58d46f5
E11 incident Incident #557: an agent read the real ledger on its own initiative (kept as the property of the rule) 2b0c07f
E12 incident Incident #274 attribution 58d46f5
E13 pointer Issue pointer #557 2b0c07f
E14 incident "the failure it was written against" (past failure as justification) 2b0c07f
E15 pointer Issue pointer #544 2a0e597
E16 used-to When ux_receipt began honoring TRADE_COACH_HOME (#269 fix, PR #275) 58d46f5
E17 date Dated "Lesson from 2026-07-20" label 58d46f5
E18 used-to "since #544 it is the documented one" 2a0e597
E19 pointer Issue pointer #357 58d46f5
E20 superseded Note that the removed --question-id flag was replaced 58d46f5
E21 pointer Issue pointer #236 as origin of the measurement 58d46f5
E22 pointer Issue pointer #293 58d46f5
E23 pointer Issue pointer #293 as the name of the mechanical check 58d46f5
E24 pointer Issue pointer #663 c217228
E25a wording Wording that tripped the history-word grep ("used to choose"); no story removed c217228
E25b used-to "now" (when verify began refusing it) c217228
E26 pointer Issue pointer #663 c217228
E27 used-to Authoring-choice narration ("rather than pointing back at Step 5 precisely because") 58d46f5
E28 incident How the differences were found ("read off a real run rather than assumed") 889c8fd
E29 pointer Issue pointer #357 used as the pre-flight's name 889c8fd
E30 incident "observed ... came back" (one past run) 889c8fd
E31 incident Past tense of one observed run ("returned ... on each candidate"); the present tense is made precise: a candidate carries grounding only when the engine has one 889c8fd
E32 pointer Issue pointer #236 889c8fd
E33 pointer Issue pointer #293 889c8fd
E34 pointer Issue pointer #530 and "observed verbatim" 889c8fd
E35 incident Past tense of one observed rerun 889c8fd
E36 incident Anecdote of the one observed run 889c8fd
E37 pointer Issue pointer #520 ("the lineage #520 added") 889c8fd
E38 date Dated owner_live audit correction as the source of the rules 58d46f5
E39 incident The 2026-07-20 deviation (under-declared card_modes, zero widget attempts, #249's card generated but the owner saw flat Markdown, the main reason card=fail) and the closing "again" 58d46f5
E40 incident "New lesson, 2026-07-21 (see the #230 comments)" label, and the mis-diagnosis incident (wrong tool reported as "host has no rendering capability", posted to GitHub, corrected afterwards). The tool-behavio… 58d46f5
E42 incident Incident: forcing en in the 2026-07-20 mock session caused mixed-language output (#262) 58d46f5
E43 pointer Issue pointer #236 ("re-measurement instrument") 58d46f5
E44 incident Dated provenance of the trap list (two consecutive walkthroughs hit one each) 58d46f5
E45 pointer PR #298 pointer ("merged") 58d46f5
E46a pointer Pointer to open issue #337 58d46f5
E46b pointer Pointer to #230 as the purpose the rule serves 58d46f5
E47 pointer Issue pointer #238 58d46f5
E48 incident Recorded wait "#236's 5–10 minute wait" 58d46f5
E49a date Dated lesson pointer (2026-07-21, #293) 58d46f5
E49b pointer Issue pointer #293 58d46f5
E50 date "since its 2026-07-29 ruling": the decision date is dropped; the id lists stay with 2026-07-29 as their as-of date (the issue's checklist has since grown to M0-T11) 2a0e597
E51a pointer Issue pointer #520 58d46f5
E51b used-to "now" (when archive began enforcing) 58d46f5
E53 superseded Deferred-work pointer to #492 ("if operational evidence ever justifies it") 58d46f5
E54 pointer Issue pointer #417 7c39ae6
E55 used-to #544 history: walking four routes used to cost four full ceremonies in four sessions 2a0e597
E56 incident Incident #543: ad hoc portfolio question answered outside any lifecycle, 34 turns and an unrequested chart 2a0e597
E57 pointer Slice/issue pointers #544 and #479 83f5cd6
E58 used-to "since #767" (when sector_display joined the check file) 757bb3a
E59a pointer Issue pointer #830 98a3c00
E59b pointer Issue pointer #739 7f460e1
E60 pointer Issue pointer #767 757bb3a
E61 superseded The retired "exploratory only" caveat and the "now"/"still" framing 83f5cd6
E62 pointer Issue pointer #230 ("core lesson") 58d46f5

Two of the removals corrected a stale claim rather than only dropping history: eval's layer "currently marks 'pending owner dogfood'" (no such marker exists in docs/eval-design.md or evals/EVALS.md) and the test_card_html remark about zero screenshots.

Rule preservation

Every old line that carries must / never / always / do not / only / refus* (or the Chinese equivalents) is mapped to its new line. Old lines with a rule keyword: 116. unchanged: 77; kept (reworded): 39; story only: 0; lost: 0.

The full map (117 rows)
old line keyword(s) old text (start) new line verdict
3 never, only, refuse description: Prepares a clean, consistent QA environment for dogfooding fomo-kernel an… 3 kept (reworded). Removed: E01: Design rationale in the frontmatter description: the purpose (kill the rework from every session testing a different environment), the mechanism list (version gate, detached worktree, simulated new user, data sources) and …. Rules kept elsewhere: the description's mechanism list duplicated body rules: version gate/"refuse when behind" -> new L74 ("If it is behind, do not go on"); worktree "never used for development" -> new L84; dogfood-only root, never the real ~/.trade-coach -> guardrail 2, new L37; "Never touches the real records in investment_note" -> guardrail 1, new L36.
10 do not **This is the mandatory standard path for every fomo-kernel dogfood (v1, fixed 2026-07… 10 kept (reworded). Removed: E02: Dated version stamp "v1, fixed 2026-07-20".
12 must **Cross-client contract source (since 2026-07-21)**: docs/qa-runbook.mdin thekol_…` 12 kept (reworded). Removed: E03a: Start date of the runbook as contract source and the PR that introduced it; E03b: Dated account of when the seventh gate was added and why.
18 always > **Its place in the eval system**: this is also the execution procedure that docs/ev…` 16 kept (reworded). Removed: E06: "had always lacked" framing, eval's current "pending owner dogfood" marking, "Since 2026-07-27 it produces one more thing", and the quoted "stable enough to keep" gloss.
26 only **Not** for reviewing a real user's trades — that is the product skill fomo-kernel. … 24 unchanged
28 do not ## Coverage (v1, fixed — do not claim beyond it) 26 kept (reworded). Removed: E07a: Version stamp "v1, fixed" in the heading.
30 do not This round verifies **L1: environment consistency + engine CLI contract + agent walkth… 28 kept (reworded). Removed: E07b: "This round" release framing; E07c: Pointer to #230 as the origin of the false-pass trap.
32 never, only - **L2 card visuals** (next round): the card HTML is never actually rendered in a brow… 30 kept (reworded). Removed: E08: "next round" roadmap note and the as-of remark about test_card_html.py.
33 only - **L3 interaction delivery** (partly an inherent ceiling): "the option buttons really… 31 kept (reworded). Removed: E09: Pointer naming #230 as the ceiling.
38 never, only 1. **Never touch the real records**: ~/Side_project/investment_note/ holds ting's re… 36 unchanged
39 always, never, only, refuses 2. **Coach state is isolated to a dogfood-only root, and the real one is not reachable… 37 kept (reworded). Removed: E11: Incident #557: an agent read the real ledger on its own initiative (kept as the property of the rule).
40 only 3. **Work only in the dogfood worktree**: every engine command runs inside the detache… 38 unchanged
41 do not 4. **Do not change product code**: QA reads, it does not edit. If the walkthrough find… 39 unchanged
42 must, never 5. **Public text passes the privacy lint first (bought by the #274 incident)**: the re… 40 kept (reworded). Removed: E12: Incident #274 attribution.
48 only Only exit 0 may be posted. On a hit, rewrite as a de-identified description ("N indivi… 46 unchanged
52 only **Cross-client execution gaps** (being in the discovery registry only guarantees the s… 50 unchanged
54 must, only - If Step 4 presents questions through Claude's native option tool (for example AskUs…` 52 unchanged
55 do not, must, only - The rendering-pipeline test mentioned in walkthrough rule 1's "try the widget once" … 53 unchanged
56 only - qa_env.sh's assumptions about the current working directory and worktree have not … 54 unchanged
60 only ### Step 0 — Isolate this shell, then the version gate (read-only) 58 unchanged
62 refuses ``qa_env.sh **refuses every command** until the shell it runs in has the account's own… 60 kept (reworded). Removed: E13: Issue pointer #557.
70 do not, only It is a bounded guarantee, and reporting it as more than that is the failure it was wr… 68 kept (reworded). Removed: E14: "the failure it was written against" (past failure as justification).
76 do not At a glance: the latest origin/main sha, how far behind the dogfood worktree is, and… 74 unchanged
78 never, only ``status also reports one extra line, **this skill's own freshness** — checking the ch… 76 unchanged
86 never Creates (or refreshes) ~/Side_project/kol_collector/fomo-kernel-dogfoodat--detach…` 84 unchanged
102 only | **Receipt** | one route-specific, append-only evidence trace | exactly one per ro… 100 unchanged
104 never **Never merge route runs into one receipt.** A first_review owes two cards and a cas… 102 unchanged
106 only **Step 0's isolate already routed the whole toolchain into the dogfood-only coach ro… 104 kept (reworded). Removed: E16: When ux_receipt began honoring TRADE_COACH_HOME (#269 fix, PR #275).
118 do not, never - **Simulate a returning user** (runs weekly-review / due-revisit): do **not** reset. … 116 kept (reworded). Removed: E17: Dated "Lesson from 2026-07-20" label.
122 do not, must, refuses **This isolation must survive into every later shell** — if commands each start a new … 120 unchanged
128 only | **Real trades** (read-only) | ~/Side_project/investment_note/trades/fomo/trades.c…` 126 unchanged
130 only | **Test-drive** | --test-drive(no CSV) | Demonstration only;persist:false, z… 128 unchanged
134 do not ### Step 4 — Walk through (follow the product's fixed lifecycle; do not rewrite it) 132 unchanged
136 always After cd .../fomo-kernel-dogfood/skills/fomo-kernel(and confirming Step 2'sexport…` 134 unchanged
145 must # "no numbers" narrative (answers.json / narrative.json must pass their schemas) 143 unchanged
156 only **The UX receipt runs through the whole walkthrough (mandatory — this is what connects… 154 unchanged
158 must, never, only Below is a **complete, directly copyable first_review trace**. The order is a contra… 156 unchanged
162 must # 0) Declare host capability right after prepare. --adapter must match the capability 160 unchanged
164 do not # universal plain_text and markdown_inline fallbacks itself — do not pass them again. 162 unchanged
170 must # it must come before the first question and the first card — recording it later is 168 unchanged
174 never # ask": a run where the user was never offered it records nothing, and verify fai… 172 unchanged
178 never # 2) One row per question asked. Question text never enters the trace: a question from a 176 unchanged
180 must # the two must appear together (this replaces the removed --question-id). 178 kept (reworded). Removed: E20: Note that the removed --question-id flag was replaced.
188 always # 4) Cards are always "artifact first, presented second", and both rows need --stage 186 unchanged
208 never, only ``--grounding-check-file points at a **transient JSON that never enters the trace** (s… 206 unchanged
220 do not, only A candidate with no groundingomits the field entirely (likecandidate_1) — **do n… 218 kept (reworded). Removed: E23: Issue pointer #293 as the name of the mechanical check.
222 do not, only **The weekly_reviewroute carries one extra opener, andverify enforces it** (the … 220 kept (reworded). Removed: E24: Issue pointer #663.
225 always, never, only, refuses - provided— always triggersadd-cash and a recompute, so the card the user actual… 223 kept (reworded). Removed: E25a: Wording that tripped the history-word grep ("used to choose"); no story removed; E25b: "now" (when verify began refusing it).
238 do not # --memory-kind flag; these do not count as the opener): 236 unchanged
242 only # The FIRST preview (holdings-only) renders and is shown here, in one message with the 240 unchanged
245 only # artifact_generated/card_presented row of its own; only the cash answer is recorded n… 243 unchanged
262 must, only, refuses The wrap-up has the same shape as Step 5, except **--memorymust bepassorfail… 260 kept (reworded). Removed: E27: Authoring-choice narration ("rather than pointing back at Step 5 precisely because").
278 never Selected when the user has a position table or screenshot and no transaction history. … 276 unchanged
289 only - **No cash anchor row.** The route's contract does not carry the #357 pre-flight, bec… 287 kept (reworded). Removed: E29: Issue pointer #357 used as the pre-flight's name.
290 do not - **No question rows.** The observed plan came back with question_queue: []andcar…` 288 kept (reworded). Removed: E30: "observed ... came back" (one past run).
338 refuses Once a book exists, a newer holdings view is **not** a review. prepare --route snapsh…` 336 kept (reworded). Removed: E34: Issue pointer #530 and "observed verbatim".
341 only this holdings view has changes only you can settle before the recorded book can 339 unchanged
345 never So the real journey is composed — record, then review — and it produces **two receipts… 343 unchanged
348 only # step 1: read-only. Writes nothing; returns the frozen diff, a summary, and 346 unchanged
361 only - **A refresh creates no session**, so the trace is keyed by the engine's own refresh…` 359 unchanged
362 refuses - **No card events at all.** verifyrefusesartifact_generated, card_presented, … 360 unchanged
363 must, never, only - **The question row depends on what the engine raised, and so does the verdict.** Ste… 361 kept (reworded). Removed: E36: Anecdote of the one observed run.
364 must, only - **That --controlschoice is not recoverable.** Recording--controls pass on a re… 362 unchanged
377 only # 1) The engine's difference, as you narrated it. change_presented carries only 375 unchanged
383 only # 2) Only when step 1 came back pending_confirmation: the one question covering 381 unchanged
407 do not, must 1. **Declare capability honestly, and try the widget once per session — with the right… 405 kept (reworded). Removed: E39: The 2026-07-20 deviation (under-declared card_modes, zero widget attempts, #249's card generated but the owner saw flat Markdown, the main reason card=fail) and the closing "again".
409 do not, only **New lesson, 2026-07-21 (see the #230 comments)**: trying the widget does not mean gr… 407 kept (reworded). Removed: E40: "New lesson, 2026-07-21 (see the #230 comments)" label, and the mis-diagnosis incident (wrong tool reported as "host has no rendering capability", posted to GitHub, corrected afterwards). The tool-behavior claim keeps its ….
411 do not, must, only **This tool-selection detail belongs here and must not be promoted into fomo-kernel's … 409 unchanged
412 always 2. **--language follows the conversation language**: a Chinese conversation always u… 410 kept (reworded). Removed: E42: Incident: forcing en in the 2026-07-20 mock session caused mixed-language output (#262).
413 must 3. **Measure "answered → card"**: the timestamp gap from answers_received to the pre… 411 kept (reworded). Removed: E43: Issue pointer #236 ("re-measurement instrument").
418 do not, must, only - artifact_generatedmust come **before** thecard_presented of the same stage, an… 416 unchanged
419 only - start --question-mode/--card-mode declare only the capabilities this client has … 417 kept (reworded). Removed: E45: PR #298 pointer ("merged").
420 only - **--adapterdefaults toplain_text**, and the plain_text adapter may declare *… 418 unchanged
421 do not, never - **The enum cannot express "widget cards but plain-text questions"** ([#337](https://… 419 kept (reworded). Removed: E46a: Pointer to open issue #337; E46b: Pointer to #230 as the purpose the rule serves.
422 must - findings_recordedmust come **before**owner_verdict: the verdict is the session… 420 unchanged
423 must, only - response_mode/response_provenance apply only to question kinds that support a pr… 421 unchanged
425 do not The QA mindset — watch for these while walking (record what you find; do not fix it he… 423 unchanged
430 must, only - **When presenting the candidate rule choice, did the agent quietly reword or invent … 428 kept (reworded). Removed: E49a: Dated lesson pointer (2026-07-21, #293); E49b: Issue pointer #293.
434 only This step ends **one route run**, not the conversation. Archive is not a stop signal: … 432 unchanged
436 always, do not 1. **Owner verdict + archive the receipt (the core output of QA; do not skip it)**: on… 434 unchanged
438 refuses **The order is hard**: findings_recorded(gate 7) first,owner_verdict second. The… 436 unchanged
460 never # Never guess, and never fill in unknown/default. 458 unchanged
466 never The --case-id values are defined by the acceptance issue, never invented here: #486'… 464 kept (reworded). Removed: E50: "since its 2026-07-29 ruling": the decision date is dropped; the id lists stay with 2026-07-29 as their as-of date (the issue's checklist has since grown to M0-T11).
468 never Archiving produces a **run manifest** (<run_id>.manifest.json) recording this dogfoo… 466 unchanged
470 only **Case and state lineage (#520)**: a receipt that verifies proves only that this sessi… 468 kept (reworded). Removed: E51a: Issue pointer #520; E51b: "now" (when archive began enforcing).
475 do not | --case-id| e.g.M0-U01, M0-U02 | Stable identifiers defined in #486; do no… 473 unchanged
477 only, refused | --parent-run-id| the previous session'srun_id| **Only** forcontinued; s… 475 unchanged
479 must, never, only ``--parent-run-id must name a manifest that **actually exists** in the receipt directo… 477 kept (reworded). Removed: E53: Deferred-work pointer to #492 ("if operational evidence ever justifies it").
482 do not, must 3. **What you found**: write each one down. If it is genuinely a bug or a gap, check …` 480 unchanged
484 must **But opening an issue is not the end — Step 6 must be finished before owner_verdict… 482 unchanged
485 do not, never, only 4. **Cleanup — only when the maintainer says the campaign is over.** Do not offer it a… 483 unchanged
488 only - Explicitly done with the worktree → ~/.claude/skills/fomo-qa/qa_env.sh down (remov… 486 unchanged
492 only **archive-receiptenforces this step**: a receipt with nofindings_recorded row ca… 490 kept (reworded). Removed: E54: Issue pointer #417.
501 do not python3 evals/run_episodes.py EP-NNN # read which checks it actually trips; do … 499 unchanged
505 must, only Then record on the receipt where each miss went — the runbook's seventh gate, enforced… 503 unchanged
508 never - **episode:EP-NNNis reconciled againstevals/episodes/** — claiming a conversion… 506 unchanged
509 must - **"This session found nothing" must be said explicitly** (--no-findings); leaving … 507 unchanged
512 only **Sessions using real data**: an episode keeps only the failure's *structure*. De-iden… 510 unchanged
516 do not After a route run is archived, **stay here**. The maintainer is now a user with a book… 514 kept (reworded). Removed: E55: #544 history: walking four routes used to cost four full ceremonies in four sessions.
539 never The response's challenge block is printed on that call's own stdout and **never stor… 537 unchanged
548 only # 1) Only when a bounded context question was actually asked (for example, 546 unchanged
562 only # read the check file before recording. The trace is append-only. 560 unchanged
567 only # invitation was shown and nothing settled; acted/declined/modified only 565 unchanged
568 never # after the user's word was recorded via consider --resolve — never a 566 unchanged
585 never ``--challenge-check-file points at a transient JSON that never enters the trace, the s… 583 kept (reworded). Removed: E58: "since #767" (when sector_display joined the check file).
595 never "detail": {"rule_id": "rule-1", "text": "One name never above a quarter of the book", … 593 unchanged
607 never {"rule_id": "rule-1", "text": "One name never above a quarter of the book", 605 unchanged
617 never "presented_text": "Not at this size. It takes SYNTH to 27.3%, which crosses a line you… 615 unchanged
623 never, refused, refuses Paste the challenge verbatim from stdout — a truncated paste is refused rather than re… 621 kept (reworded). Removed: E59a: Issue pointer #830; E59b: Issue pointer #739.
625 never, only What the machine half checks, and what stays with the owner's comprehension verdict:… 623 kept (reworded). Removed: E60: Issue pointer #767.
627 only Archive like any other continued route run in this campaign: --state-mode continued -…` 625 kept (reworded). Removed: E61: The retired "exploratory only" caveat and the "now"/"still" framing.
629 do not, never, only The resolution boundary: actedrecords the user's own word viaconsider --resolve,… 627 unchanged
633 never, only The campaign ends only when the maintainer explicitly says so ("that's enough", "we're… 631 unchanged
642 only | agent_with_owner_verdict | The AI walked the flow; you gave only the final verdi… 640 unchanged
643 only | agent_simulated (**default**) | Fully AI-simulated, no human | ❌ Contract only,… 641 unchanged
653 do not, only The report keeps owner_liveandagent_simulated **strictly separate**, and buckets… 651 unchanged

Review

A whole-diff "did any rule disappear?" pass is not sensitive enough, so the review is per item and delta-first. A token-level diff of the old and new file gave 114 items. Each item is one piece of old text that is no longer there, cut at sentence boundaries, with its replacement and a little context. The reviewer returns one verdict per item (story, rule_kept, rule_changed, rule_lost, borderline) with a verbatim quote of the live fragment and of the new text that carries it, and gets the complete new file to look up where a rule went. Two reviewers from two non-Claude families, cold start, review-only, material inline, empty working directory.

Calibration: a copy with one live rule deleted

The planted deletion is the sentence "Confirm the tool's contract (does it preserve the original CSS?) before concluding that the host cannot render rich HTML." It closes a paragraph whose story ending is removed too, which is the easiest place for a deleted rule to hide.

pack design Codex gpt-6-luna Antigravity
35 blocks (nearby hunks merged) caught (block flagged rule_changed) gemini-3.7-flash-high missed
61 line-level items missed gemini-3.7-flash-high missed
117 delta items, deleted spans cut at sentences caught (rule_changed on the sentence, rule_lost on its tail) 3.7 Flash missed; gemini-3.1-pro-high missed; gemini-3.8-flash-high caught (rule_lost twice)

So the real review uses the delta design with the two reviewers that caught the planted deletion on it: Codex gpt-6-luna and gemini-3.8-flash-high. The contract's default gemini-3.7-flash-high did not catch it on any design, and neither did Pro; a coarser pack let even Codex miss it once.

Receipts

run family, model command shape exit time result
real Codex codex-cli 0.156.1, gpt-6-luna (the config default gpt-6.1-sol is rejected by the account) codex exec -m gpt-6-luna --sandbox read-only --json --ephemeral --skip-git-repo-check -C <empty dir> -, prompt on stdin, 109,731 bytes 0 181 s marker echoed, 114/114 answered: story 76, rule_kept 30, rule_changed 7, rule_lost 1
real Antigravity agy 1.2.16, gemini-3.8-flash-high agy -p <prompt as argv> --model gemini-3.8-flash-high --effort high, empty cwd 0 134 s marker echoed, 114/114 answered: story 111, rule_kept 3, nothing flagged
calibration Codex gpt-6-luna same 0 165 s caught
calibration gemini-3.8-flash-high same 0 132 s caught

Findings and dispositions

Every quote was checked against the files; only Codex flagged anything.

item reviewer verdict disposition
22: eval "currently marks 'pending owner dogfood'" Codex rule_changed stays deleted: it was a stale fact, no such marker exists in docs/eval-design.md or evals/EVALS.md
32: the test_card_html screenshot remark Codex rule_lost stays deleted: it was supporting evidence ("today ..."); the limit it supports is still stated in the same sentence and nothing is misread without it
38: "judged" to "finds" (guardrail 2) Codex rule_changed stays: an event restated as the property it showed; the rule (TRADE_COACH_HOME routes writers only, isolate closes the read path) is unchanged
65, 72, 73: one observed run to present tense Codex rule_changed stays: each statement is guaranteed by code or a test (QUESTION_POLICY["snapshot_review"] is min 0 / max 0 in engine/review.py, and plan["question_queue"] == [] is asserted in tests/test_review_v2.py lines 461 and 2069; a declaration that agrees with the recorded book prepares as reconciled in tests/test_review_v2.py line 1756, and a re-run after adoption reports reconciled in tests/test_book_refresh.py line 1230)
100: "three things" to "these arguments" Codex rule_changed fixed: only "now" is removed, the count stays as written
109: "not yet read" to "not read" Codex rule_changed fixed: edit dropped, "not yet" restored
(own check of the observed-to-present rewrites, not flagged by a reviewer) fixed: "a grounding sentence on each candidate" would over-claim, because engine/review.py attaches grounding only when it is non-empty; the text now says "each with a grounding sentence when the engine has one"

The calibration passes ran on the real, unmutated content too and flagged two more things, both acted on: the exact M0-F/M0-T id ranges had been dropped (they stay, dated as of 2026-07-29) and the --case-id example ids in the lineage table had been changed (reverted, example values stay).

Backstop: every deleted span with a rule-ish word (must, never, only, confirm, record ...) that both reviewers called story was read by hand. There are four, and all are plain stories: "had always lacked", the #230 pointer, the mis-diagnosis incident, "the old exploratory-only caveat is gone".

After the review the three adjustments above were applied. The final text differs from the reviewed text in exactly those three lines; each restores old wording or states an engine fact.

Verification

python3 tests/run_all.py (group all, Python 3.14.7) on the head commit: PASS: all 60 suites passed, exit 0, 5 min 24 s. The pinned suites (qa/tests/test_skill_commands.py, test_receipts.py, test_isolation_gate.py, tests/test_doc_language.py, tests/test_repo_hygiene.py) were also run alone while editing. CI on the head: product-contract and qa-eval-tooling pass on Python 3.11 and 3.12 (qa/ is QA/eval-owned, so qa-eval-tooling blocks).

Acceptance greps (grep exit status: 0 match, 1 none, 2 or more error): the date grep exits 0 with the four lines above, the history-word grep exits 0 with the two legacy-unattributed lines; neither exits 2 or more.

🤖 Generated with Claude Code

qa/SKILL.md is the fomo-qa skill that Claude, Codex, Antigravity and Grok load
as instructions. Next to its rules it carried dated decisions, incident and
audit write-ups, issue and PR pointers used as provenance, "used to" and "now"
notes, a retired-flag remark and superseded examples. Each is rewritten in
place so the rule stands alone; nothing is added, and no rule is dropped or
weakened.

- frontmatter description: 956 -> 450 characters; what the skill does, its
  triggers, and the pointer to the product skill fomo-kernel. The rationale
  half (purpose, mechanism list) is gone; every rule in it is stated in the
  body.
- 68 in-place edits across the body. Past-tense observations of one run become
  present-tense statements of the route's behaviour only where the engine or a
  test guarantees them (QUESTION_POLICY, test_review_v2's reconciled status);
  a candidate's `grounding` is stated as present when the engine has one.
- two dates stay because deleting them would change what an agent does: the
  as-of stamp on how generic visualization tools treat a supplied <style>
  (2026-07-21) and on the case-id lists of acceptance issue #486 (2026-07-29).
  Example values in templates are untouched.
- the issue and PR numbers that only said where a rule came from are gone;
  #486, the live acceptance issue an operator must read, stays.

Locks kept: the gate count and the seven-name roster, CJK only on the
description line and the trigger bullet, every `# qa-trace:` fence and its
replay, the placeholder map, the "Continuing in the same campaign" section.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@atomchung
atomchung merged commit aa1c62c into main Oct 5, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant