Repository navigation
docs(qa): refresh qa/SKILL.md, keep every rule and drop the history - #855
Merged
Merged
Conversation
qa/SKILL.md is the fomo-qa skill that Claude, Codex, Antigravity and Grok load as instructions. Next to its rules it carried dated decisions, incident and audit write-ups, issue and PR pointers used as provenance, "used to" and "now" notes, a retired-flag remark and superseded examples. Each is rewritten in place so the rule stands alone; nothing is added, and no rule is dropped or weakened. - frontmatter description: 956 -> 450 characters; what the skill does, its triggers, and the pointer to the product skill fomo-kernel. The rationale half (purpose, mechanism list) is gone; every rule in it is stated in the body. - 68 in-place edits across the body. Past-tense observations of one run become present-tense statements of the route's behaviour only where the engine or a test guarantees them (QUESTION_POLICY, test_review_v2's reconciled status); a candidate's `grounding` is stated as present when the engine has one. - two dates stay because deleting them would change what an agent does: the as-of stamp on how generic visualization tools treat a supplied <style> (2026-07-21) and on the case-id lists of acceptance issue #486 (2026-07-29). Example values in templates are untouched. - the issue and PR numbers that only said where a rule came from are gone; #486, the live acceptance issue an operator must read, stays. Locks kept: the gate count and the seven-name roster, CJK only on the description line and the trigger bullet, every `# qa-trace:` fence and its replay, the placeholder map, the "Continuing in the same campaign" section. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
qa/SKILL.mdis the fomo-qa skill that Claude, Codex, Antigravity and Grok load as instructions (all four discovery links point at this directory). Next to its rules it carried dated decisions, incident and audit write-ups, issue and PR pointers used as provenance, "used to" and "now" notes, a remark about a removed flag, and superseded examples. This rewrites each in place so the rule stands alone. No rule is added, dropped or weakened, and nothing outsideqa/SKILL.mdchanges. The stories move into this description.Numbers
wc -l)descriptiongrep -nE '20[0-9]{2}[-/][0-9]{2}'grep -nEi 'historical|legacy|previously|used to|deprecated|歷史|舊版|以前|曾經|拍板'legacy-unattributed, below)#486)The description now says what the skill does, its triggers (the Chinese trigger phrases stay on that one line) and the out-of-scope boundary with the pointer to the product skill
fomo-kernel.The workspace skill lint reports 0 errors and 5 warnings (before: 13 errors and 5 warnings). The warnings are the length, the two as-of dates and the two
legacy-unattributedlines, all explained below.Length. 659 lines against the checklist's 200. 299 of them are fenced commands, and 209 of those sit in seven
# qa-trace:fences thatqa/tests/test_skill_commands.pyreplays throughux_receipt.py verify. The order of those commands is a contract (a wrong order voids the run and the trace is append-only), and the prose between them is the per-route rule behind that order, so none of it can go without changing what an operator does. Splitting the file into references is a separate decision, not part of a history refresh.Repository rules checked first
docs/expression-contract.md§3.4 ("Where a local answer-shape phrasing was superseded by this chapter, the surface document says so rather than deleting its own history") does not cover this file. It governs answer-shape phrasing on the contract's own surfaces (§6: review card,consider, freeform, no-book, weekly read, plusoutput-voice.md's V1), and asks only that a surface that once had a local answer-first phrasing say it was superseded.qa/SKILL.mdis a maintainer procedure with no answer-shape phrasing and is not a surface. I found no other repository rule that asks a skill, flow or procedure to keep its history; the history-bearing documents are ruling logs (docs/expression-contract.md§8,docs/output-contract.md), the maintainer guide's receipt rows, and V1 indocs/output-voice.md, none of which is a procedure an agent follows.qa/tests/test_skill_commands.py,docs/language-policy.md(qa/is English; CJK is allowed only on thedescription:line and the- The user saysbullet),docs/development-guide.md§6 (find the enforcing test before cleaning; read and classify "never" lines instead of counting them) and the maintainer-guide row for the isolation gate, which names this file's Step 0 and guardrail 2 as readers. All of those stay intact: the gate count and its seven-name roster, every# qa-trace:fence and its replay, the placeholder map, the "Continuing in the same campaign" section, guardrails 1 to 5.Kept on purpose
As of 2026-07-21<style>); it tells an agent when to re-verifyas of 2026-07-29M0-T11), so the lists read as a dated example and the next sentence sends the operator to the issue2026-08-14--challenge-check-filetemplate; a test replays this JSONlegacy-unattributedreportprints (qa/receipts.py), a live rule#486Left untouched because they are not history: the
M0-U01,M0-U02example ids in the lineage table (they disagree with theM0-Fids two paragraphs above; example values stay) and the count "three things" before a four-row table.What was removed
Source commit = the oldest commit whose diff adds the text (
git log -S, restricted toqa/SKILL.md).58d46f5is the commit that translated the Chinese import0957af0into English, so the oldest stories (the 2026-07-19 to 2026-07-27 dates, the #230/#236/#238/#293 notes) were already in that import.58d46f558d46f558d46f558d46f558d46f558d46f558d46f558d46f558d46f558d46f558d46f558d46f558d46f52b0c07f58d46f52b0c07f2b0c07f2a0e59758d46f558d46f52a0e59758d46f558d46f558d46f558d46f558d46f5c217228c217228c217228c21722858d46f5889c8fd889c8fd889c8fdgroundingonly when the engine has one889c8fd889c8fd889c8fd889c8fd889c8fd889c8fd889c8fd58d46f558d46f558d46f5enin the 2026-07-20 mock session caused mixed-language output (#262)58d46f558d46f558d46f558d46f558d46f558d46f558d46f558d46f558d46f558d46f52a0e59758d46f558d46f558d46f57c39ae62a0e5972a0e59783f5cd6757bb3a98a3c007f460e1757bb3a83f5cd658d46f558d46f52026-07-28: fix(qa): /fomo-qa commands run as written, and archived runs name their case (closes [feat·QA·M0] Bind owner-live receipts to named acceptance cases and state lineage #520) (closes [process·qa] SKILL.md and qa-runbook.md still hand-mirror their prose halves — #520 gated only the commands #527) (fix(qa): /fomo-qa commands run as written, and archived runs name their case (closes #520) #524)2b0c07f2026-07-29: fix(qa): a dogfood run refuses to start while the account's own coach root is still reachable (closes [bug·privacy·qa] TRADE_COACH_HOME routes the engine, not the agent — a probe reached the real coach root from an isolated run #557) (fix(qa): a dogfood run refuses to start while the account's own coach root is still reachable (closes #557) #568)2a0e5972026-07-29: docs(qa): a campaign outlives the run it archived (refs [process·qa] /fomo-qa needs one continuous host campaign across review, refresh, and consider #544) (docs(qa): a campaign outlives the run it archived (refs #544) #545)889c8fd2026-07-29: docs(qa): the snapshot-then-refresh journey has walkthroughs, observed rather than composed (closes [process·qa] /fomo-qa has no verified walkthrough for snapshot_review, the route M0-U02 starts on #526) (docs(qa): the snapshot-then-refresh journey has walkthroughs, observed rather than composed (closes #526) #537)83f5cd62026-07-30: feat(qa): consider becomes a receipted card-free route — challenge delivery proven, not assumed (refs [process·qa] /fomo-qa needs one continuous host campaign across review, refresh, and consider #544) (refs [feat·M1] Bounded TradeEvaluation with visible two-sided challenge #479) (feat(qa): consider becomes a receipted card-free route — challenge delivery proven, not assumed (refs #544) (refs #479) #586)c2172282026-08-01: fix(qa): the preview receipt is the settled pre-commitment card, not an intermediate render (closes [process·qa] Cash-anchor recompute mid-first_review has no documented receipt slot for its second preview card #663) (fix(qa): the preview receipt is the settled pre-commitment card, not an intermediate render (closes #663) #696)757bb3a2026-08-02: feat(qa): check a consider answer's sector name against its own sector_display (closes [design·gate] Move #746's guarantee off the documentation layer — refuse an answer naming the sector against the response's own sector_display #767) (feat(qa): check a consider answer's sector name against its own sector_display (closes #767) #770)7f460e12026-08-02: fix(consider): default premise.price from the observed close, localize disclosure keys, and freeze the position cap across add-cash (closes [dx·consider·P3] consider prices the premise instrument itself but will not surface that observation, while requiring a price the caller must source elsewhere #777, closes [copy·i18n·UX] Engine disclosure keys and error payloads leak raw English enums into localized zh-TW / zh-CN consider turns #739, closes [bug·docs·P2] set-cap before add-cash inside one pending session trips the drift gate — two same-beat documented actions are order-sensitive with no warning #758) (fix(consider): default price from observed close, localize disclosure keys, freeze cap across add-cash #795)98a3c002026-08-20: feat(challenge): delete the obligation whitelist, keep the seven facts that earn their place (closes [design·M1] Owner-live: answers exceed the reading budget — verbosity is the next usefulness bottleneck after #827 #830) (feat(challenge): delete the obligation whitelist, keep the seven facts that earn their place (closes #830) #831)7c39ae62026-10-05: docs(instructions): drop history narratives from the instruction files (docs(instructions): drop history narratives from the instruction files #851)Two of the removals corrected a stale claim rather than only dropping history: eval's layer "currently marks 'pending owner dogfood'" (no such marker exists in
docs/eval-design.mdorevals/EVALS.md) and the test_card_html remark about zero screenshots.Rule preservation
Every old line that carries
must / never / always / do not / only / refus*(or the Chinese equivalents) is mapped to its new line. Old lines with a rule keyword: 116. unchanged: 77; kept (reworded): 39; story only: 0; lost: 0.The full map (117 rows)
description: Prepares a clean, consistent QA environment for dogfooding fomo-kernel an…~/.trade-coach-> guardrail 2, new L37; "Never touches the real records in investment_note" -> guardrail 1, new L36.**This is the mandatory standard path for every fomo-kernel dogfood (v1, fixed 2026-07…**Cross-client contract source (since 2026-07-21)**:docs/qa-runbook.mdin thekol_…`> **Its place in the eval system**: this is also the execution procedure thatdocs/ev…`**Not** for reviewing a real user's trades — that is the product skillfomo-kernel. …## Coverage (v1, fixed — do not claim beyond it)This round verifies **L1: environment consistency + engine CLI contract + agent walkth…- **L2 card visuals** (next round): the card HTML is never actually rendered in a brow…- **L3 interaction delivery** (partly an inherent ceiling): "the option buttons really…1. **Never touch the real records**:~/Side_project/investment_note/holds ting's re…2. **Coach state is isolated to a dogfood-only root, and the real one is not reachable…3. **Work only in the dogfood worktree**: every engine command runs inside the detache…4. **Do not change product code**: QA reads, it does not edit. If the walkthrough find…5. **Public text passes the privacy lint first (bought by the #274 incident)**: the re…Only exit 0 may be posted. On a hit, rewrite as a de-identified description ("N indivi…**Cross-client execution gaps** (being in the discovery registry only guarantees the s…- If Step 4 presents questions through Claude's native option tool (for exampleAskUs…`- The rendering-pipeline test mentioned in walkthrough rule 1's "try the widget once" …-qa_env.sh's assumptions about the current working directory and worktree have not …### Step 0 — Isolate this shell, then the version gate (read-only)**refuses every command** until the shell it runs in has the account's own…It is a bounded guarantee, and reporting it as more than that is the failure it was wr…At a glance: the latestorigin/mainsha, how far behind the dogfood worktree is, and…also reports one extra line, **this skill's own freshness** — checking the ch…Creates (or refreshes)~/Side_project/kol_collector/fomo-kernel-dogfoodat--detach…`| **Receipt** | one route-specific, append-only evidence trace | exactly one per ro…**Never merge route runs into one receipt.** Afirst_reviewowes two cards and a cas…**Step 0'sisolatealready routed the whole toolchain into the dogfood-only coach ro…- **Simulate a returning user** (runs weekly-review / due-revisit): do **not** reset. …**This isolation must survive into every later shell** — if commands each start a new …| **Real trades** (read-only) |~/Side_project/investment_note/trades/fomo/trades.c…`| **Test-drive** |--test-drive(no CSV) | Demonstration only;persist:false, z…### Step 4 — Walk through (follow the product's fixed lifecycle; do not rewrite it)Aftercd .../fomo-kernel-dogfood/skills/fomo-kernel(and confirming Step 2'sexport…`# "no numbers" narrative (answers.json / narrative.json must pass their schemas)**The UX receipt runs through the whole walkthrough (mandatory — this is what connects…Below is a **complete, directly copyablefirst_reviewtrace**. The order is a contra…# 0) Declare host capability right after prepare. --adapter must match the capability# universal plain_text and markdown_inline fallbacks itself — do not pass them again.# it must come before the first question and the first card — recording it later is# ask": a run where the user was never offered it records nothing, andverifyfai…# 2) One row per question asked. Question text never enters the trace: a question from a# the two must appear together (this replaces the removed --question-id).# 4) Cards are always "artifact first, presented second", and both rows need --stagepoints at a **transient JSON that never enters the trace** (s…A candidate with nogroundingomits the field entirely (likecandidate_1) — **do n…**Theweekly_reviewroute carries one extra opener, andverifyenforces it** (the …-provided— always triggersadd-cashand a recompute, so the card the user actual…# --memory-kind flag; these do not count as the opener):# The FIRST preview (holdings-only) renders and is shown here, in one message with the# artifact_generated/card_presented row of its own; only the cash answer is recorded n…The wrap-up has the same shape as Step 5, except **--memorymust bepassorfail…Selected when the user has a position table or screenshot and no transaction history. …- **No cash anchor row.** The route's contract does not carry the #357 pre-flight, bec…- **No question rows.** The observed plan came back withquestion_queue: []andcar…`Once a book exists, a newer holdings view is **not** a review.prepare --route snapsh…`this holdings view has changes only you can settle before the recorded book canSo the real journey is composed — record, then review — and it produces **two receipts…# step 1: read-only. Writes nothing; returns the frozen diff, a summary, and- **A refresh creates no session**, so the trace is keyed by the engine's ownrefresh…`- **No card events at all.**verifyrefusesartifact_generated,card_presented, …- **The question row depends on what the engine raised, and so does the verdict.** Ste…- **That--controlschoice is not recoverable.** Recording--controls passon a re…# 1) The engine's difference, as you narrated it. change_presented carries only# 2) Only when step 1 came backpending_confirmation: the one question covering1. **Declare capability honestly, and try the widget once per session — with the right…**New lesson, 2026-07-21 (see the #230 comments)**: trying the widget does not mean gr…**This tool-selection detail belongs here and must not be promoted into fomo-kernel's …2. **--languagefollows the conversation language**: a Chinese conversation always u…enin the 2026-07-20 mock session caused mixed-language output (#262).3. **Measure "answered → card"**: the timestamp gap fromanswers_receivedto the pre…-artifact_generatedmust come **before** thecard_presentedof the same stage, an…-start --question-mode/--card-modedeclare only the capabilities this client has …- **--adapterdefaults toplain_text**, and theplain_textadapter may declare *…- **The enum cannot express "widget cards but plain-text questions"** ([#337](https://…-findings_recordedmust come **before**owner_verdict: the verdict is the session…-response_mode/response_provenanceapply only to question kinds that support a pr…The QA mindset — watch for these while walking (record what you find; do not fix it he…- **When presenting the candidate rule choice, did the agent quietly reword or invent …This step ends **one route run**, not the conversation. Archive is not a stop signal: …1. **Owner verdict + archive the receipt (the core output of QA; do not skip it)**: on…**The order is hard**:findings_recorded(gate 7) first,owner_verdictsecond. The…# Never guess, and never fill in unknown/default.The--case-idvalues are defined by the acceptance issue, never invented here: #486'…Archiving produces a **run manifest** (<run_id>.manifest.json) recording this dogfoo…**Case and state lineage (#520)**: a receipt that verifies proves only that this sessi…|--case-id| e.g.M0-U01,M0-U02| Stable identifiers defined in #486; do no…|--parent-run-id| the previous session'srun_id| **Only** forcontinued; s…must name a manifest that **actually exists** in the receipt directo…3. **What you found**: write each one down. If it is genuinely a bug or a gap, check…`**But opening an issue is not the end — Step 6 must be finished beforeowner_verdict…4. **Cleanup — only when the maintainer says the campaign is over.** Do not offer it a…- Explicitly done with the worktree →~/.claude/skills/fomo-qa/qa_env.sh down(remov…**archive-receiptenforces this step**: a receipt with nofindings_recordedrow ca…python3 evals/run_episodes.py EP-NNN # read which checks it actually trips; do …Then record on the receipt where each miss went — the runbook's seventh gate, enforced…- **episode:EP-NNNis reconciled againstevals/episodes/** — claiming a conversion…- **"This session found nothing" must be said explicitly** (--no-findings); leaving …**Sessions using real data**: an episode keeps only the failure's *structure*. De-iden…After a route run is archived, **stay here**. The maintainer is now a user with a book…The response'schallengeblock is printed on that call's own stdout and **never stor…# 1) Only when a bounded context question was actually asked (for example,# read the check file before recording. The trace is append-only.# invitation was shown and nothing settled; acted/declined/modified only# after the user's word was recorded viaconsider --resolve— never apoints at a transient JSON that never enters the trace, the s…"detail": {"rule_id": "rule-1", "text": "One name never above a quarter of the book", …{"rule_id": "rule-1", "text": "One name never above a quarter of the book","presented_text": "Not at this size. It takes SYNTH to 27.3%, which crosses a line you…Paste the challenge verbatim from stdout — a truncated paste is refused rather than re…What the machine half checks, and what stays with the owner'scomprehensionverdict:…Archive like any other continued route run in this campaign:--state-mode continued -…`The resolution boundary:actedrecords the user's own word viaconsider --resolve,…The campaign ends only when the maintainer explicitly says so ("that's enough", "we're…|agent_with_owner_verdict| The AI walked the flow; you gave only the final verdi…|agent_simulated(**default**) | Fully AI-simulated, no human | ❌ Contract only,…The report keepsowner_liveandagent_simulated**strictly separate**, and buckets…Review
A whole-diff "did any rule disappear?" pass is not sensitive enough, so the review is per item and delta-first. A token-level diff of the old and new file gave 114 items. Each item is one piece of old text that is no longer there, cut at sentence boundaries, with its replacement and a little context. The reviewer returns one verdict per item (
story,rule_kept,rule_changed,rule_lost,borderline) with a verbatim quote of the live fragment and of the new text that carries it, and gets the complete new file to look up where a rule went. Two reviewers from two non-Claude families, cold start, review-only, material inline, empty working directory.Calibration: a copy with one live rule deleted
The planted deletion is the sentence "Confirm the tool's contract (does it preserve the original CSS?) before concluding that the host cannot render rich HTML." It closes a paragraph whose story ending is removed too, which is the easiest place for a deleted rule to hide.
gpt-6-lunarule_changed)gemini-3.7-flash-highmissedgemini-3.7-flash-highmissedrule_changedon the sentence,rule_loston its tail)gemini-3.1-pro-highmissed;gemini-3.8-flash-highcaught (rule_losttwice)So the real review uses the delta design with the two reviewers that caught the planted deletion on it: Codex
gpt-6-lunaandgemini-3.8-flash-high. The contract's defaultgemini-3.7-flash-highdid not catch it on any design, and neither did Pro; a coarser pack let even Codex miss it once.Receipts
codex-cli 0.156.1,gpt-6-luna(the config defaultgpt-6.1-solis rejected by the account)codex exec -m gpt-6-luna --sandbox read-only --json --ephemeral --skip-git-repo-check -C <empty dir> -, prompt on stdin, 109,731 bytesagy 1.2.16,gemini-3.8-flash-highagy -p <prompt as argv> --model gemini-3.8-flash-high --effort high, empty cwdgpt-6-lunagemini-3.8-flash-highFindings and dispositions
Every quote was checked against the files; only Codex flagged anything.
rule_changeddocs/eval-design.mdorevals/EVALS.mdrule_lostrule_changedTRADE_COACH_HOMEroutes writers only,isolatecloses the read path) is unchangedrule_changedQUESTION_POLICY["snapshot_review"]is min 0 / max 0 inengine/review.py, andplan["question_queue"] == []is asserted intests/test_review_v2.pylines 461 and 2069; a declaration that agrees with the recorded book prepares asreconciledintests/test_review_v2.pyline 1756, and a re-run after adoption reportsreconciledintests/test_book_refresh.pyline 1230)rule_changedrule_changedgroundingsentence on each candidate" would over-claim, becauseengine/review.pyattachesgroundingonly when it is non-empty; the text now says "each with agroundingsentence when the engine has one"The calibration passes ran on the real, unmutated content too and flagged two more things, both acted on: the exact
M0-F/M0-Tid ranges had been dropped (they stay, dated as of 2026-07-29) and the--case-idexample ids in the lineage table had been changed (reverted, example values stay).Backstop: every deleted span with a rule-ish word (
must,never,only,confirm,record...) that both reviewers calledstorywas read by hand. There are four, and all are plain stories: "had always lacked", the #230 pointer, the mis-diagnosis incident, "the old exploratory-only caveat is gone".After the review the three adjustments above were applied. The final text differs from the reviewed text in exactly those three lines; each restores old wording or states an engine fact.
Verification
python3 tests/run_all.py(group all, Python 3.14.7) on the head commit: PASS: all 60 suites passed, exit 0, 5 min 24 s. The pinned suites (qa/tests/test_skill_commands.py,test_receipts.py,test_isolation_gate.py,tests/test_doc_language.py,tests/test_repo_hygiene.py) were also run alone while editing. CI on the head:product-contractandqa-eval-toolingpass on Python 3.11 and 3.12 (qa/is QA/eval-owned, soqa-eval-toolingblocks).Acceptance greps (
grepexit status: 0 match, 1 none, 2 or more error): the date grep exits 0 with the four lines above, the history-word grep exits 0 with the twolegacy-unattributedlines; neither exits 2 or more.🤖 Generated with Claude Code