Skip to content

feat(evals): add the #715 generic-parity fast probe (recovered from an unmerged local commit) - #848

Draft
atomchung wants to merge 1 commit into
mainfrom
claude/issue-715-generic-parity
Draft

atomchung wants to merge 1 commit into
mainfrom
claude/issue-715-generic-parity

Conversation

@atomchung

Copy link
Copy Markdown
Owner

Why this is a draft

#715 is marked "Research / Discussion only — product decision recorded, not an active implementation front", and says not to merge ahead of the reliability fronts in #27. This PR adds no product behaviour — it is opt-in evaluation tooling — but the merge call is the owner's, so it opens as a draft rather than presuming the gate.

What this recovers

The probe was written on 2026-08-01 on codex/issue-715-generic-parity, the branch that had already carried PR #727, and never reached a PR of its own. It has been sitting only in a local commit ever since:

File Lines
evals/generic_parity.py 196
evals/run_generic_parity.py 214
tests/test_generic_parity.py 128
evals/judge_episodes.py (two shared text-call helpers) +75
evals/generic-parity-cases.json, evals/generic-parity-witnesses.json 35
docs/eval-design.md, evals/EVALS.md, tests/run_all.py +28

run_generic_parity.py is an opt-in, host-side A01/A07/A10 probe for the no-book route. It is not an engine route, not a product-surface capture, and not a Generic Parity acceptance gate: its receipt always says owner_unreviewed / not_granted, and candidate_pass is iteration evidence, never merge authority. --plan makes zero model calls.

What was dropped in the rebase

The original commit also touched skills/fomo-kernel/references/decision-framing.md and research-priors.md. Both are dropped — main's versions are the later rewrite that #727 landed, and they already carry the strategy-class map the older text proposed (main renamed the section to ## Research-aware strategy framing and added the #716 cross-route note). Keeping the commit's halves would have reverted that work.

tests/run_all.py registers the new suite in main's three-tuple lane format (qa-eval), not the two-tuple format the original commit predated.

Verification

python3 tests/run_all.py → PASS: all 61 suites passed (group: all), including the newly registered tests/test_generic_parity.py (6/6).

Refs #715. Does not close it — the design question the issue tracks is unchanged by this PR.

🤖 Generated with Claude Code

The probe was written on 2026-08-01 on a branch that had already carried
PR #727, and never reached a PR of its own. `evals/generic_parity.py`,
`evals/run_generic_parity.py`, `tests/test_generic_parity.py` and the two
case/witness fixtures existed only in one local commit while #715 stayed
open.

Rebased onto main. The original commit's `decision-framing.md` and
`research-priors.md` halves are dropped: main's versions are the later
rewrite #727 landed, and they already carry the strategy-class map the
older text proposed. `tests/run_all.py` registers the new suite in main's
three-tuple lane format.

61/61 offline suites pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant