test: ties, NULLs, extreme values and big tables in the v2 soak (M4 7) - #91
Merged
Merged
Conversation
This was referenced Sep 29, 2026
alxmrs
added a commit
that referenced
this pull request
Sep 29, 2026
From the v2 soak (#89, #91): - MAX and MIN sent the cotangent to the recomputed rows equal to the saved extreme. A recomputation need not match the forward pass to the bit (DataFusion's grouped SUM over several partitions adds in arrival order), so about half the time no row attained the saved maximum: a silent zero gradient, or NaN/inf from dividing by zero attaining rows. The extreme, the count of rows attaining it, and AVG's count are now windows over the recomputed rows themselves, partitioned by the group's keys, so the comparison is always with the values being compared. The saved aggregate is no longer read back, and neither is a second copy of the region. - AVG, MAX and MIN also get the NULL-row mask SUM got in #74 (applied in the rebase of the rules commit): no gradient through a row the aggregate skips. - Tests for both, and for a rank filter over a rank filter (fixed in #73). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
September 29, 2026 17:26
cc012e5 to
0406589
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
from
September 29, 2026 17:26
e224dcc to
6a846d4
Compare
Member
Author
|
The MAX recomputation bug was real and is fixed in #76. The rows attaining an extreme, and how many there are, are now windows over the recomputed rows themselves, never a comparison with the saved value. A later ties-mode soak (seed 2101304) showed the same jitter can also flip which of two exactly-tied groups wins. So a row now attains a finite extreme if it's equal to it or within 8 ulps of it; a tie that rounding breaks is shared, the same way every run. Other changes on this branch, all from soaks of the fixed stack:
🤖 Generated with Claude Code |
alxmrs
added a commit
that referenced
this pull request
Sep 30, 2026
From the v2 soak (#89, #91): - MAX and MIN sent the cotangent to the recomputed rows equal to the saved extreme. A recomputation need not match the forward pass to the bit (DataFusion's grouped SUM over several partitions adds in arrival order), so about half the time no row attained the saved maximum: a silent zero gradient, or NaN/inf from dividing by zero attaining rows. The extreme, the count of rows attaining it, and AVG's count are now windows over the recomputed rows themselves, partitioned by the group's keys, so the comparison is always with the values being compared. The saved aggregate is no longer read back, and neither is a second copy of the region. - AVG, MAX and MIN also get the NULL-row mask SUM got in #74 (applied in the rebase of the rules commit): no gradient through a row the aggregate skips. - Tests for both, and for a rank filter over a rank filter (fixed in #73). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
September 30, 2026 02:47
0406589 to
64a3df4
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
from
September 30, 2026 02:47
6a846d4 to
74c66e6
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
September 30, 2026 09:26
64a3df4 to
7466375
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
from
September 30, 2026 09:26
74c66e6 to
70aba13
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
September 30, 2026 21:58
7466375 to
e9cf344
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
2 times, most recently
from
October 2, 2026 22:51
45aff4b to
a87696b
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
October 2, 2026 22:51
e9cf344 to
3f2ad9b
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
2 times, most recently
from
October 3, 2026 01:55
fe910eb to
7e54771
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
October 3, 2026 01:55
41ddc2d to
6b3b5fa
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
October 3, 2026 02:16
6b3b5fa to
ffa4ffa
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
from
October 3, 2026 02:16
7e54771 to
3f2d42e
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
October 3, 2026 02:22
ffa4ffa to
cec2584
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
from
October 3, 2026 02:22
3f2d42e to
6b6bea4
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
October 3, 2026 02:24
cec2584 to
c69d03e
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
from
October 3, 2026 02:24
6b6bea4 to
5c74898
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
October 3, 2026 02:26
c69d03e to
6410f81
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
from
October 3, 2026 02:26
5c74898 to
4d887fe
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
October 3, 2026 03:30
6410f81 to
bbbd2c8
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
October 3, 2026 19:19
477b716 to
c3431e8
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
from
October 3, 2026 19:57
abf02dd to
5dc960b
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
2 times, most recently
from
October 3, 2026 20:20
22da557 to
77decc8
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
from
October 3, 2026 20:20
5dc960b to
0c4d9fe
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
October 3, 2026 20:27
77decc8 to
f4045cf
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
from
October 3, 2026 20:27
0c4d9fe to
3159d21
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
October 3, 2026 20:38
f4045cf to
5dd51e5
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
2 times, most recently
from
October 3, 2026 20:40
8b33ade to
94e28d3
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
October 3, 2026 20:40
5dd51e5 to
15fea2e
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
from
October 3, 2026 20:57
94e28d3 to
fa044fb
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
from
October 3, 2026 20:57
15fea2e to
8d3d3df
Compare
alxmrs
commented
Oct 3, 2026
|
|
||
| /// Extreme values a case can hold. | ||
| #[derive(Clone, Copy, Debug, PartialEq, Eq)] | ||
| enum Extreme { |
Member
Author
There was a problem hiding this comment.
I like the idea of varying data, good work.
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
from
October 3, 2026 21:22
fa044fb to
b7781f4
Compare
alxmrs
force-pushed
the
m4/v2-plan-shapes
branch
2 times, most recently
from
October 3, 2026 21:34
b6c57ab to
411926b
Compare
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
from
October 3, 2026 21:34
b7781f4 to
41c8b3b
Compare
…he v2 soak Data modes, each drawn from its own stream seeded by the case and applied after generation, so a seed without one replays as before: - ties: parameters from four values, so MAX, MIN and rankings tie exactly. At a tie the loss need not be differentiable (the median of three tied values moves at median(d) along any d, on both sides), so only directions that move each value by a function of itself, keeping ties tied, are compared; they check a tie's shared cotangent adds up. - nulls: a third of all values NULL, and NULL keys in data. - extreme: huge, tiny, -0.0, NaN and infinite values. - big: ten times the rows, in the soak only. Every table is now registered in random partitions of random batches. Near a kink the finite difference is no derivative, but ⟨∇L, d⟩ must still lie between the one-sided derivatives, so those points are now checked rather than skipped. The kink test compares a(h) − 4·a(h/2), which a dominant quadratic term cannot hide, and the tolerance allows Richardson's O(h) error at a C¹ kink. DDX_V2_MODES forces modes; DDX_V2_DEBUG prints a replayed case's steps and flags any relation that is not bit-reproducible. Found, and pinned in ad_findings.rs: the MAX rule finds its rows by float equality between the saved maximum and a recomputation of the region beneath it, constant subtrees included; a grouped SUM over several partitions is not bit-reproducible, so about half the time no row attains the maximum and the gradient is silently 0 everywhere. Ties, extreme values and NULL-focused cases found nothing new; the NULL leak accounts for every NULL-mode failure. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu
…re it max_finds_its_row_when_the_recomputed_values_jitter passes since #76 finds the rows attaining an extreme with windows over the recomputed rows themselves, so it runs as an ordinary test. Big mode's comment no longer calls that bug the reason it stays out of the PR gate. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
…ingless An extreme-values soak on the fixed stack failed three checks, each on a point the comparison could not judge rather than a wrong gradient: - -0.0 values tie exactly (a parameter at -0 equals data at -0), so several greatest() kinks meet and the one-sided bracket is unsound there, as it already is at ties; such a point is now screened the same way (seed 300105). - Over NaN data a MAX can depend on the order it meets rows in, so a permuted or repartitioned table can change the loss itself. The partition and row-order checks now compare gradients only where the engine computes the same loss (seed 300030). - Huge values take the calculus transforms past f64: the square of a loss near 1e270 overflows, and sin of a loss past 1e15 is rounding. Those transforms are skipped there (seed 300079). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
… expectations From an extreme-values soak on the fixed stack: - DataFusion's grouped MAX skips a NaN where its window and ungrouped MAX return it (seed 600845). ddx's MAX rule now copes (#76); the engine's inconsistency is pinned in ad_findings.rs's upstream section, ignored, with ddx not involved. - A metamorphic comparison whose expected value is past f64 (an overflowed or NaN gradient scaled by a chain factor) is not a comparison: huge values reach it by different roundings on each side (seed 600912). compare() now skips such entries. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
… ties mode An extreme-values soak on the fixed stack failed two finite-difference checks in -0.0 mode where exact ties meet (seeds 1300234, 1300617): MIN(abs(u)) with three parameters at -0, moved one at a time, and a greatest() tie inside a MIN tie that together make a smooth loss. At an exact tie any subgradient is a convention, and a direction that breaks the tie proves nothing, which is why ties mode moves every value by a function of itself. -0.0 values tie exactly too, so they now get the same directions, and a parameter at zero stays put: it ties with zero data and makes computed values tie (u + z against u when z is -0). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs
force-pushed
the
m4/v2-data-shapes
branch
from
October 3, 2026 21:50
41c8b3b to
c4ab48b
Compare
alxmrs
added a commit
that referenced
this pull request
Oct 3, 2026
Stacked on #91. This PR tests the adapter's SQL surface (`ad::sql`, `ad::sql_all`) the way users will reach it. ## What it adds - **`names`**: `grad(loss, t.col)` must equal the program's gradient in each of these setups: - the table read as `s1.t` or `"T t"`, while a **decoy** under its old name holds different values, so reading the wrong table can't pass by coincidence; - a CTE shadowing the table's name; - the call inside a subquery or a later CTE; - two losses in one statement (`grad(2L) − 2·grad(L) = 0`); - user tables named like ddx's steps (`value`, `saved_0`, `grad_0_w`, …), which must come through untouched. - **`train`**: three SGD steps written in SQL, one statement per table through `sql_all`. Each must move θ to exactly `θ − lr·∇L` as a fresh `ad::grad` program computes it. Some step along `−∇L` must lower a smooth loss (kinked losses are exempt). The context must be clean after every step. - **Stricter oracle**: a NaN or infinite gradient where the loss is finite and smooth along the direction is now a failure. It used to be skipped. - A refusal only the SQL surface makes (it plans `SELECT * FROM loss`, a slightly different plan from the program API's) is **tallied as coverage**, not failed. ## What it found About 4,900 cases: - **A table whose name has a capital has no gradient in SQL.** Pinned in `ad_findings.rs`. The gradient step is named after the table (`…_grad_0_W`). `ctx.register_table` folds that unquoted name to lower case, and `ad::sql` then reads it back quoted, so it isn't found. Any `"Weights"`-style table name hits this. Only the soak tries this variant until it's fixed, so the PR gate stays green. - **The MAX recomputation bug from #91 also produces NaN and inf**, not just zeros: the even split divides by zero attaining rows. The stricter check catches these now. - Minor, noted rather than pinned: with NaN or infinite *data*, ddx-core's quotient-rule expression evaluates `∞/∞ = NaN` where the partial's limit is 0. That's a v1 expression-form limit only infinite inputs reach, and those cases are exempt from the NaN check. Nothing wrong in schema-qualified names, shadowing CTEs, calls in subqueries or later CTEs, two losses per statement, look-alike user tables, or the SGD loop itself. Next and last in the stack: mutation testing of `ddx-ad`, to measure how fast this harness catches seeded bugs. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Open
alxmrs
added a commit
that referenced
this pull request
Oct 3, 2026
…(M4 9) (#93) Stacked on #92. A clean soak proves only as much as the soak can see. This PR measures that: it seeds one bug at a time into `ddx-ad` and the DataFusion adapter and times how long the v2 soak takes to catch each. ## What it adds - **`.github/scripts/mutation_test.py`**: 15 mutants, one per rule or contract: each reduce rule, tie sharing, the map rule, fan-in, the unreached-0 and NULL-row conventions, stop-gradient, the dims and ranking checks, step naming, `release`, and `grad` in SQL's dims/values split. Each is an exact text replacement that must match its file exactly once, so a stale mutant fails loudly instead of silently testing nothing. The source is restored afterwards, even on an interrupt. - **A clean baseline.** The script first runs the soak unmutated over the same seeds and skips every seed that fails there. The variants that always meet a known, pinned bug are turned off (`DDX_V2_SKIP_KNOWN`). So a kill is the mutant's doing. (The first attempt didn't have this; a known bug "killed" a mutant.) - **Kills by refusal.** If a mutant makes ddx accept 5 points fewer of the generated cases than the baseline, it counts as killed. A `ddx_stop_gradient` ddx no longer understands gets *refused*, not answered wrongly, so the only symptom is the refusal rate. - **`.github/workflows/mutation.yml`**: weekly and manually triggerable. The table goes to the job summary; a survivor fails the job but files no issue, since it's a gap in the tests, not a bug in ddx. - `DDX_SOAK_STOP_ON_FAIL` for the runner. ## Results The first sweep killed 10 of 15. Each survivor was a blind spot, now closed: | survivor | why | fix | |---|---|---| | `null-row-zero` | the clean baseline had turned NULLs off | keep NULLs on; the baseline skips the known NULL leak | | `stop-gradient-ignored` | the damage shows as refusals, not wrong answers | refusal-rate kill | | `ranking-check-off` | the generator only wrote total rankings | non-total rankings over what a `wrt` table feeds; ddx must refuse them | | `gradient-names-collide` | `w`, `b`, `u` never shared a name | `wrt` tables `t` and `s1.t` | | `sql-case-sensitive` | one value column per table hid the wrong dims/values split | a two-value table, `grad(loss, t2.val, t2.VAL2)` read with `SELECT *` | Now all 15 die: | mutant | time to kill | first failure | |---|---|---| | sum-halved | 1 case | vjp-vs-weighted-grad | | mean-undivided, mean-over-sum, partial-is-one | 1 case | finite-diff | | fan-in-first | 1 case | read-twice | | dims-check-off | 1 case | dims-check | | release-keeps-value | 1 case | catalog | | extreme-everyone | 5 cases, 2s | finite-diff | | sql-case-sensitive | 8 cases, 2s | names | | gradient-names-collide | 14 cases, 6s | names | | tie-unshared (wrong only at ties) | 16 cases, 8s | finite-diff | | null-row-zero | 19 cases, 8s | null | | unreached-one | 91 cases, 43s | finite-diff | | ranking-check-off | 82 cases, 44s | contract | | stop-gradient-ignored | 120s | accepts 73% of cases vs 80% unmutated | The slowest three are where the soak is thinnest: rows no gradient reaches, non-total rankings, and stop-gradient. **A finding along the way.** The non-total ranking check first fired on a ranking over *constant* data, which ddx doesn't need to refuse. But it shows a latent risk in the same family as #91's MAX bug: constant subtrees are recomputed in each backward step, unread and unchecked. A tie-breaking-free ranking over constant data could keep a different row on the way back, and a `CASE` on it would then take a different branch. Recorded in `ad_findings.rs` under the recomputation finding. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #90. This PR varies the data the v2 soak generates, where #90 varied the plan.
What it adds
Data modes. Each is drawn from its own random stream seeded by the case and applied after generation, so a seed without a mode replays exactly as before.
DDX_V2_MODES=…forces modes on every case.median(d)along anyd, the same on both sides, yet has no gradient, so ddx's convention is one valid subgradient among many. Only tie-preserving directions are compared: each value moves by a function of itself, so equal values stay equal. These check that a tie's shared cotangent adds up.-0.0, and NaN or ±inf in data.Every table is now also registered across random partitions and batches.
Sharper oracle:
⟨∇L, d⟩must lie between the one-sided derivatives, so those points are now checked instead of skipped. That holds for any convention ddx pins, but only for one kink at a time. At an exact tie several kinks meet, so ties mode screens those points instead.a(h) − 4·a(h/2), which a dominant quadratic term can't hide.greatest(v, 0)²at 0).sinof a sum near 1e8).Triage.
DDX_V2_DEBUG=1withreplay_one_seedprints each step's table and flags any relation in the case that isn't bit-reproducible across 30 runs. That's how the bug below was found.What it found
MAX finds its rows by float equality with a recomputation that isn't bit-reproducible. ddx saves the MAX and recomputes the region beneath it, constant subtrees included, then sends the cotangent to the rows whose recomputed value equals the saved maximum. DataFusion's grouped
SUMover several partitions adds partial sums in arrival order: 18 distinct bit patterns in 100 runs in the repro. So about half the time no row attains the maximum, and the whole gradient is silently 0. This happens on a default context with ordinary SQL; the only requirement is a table spread over several partitions, which real tables are. It's pinned inad_findings.rs. MIN is built the same way. Possible fixes: save what a MAX/MIN reads instead of recomputing it, or find the attaining rows in the same pass that computes the maximum.In total, about 7,500 cases across the mode-focused soaks:
Big mode stays out of the bounded suites: the bug it finds is nondeterministic, and a PR gate has to be reproducible.
Next in the stack:
ad::sqlname collisions and SGD training loops, then mutation testing of ddx-ad.🤖 Generated with Claude Code
https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu