Skip to content

test: ties, NULLs, extreme values and big tables in the v2 soak (M4 7) - #91

Merged
alxmrs merged 5 commits into
mainfrom
m4/v2-data-shapes
Oct 3, 2026
Merged

alxmrs merged 5 commits into
mainfrom
m4/v2-data-shapes

Conversation

@alxmrs

@alxmrs alxmrs commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Stacked on #90. This PR varies the data the v2 soak generates, where #90 varied the plan.

What it adds

Data modes. Each is drawn from its own random stream seeded by the case and applied after generation, so a seed without a mode replays exactly as before. DDX_V2_MODES=… forces modes on every case.

  • ties (15%): parameters take one of four values, so MAX, MIN and rankings tie exactly. At a tie the loss need not be differentiable. For example, the median of three tied values moves at median(d) along any d, the same on both sides, yet has no gradient, so ddx's convention is one valid subgradient among many. Only tie-preserving directions are compared: each value moves by a function of itself, so equal values stay equal. These check that a tie's shared cotangent adds up.
  • nulls (10%): about a third of all values are NULL, with NULL keys in data tables.
  • extreme (8%): huge, tiny or subnormal values, -0.0, and NaN or ±inf in data.
  • big (4%, soak only): ten times the rows per relation.

Every table is now also registered across random partitions and batches.

Sharper oracle:

  • Near a kink, ⟨∇L, d⟩ must lie between the one-sided derivatives, so those points are now checked instead of skipped. That holds for any convention ddx pins, but only for one kink at a time. At an exact tie several kinks meet, so ties mode screens those points instead.
  • The kink test uses a(h) − 4·a(h/2), which a dominant quadratic term can't hide.
  • The tolerance allows Richardson extrapolation's O(h) error at a C¹ kink (greatest(v, 0)² at 0).
  • A step too large for the loss to be linear over is skipped (e.g. sin of a sum near 1e8).

Triage. DDX_V2_DEBUG=1 with replay_one_seed prints each step's table and flags any relation in the case that isn't bit-reproducible across 30 runs. That's how the bug below was found.

What it found

MAX finds its rows by float equality with a recomputation that isn't bit-reproducible. ddx saves the MAX and recomputes the region beneath it, constant subtrees included, then sends the cotangent to the rows whose recomputed value equals the saved maximum. DataFusion's grouped SUM over several partitions adds partial sums in arrival order: 18 distinct bit patterns in 100 runs in the repro. So about half the time no row attains the maximum, and the whole gradient is silently 0. This happens on a default context with ordinary SQL; the only requirement is a table spread over several partitions, which real tables are. It's pinned in ad_findings.rs. MIN is built the same way. Possible fixes: save what a MAX/MIN reads instead of recomputing it, or find the attaining rows in the same pass that computes the maximum.

In total, about 7,500 cases across the mode-focused soaks:

mode cases result
ties ~760 with the final oracle 1,848 exact tie-preserving comparisons, no failures
extreme ~1,400 nothing new
big ~900 the MAX bug
nulls ~700 only the known NULL leak
fd only, NULLs off ~4,700 14,558 agreements, only the known stacked-rank bug

Big mode stays out of the bounded suites: the bug it finds is nondeterministic, and a PR gate has to be reproducible.

Next in the stack: ad::sql name collisions and SGD training loops, then mutation testing of ddx-ad.

🤖 Generated with Claude Code

https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu

alxmrs added a commit that referenced this pull request Sep 29, 2026
From the v2 soak (#89, #91):

- MAX and MIN sent the cotangent to the recomputed rows equal to the
  saved extreme. A recomputation need not match the forward pass to the
  bit (DataFusion's grouped SUM over several partitions adds in arrival
  order), so about half the time no row attained the saved maximum: a
  silent zero gradient, or NaN/inf from dividing by zero attaining rows.
  The extreme, the count of rows attaining it, and AVG's count are now
  windows over the recomputed rows themselves, partitioned by the group's
  keys, so the comparison is always with the values being compared. The
  saved aggregate is no longer read back, and neither is a second copy of
  the region.
- AVG, MAX and MIN also get the NULL-row mask SUM got in #74 (applied in
  the rebase of the rules commit): no gradient through a row the
  aggregate skips.
- Tests for both, and for a rank filter over a rank filter (fixed in #73).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
@alxmrs

alxmrs commented Sep 29, 2026

Copy link
Copy Markdown
Member Author

The MAX recomputation bug was real and is fixed in #76. The rows attaining an extreme, and how many there are, are now windows over the recomputed rows themselves, never a comparison with the saved value. A later ties-mode soak (seed 2101304) showed the same jitter can also flip which of two exactly-tied groups wins. So a row now attains a finite extreme if it's equal to it or within 8 ulps of it; a tie that rounding breaks is shared, the same way every run.

Other changes on this branch, all from soaks of the fixed stack:

  • Un-ignored: the MAX finding. Big mode no longer describes that bug as the reason it stays out of the PR gate.
  • New upstream bug, pinned: DataFusion's grouped MAX skips a NaN, while its window and ungrouped MAX return it (seed 600845). feat(ddx-ad): AVG, MAX and MIN rules; rank-select and stop-gradient (M3 8) #76 now leaves NaN arguments out of the attainment window, so ddx agrees with any MAX that produced a number.
  • Screens for points the comparison can't judge:
    • −0.0 values tie exactly, so that mode gets ties mode's tie-preserving directions, and a parameter at zero stays fixed (seeds 300105, 1300234, 1300617).
    • Over NaN data a MAX can depend on row order, so the partition and row-order checks only compare where the loss is unchanged (300030).
    • Huge values push the calculus transforms past f64 (300079).
    • Expected values that aren't finite are skipped (600912).

🤖 Generated with Claude Code

alxmrs added a commit that referenced this pull request Sep 30, 2026
From the v2 soak (#89, #91):

- MAX and MIN sent the cotangent to the recomputed rows equal to the
  saved extreme. A recomputation need not match the forward pass to the
  bit (DataFusion's grouped SUM over several partitions adds in arrival
  order), so about half the time no row attained the saved maximum: a
  silent zero gradient, or NaN/inf from dividing by zero attaining rows.
  The extreme, the count of rows attaining it, and AVG's count are now
  windows over the recomputed rows themselves, partitioned by the group's
  keys, so the comparison is always with the values being compared. The
  saved aggregate is no longer read back, and neither is a second copy of
  the region.
- AVG, MAX and MIN also get the NULL-row mask SUM got in #74 (applied in
  the rebase of the rules commit): no gradient through a row the
  aggregate skips.
- Tests for both, and for a rank filter over a rank filter (fixed in #73).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch 2 times, most recently from 45aff4b to a87696b Compare October 2, 2026 22:51
@alxmrs
alxmrs force-pushed the m4/v2-plan-shapes branch from e9cf344 to 3f2ad9b Compare October 2, 2026 22:51
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch 2 times, most recently from fe910eb to 7e54771 Compare October 3, 2026 01:55
@alxmrs
alxmrs force-pushed the m4/v2-plan-shapes branch from 41ddc2d to 6b3b5fa Compare October 3, 2026 01:55
@alxmrs
alxmrs force-pushed the m4/v2-plan-shapes branch from 6b3b5fa to ffa4ffa Compare October 3, 2026 02:16
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from 7e54771 to 3f2d42e Compare October 3, 2026 02:16
@alxmrs
alxmrs force-pushed the m4/v2-plan-shapes branch from ffa4ffa to cec2584 Compare October 3, 2026 02:22
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from 3f2d42e to 6b6bea4 Compare October 3, 2026 02:22
@alxmrs
alxmrs force-pushed the m4/v2-plan-shapes branch from cec2584 to c69d03e Compare October 3, 2026 02:24
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from 6b6bea4 to 5c74898 Compare October 3, 2026 02:24
@alxmrs
alxmrs force-pushed the m4/v2-plan-shapes branch from c69d03e to 6410f81 Compare October 3, 2026 02:26
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from 5c74898 to 4d887fe Compare October 3, 2026 02:26
@alxmrs
alxmrs force-pushed the m4/v2-plan-shapes branch from 6410f81 to bbbd2c8 Compare October 3, 2026 03:30
@alxmrs
alxmrs force-pushed the m4/v2-plan-shapes branch from 477b716 to c3431e8 Compare October 3, 2026 19:19
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from abf02dd to 5dc960b Compare October 3, 2026 19:57
@alxmrs
alxmrs force-pushed the m4/v2-plan-shapes branch 2 times, most recently from 22da557 to 77decc8 Compare October 3, 2026 20:20
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from 5dc960b to 0c4d9fe Compare October 3, 2026 20:20
@alxmrs
alxmrs force-pushed the m4/v2-plan-shapes branch from 77decc8 to f4045cf Compare October 3, 2026 20:27
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from 0c4d9fe to 3159d21 Compare October 3, 2026 20:27
@alxmrs
alxmrs force-pushed the m4/v2-plan-shapes branch from f4045cf to 5dd51e5 Compare October 3, 2026 20:38
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch 2 times, most recently from 8b33ade to 94e28d3 Compare October 3, 2026 20:40
@alxmrs
alxmrs force-pushed the m4/v2-plan-shapes branch from 5dd51e5 to 15fea2e Compare October 3, 2026 20:40
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from 94e28d3 to fa044fb Compare October 3, 2026 20:57
@alxmrs
alxmrs force-pushed the m4/v2-plan-shapes branch from 15fea2e to 8d3d3df Compare October 3, 2026 20:57

@alxmrs alxmrs left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM.


/// Extreme values a case can hold.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
enum Extreme {

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like the idea of varying data, good work.

@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from fa044fb to b7781f4 Compare October 3, 2026 21:22
@alxmrs
alxmrs force-pushed the m4/v2-plan-shapes branch 2 times, most recently from b6c57ab to 411926b Compare October 3, 2026 21:34
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from b7781f4 to 41c8b3b Compare October 3, 2026 21:34
Base automatically changed from m4/v2-plan-shapes to main October 3, 2026 21:50
alxmrs and others added 5 commits October 3, 2026 14:50
…he v2 soak

Data modes, each drawn from its own stream seeded by the case and
applied after generation, so a seed without one replays as before:

- ties: parameters from four values, so MAX, MIN and rankings tie
  exactly. At a tie the loss need not be differentiable (the median of
  three tied values moves at median(d) along any d, on both sides), so
  only directions that move each value by a function of itself, keeping
  ties tied, are compared; they check a tie's shared cotangent adds up.
- nulls: a third of all values NULL, and NULL keys in data.
- extreme: huge, tiny, -0.0, NaN and infinite values.
- big: ten times the rows, in the soak only.

Every table is now registered in random partitions of random batches.
Near a kink the finite difference is no derivative, but ⟨∇L, d⟩ must
still lie between the one-sided derivatives, so those points are now
checked rather than skipped. The kink test compares a(h) − 4·a(h/2),
which a dominant quadratic term cannot hide, and the tolerance allows
Richardson's O(h) error at a C¹ kink. DDX_V2_MODES forces modes;
DDX_V2_DEBUG prints a replayed case's steps and flags any relation that
is not bit-reproducible.

Found, and pinned in ad_findings.rs: the MAX rule finds its rows by float
equality between the saved maximum and a recomputation of the region
beneath it, constant subtrees included; a grouped SUM over several
partitions is not bit-reproducible, so about half the time no row
attains the maximum and the gradient is silently 0 everywhere. Ties,
extreme values and NULL-focused cases found nothing new; the NULL leak
accounts for every NULL-mode failure.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu
…re it

max_finds_its_row_when_the_recomputed_values_jitter passes since #76
finds the rows attaining an extreme with windows over the recomputed rows
themselves, so it runs as an ordinary test. Big mode's comment no longer
calls that bug the reason it stays out of the PR gate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
…ingless

An extreme-values soak on the fixed stack failed three checks, each on a
point the comparison could not judge rather than a wrong gradient:

- -0.0 values tie exactly (a parameter at -0 equals data at -0), so several
  greatest() kinks meet and the one-sided bracket is unsound there, as it
  already is at ties; such a point is now screened the same way (seed
  300105).
- Over NaN data a MAX can depend on the order it meets rows in, so a
  permuted or repartitioned table can change the loss itself. The
  partition and row-order checks now compare gradients only where the
  engine computes the same loss (seed 300030).
- Huge values take the calculus transforms past f64: the square of a loss
  near 1e270 overflows, and sin of a loss past 1e15 is rounding. Those
  transforms are skipped there (seed 300079).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
… expectations

From an extreme-values soak on the fixed stack:

- DataFusion's grouped MAX skips a NaN where its window and ungrouped MAX
  return it (seed 600845). ddx's MAX rule now copes (#76); the engine's
  inconsistency is pinned in ad_findings.rs's upstream section, ignored,
  with ddx not involved.
- A metamorphic comparison whose expected value is past f64 (an
  overflowed or NaN gradient scaled by a chain factor) is not a
  comparison: huge values reach it by different roundings on each side
  (seed 600912). compare() now skips such entries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
… ties mode

An extreme-values soak on the fixed stack failed two finite-difference
checks in -0.0 mode where exact ties meet (seeds 1300234, 1300617):
MIN(abs(u)) with three parameters at -0, moved one at a time, and a
greatest() tie inside a MIN tie that together make a smooth loss. At an
exact tie any subgradient is a convention, and a direction that breaks
the tie proves nothing, which is why ties mode moves every value by a
function of itself. -0.0 values tie exactly too, so they now get the same
directions, and a parameter at zero stays put: it ties with zero data and
makes computed values tie (u + z against u when z is -0).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from 41c8b3b to c4ab48b Compare October 3, 2026 21:50
@alxmrs
alxmrs merged commit 73fb319 into main Oct 3, 2026
9 checks passed
@alxmrs
alxmrs deleted the m4/v2-data-shapes branch October 3, 2026 22:02
alxmrs added a commit that referenced this pull request Oct 3, 2026
Stacked on #91. This PR tests the adapter's SQL surface (`ad::sql`,
`ad::sql_all`) the way users will reach it.

## What it adds

- **`names`**: `grad(loss, t.col)` must equal the program's gradient in
each of these setups:
- the table read as `s1.t` or `"T t"`, while a **decoy** under its old
name holds different values, so reading the wrong table can't pass by
coincidence;
  - a CTE shadowing the table's name;
  - the call inside a subquery or a later CTE;
  - two losses in one statement (`grad(2L) − 2·grad(L) = 0`);
- user tables named like ddx's steps (`value`, `saved_0`, `grad_0_w`,
…), which must come through untouched.
- **`train`**: three SGD steps written in SQL, one statement per table
through `sql_all`. Each must move θ to exactly `θ − lr·∇L` as a fresh
`ad::grad` program computes it. Some step along `−∇L` must lower a
smooth loss (kinked losses are exempt). The context must be clean after
every step.
- **Stricter oracle**: a NaN or infinite gradient where the loss is
finite and smooth along the direction is now a failure. It used to be
skipped.
- A refusal only the SQL surface makes (it plans `SELECT * FROM loss`, a
slightly different plan from the program API's) is **tallied as
coverage**, not failed.

## What it found

About 4,900 cases:

- **A table whose name has a capital has no gradient in SQL.** Pinned in
`ad_findings.rs`. The gradient step is named after the table
(`…_grad_0_W`). `ctx.register_table` folds that unquoted name to lower
case, and `ad::sql` then reads it back quoted, so it isn't found. Any
`"Weights"`-style table name hits this. Only the soak tries this variant
until it's fixed, so the PR gate stays green.
- **The MAX recomputation bug from #91 also produces NaN and inf**, not
just zeros: the even split divides by zero attaining rows. The stricter
check catches these now.
- Minor, noted rather than pinned: with NaN or infinite *data*,
ddx-core's quotient-rule expression evaluates `∞/∞ = NaN` where the
partial's limit is 0. That's a v1 expression-form limit only infinite
inputs reach, and those cases are exempt from the NaN check.

Nothing wrong in schema-qualified names, shadowing CTEs, calls in
subqueries or later CTEs, two losses per statement, look-alike user
tables, or the SGD loop itself.

Next and last in the stack: mutation testing of `ddx-ad`, to measure how
fast this harness catches seeded bugs.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot mentioned this pull request Oct 3, 2026
alxmrs added a commit that referenced this pull request Oct 3, 2026
…(M4 9) (#93)

Stacked on #92. A clean soak proves only as much as the soak can see.
This PR measures that: it seeds one bug at a time into `ddx-ad` and the
DataFusion adapter and times how long the v2 soak takes to catch each.

## What it adds

- **`.github/scripts/mutation_test.py`**: 15 mutants, one per rule or
contract: each reduce rule, tie sharing, the map rule, fan-in, the
unreached-0 and NULL-row conventions, stop-gradient, the dims and
ranking checks, step naming, `release`, and `grad` in SQL's dims/values
split. Each is an exact text replacement that must match its file
exactly once, so a stale mutant fails loudly instead of silently testing
nothing. The source is restored afterwards, even on an interrupt.
- **A clean baseline.** The script first runs the soak unmutated over
the same seeds and skips every seed that fails there. The variants that
always meet a known, pinned bug are turned off (`DDX_V2_SKIP_KNOWN`). So
a kill is the mutant's doing. (The first attempt didn't have this; a
known bug "killed" a mutant.)
- **Kills by refusal.** If a mutant makes ddx accept 5 points fewer of
the generated cases than the baseline, it counts as killed. A
`ddx_stop_gradient` ddx no longer understands gets *refused*, not
answered wrongly, so the only symptom is the refusal rate.
- **`.github/workflows/mutation.yml`**: weekly and manually triggerable.
The table goes to the job summary; a survivor fails the job but files no
issue, since it's a gap in the tests, not a bug in ddx.
- `DDX_SOAK_STOP_ON_FAIL` for the runner.

## Results

The first sweep killed 10 of 15. Each survivor was a blind spot, now
closed:

| survivor | why | fix |
|---|---|---|
| `null-row-zero` | the clean baseline had turned NULLs off | keep NULLs
on; the baseline skips the known NULL leak |
| `stop-gradient-ignored` | the damage shows as refusals, not wrong
answers | refusal-rate kill |
| `ranking-check-off` | the generator only wrote total rankings |
non-total rankings over what a `wrt` table feeds; ddx must refuse them |
| `gradient-names-collide` | `w`, `b`, `u` never shared a name | `wrt`
tables `t` and `s1.t` |
| `sql-case-sensitive` | one value column per table hid the wrong
dims/values split | a two-value table, `grad(loss, t2.val, t2.VAL2)`
read with `SELECT *` |

Now all 15 die:

| mutant | time to kill | first failure |
|---|---|---|
| sum-halved | 1 case | vjp-vs-weighted-grad |
| mean-undivided, mean-over-sum, partial-is-one | 1 case | finite-diff |
| fan-in-first | 1 case | read-twice |
| dims-check-off | 1 case | dims-check |
| release-keeps-value | 1 case | catalog |
| extreme-everyone | 5 cases, 2s | finite-diff |
| sql-case-sensitive | 8 cases, 2s | names |
| gradient-names-collide | 14 cases, 6s | names |
| tie-unshared (wrong only at ties) | 16 cases, 8s | finite-diff |
| null-row-zero | 19 cases, 8s | null |
| unreached-one | 91 cases, 43s | finite-diff |
| ranking-check-off | 82 cases, 44s | contract |
| stop-gradient-ignored | 120s | accepts 73% of cases vs 80% unmutated |

The slowest three are where the soak is thinnest: rows no gradient
reaches, non-total rankings, and stop-gradient.

**A finding along the way.** The non-total ranking check first fired on
a ranking over *constant* data, which ddx doesn't need to refuse. But it
shows a latent risk in the same family as #91's MAX bug: constant
subtrees are recomputed in each backward step, unread and unchecked. A
tie-breaking-free ranking over constant data could keep a different row
on the way back, and a `CASE` on it would then take a different branch.
Recorded in `ad_findings.rs` under the recomputation finding.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant