Skip to content

test: a simulation soak for query-level AD, and three bugs it found (M4 5) - #89

Merged
alxmrs merged 5 commits into
mainfrom
m4/v2-simulation
Oct 3, 2026
Merged

alxmrs merged 5 commits into
mainfrom
m4/v2-simulation

Conversation

@alxmrs

@alxmrs alxmrs commented Sep 29, 2026

Copy link
Copy Markdown
Member

Adversarial property/fuzz testing for the M3/M4 stack (#69–#83): a generator of loss queries, a finite-difference oracle that shares nothing with ddx, a set of metamorphic relations, and a nightly soak. Stacked on #82.

What it checks

crates/ddx-datafusion/tests/ad_simulation.rs generates loss queries as DAGs of CTEs over small random tables. It uses every primitive in design.md §4.3 and the shapes around them: maps (including about 10% drawn from ddx-core's own gen_expr), joins on shared dims and cross joins, LEFT JOIN with COALESCE, CASE on data, grouped SUM/AVG/MAX/MIN, repeated identical aggregates, HAVING, dim and value filters, IN/NOT IN/EXISTS semi-joins, rank filters, ORDER BY … LIMIT, UNION ALL, DISTINCT, softmax with and without ddx_stop_gradient on the shift, and scalar subqueries over data. Tables get NULLs, orphan rows, repeated keys in data, shuffled rows, and BIGINT/INT/VARCHAR dims.

Oracle. The loss is a query DataFusion can run without ddx, so each gradient is checked against (L(θ+hd) − L(θ−hd))/2h. It uses random directions and single entries. The loss is screened at kinks (the second difference must shrink like h²) and Richardson-extrapolated. A screened point is counted, and the bounded suites assert how many cases were accepted.

Metamorphic relations, where ddx is compared with itself:

  • the calculus: ∇(cL), ∇(L+c), ∇(L²), ∇sin L, a loss CTE read twice, L·sg(L), L + sg(L);
  • grad = vjp seeded with 1 (and 2.5); vjp(R, c) = grad(Σ R·c); vjp is linear in c;
  • invariance to CTE vs inline subqueries, the unoptimized plan, wrt order, case and splitting, row order, and partitions 1/4/7;
  • the program contract: a program built at θ₀ and run at θ₁ (sometimes with extra rows) equals a fresh one; repeated dims are refused by the checks; after run only the value and gradients remain, after release nothing does, and user tables are untouched; three programs spawned concurrently on one context don't interfere (S11);
  • the SQL surface: grad(loss, t.val) equals the program, and an SGD join equals θ − 0.1·∇L. The statement is spelled with GRAD/Grad (, comments holding grad( and multibyte text, a quoted "Loss Fn" CTE, W.VAL, and extra CTEs around the loss;
  • shape: one gradient row per table row, Float64, NULL exactly where the value is NULL.

A refusal is always allowed and is tallied by kind. A panic, an Internal or InvalidPlan error, a program that is accepted but fails to run, or a wrong number is a failure.

Running it. Six bounded suites run under cargo test (24 seeds each, about 16s in debug). soak_v2_query_ad is #[ignore]d. It uses the same DDX_SOAK_* knobs and log format as ddx-core's soak, and now runs nightly as a second job. .github/scripts/report_fuzz_findings.sh now reads each soak's own REPRO line, and no longer counts a top-level log twice. DDX_V2_SEED=<n> with replay_one_seed prints a case's SQL and tables. DDX_V2_NULL_PCT=0 turns off NULL injection, to hunt past the NULL bug; it replays the same case minus its NULLs, which is how every NULL failure below was attributed.

What it found

Soaks totalled about 6,500 cases, of which about 84% were accepted. That gave about 17k finite-difference comparisons (fewer than 0.1% screened) and about 100k metamorphic comparisons. Every failure maps to one of three bugs, pinned in tests/ad_findings.rs as #[ignore = "known bug: …"] repros with hand-worked gradients (CONTRIBUTING's failing-test-first convention; -- --ignored shows all four failing):

  1. Silently wrong gradient at NULL rows. SUM, AVG, MAX and MIN skip a row whose argument is NULL, but ddx still sends the group's cotangent through that row's other inputs. In SUM(p.val + q.val) with q NULL at row 1, ddx gives ∂/∂p(1) = 1 where the answer is 0. This happens whether the NULL is in a parameter or in constant data, and it accounts for every finite-difference failure (all pass once the case's NULLs are removed).
  2. Two rank filters in one recomputed region are accepted, then fail to run. The optimizer gives both window columns the same name, and the rebuilt region keeps both.
  3. An unoptimized plan with a constant join condition under a CASE is accepted by grad_plan (which takes any LogicalPlan, a DataFrame's included), and then its backward step fails DataFusion's schema check.

No failures came from the calculus identities, vjp, wrt handling, the SQL rewriter, dim types, program reuse, concurrency or catalog hygiene.

Also worth knowing, though not a failure: ddx refuses some rankings the generator makes total over the CTE's dims ("a window function whose PARTITION BY and ORDER BY…", "over rows joined to…"). These are conservative refusals, not wrong answers, but they are coverage the rank-select rule could have.

Not covered

The Python ddxdb.ad layer (it wraps the same Rust calls), Float32 value columns (the finite-difference oracle would need f32 tolerances), and WITH RECURSIVE, which ddx refuses.

🤖 Generated with Claude Code

https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu

alxmrs added a commit that referenced this pull request Sep 29, 2026
… collide

From the v2 soak (#89): a region is rebuilt keeping every column, so a rank
filter over a rank filter gave DataFusion a window over rows that already
carried an identically printed window column, and the step failed with
DuplicateUnqualifiedField after the program was accepted. Each window
column is now replaced in its place by `CASE WHEN true THEN w ELSE w END`
as soon as it is computed: the same values and column number, under a
name no later window can take.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 29, 2026
…o coalesce

From the v2 soak (#89, #90, #92):

- SUM skips a row whose argument is NULL, but the reduce rule still gave
  that row the group's cotangent, so its other inputs got gradient: the p
  of SUM(p + q) with q NULL, or SUM(d - p) with a NULL in constant data.
  The seed is now NULL where the argument is.
- A gradient step was named after its table case and all (`…_grad_0_W`);
  DataFusion folds the unquoted name it is registered under, so it could
  not be found again. Step names are lower case.
- The gradient step used coalesce, which DataFusion 54 runs only after
  SimplifyExpressions has rewritten it; it is now a CASE, so a program runs
  on a context without that rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 29, 2026
From the v2 soak (#89, #91):

- MAX and MIN sent the cotangent to the recomputed rows equal to the
  saved extreme. A recomputation need not match the forward pass to the
  bit (DataFusion's grouped SUM over several partitions adds in arrival
  order), so about half the time no row attained the saved maximum: a
  silent zero gradient, or NaN/inf from dividing by zero attaining rows.
  The extreme, the count of rows attaining it, and AVG's count are now
  windows over the recomputed rows themselves, partitioned by the group's
  keys, so the comparison is always with the values being compared. The
  saved aggregate is no longer read back, and neither is a second copy of
  the region.
- AVG, MAX and MIN also get the NULL-row mask SUM got in #74 (applied in
  the rebase of the rules commit): no gradient through a row the
  aggregate skips.
- Tests for both, and for a rank filter over a rank filter (fixed in #73).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 29, 2026
From the v2 soak (#89, #90): grad_plan takes any LogicalPlan, a
DataFrame's unanalyzed one included. There a CASE between a BIGINT and a
DOUBLE branch is typed before type coercion, so the step's logical schema
disagreed with the batches it produced and run failed with "Mismatch
between schema and batches" after the program was accepted. run_step now
registers each step under the schema of the physical plan that ran.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 29, 2026
From the v2 soak (#89, #92), fixed in ddx-ad (#74): pinned in Python too,
since ddxdb.ad reads gradient steps back by name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
@alxmrs

alxmrs commented Sep 29, 2026

Copy link
Copy Markdown
Member Author

Thanks: all three findings were real ddx bugs, and they're fixed in the stack below. This branch is rebased onto it, and ad_findings.rs now runs the repros as ordinary tests, each naming its fix.

Finding Fix
NULL rows leak gradient #74, #76: each reduce rule's seed is NULL where its argument is NULL. Later soaks found three more ways the same leak got through, all fixed: vjp's seed (#74), a NULL term added to a real one in a multi-aggregate loss (#74), and a MAX/MIN row that doesn't attain the extreme (#76).
Two rank filters collide #73: each window column is renamed in place as soon as it's computed
Unoptimized CASE over integer data #79: a step's table takes the schema of the physical plan that ran

My commits on this branch:

  • un-ignore: takes the three repros off #[ignore].
  • stopped-softmax fix: a softmax whose shift is stopped no longer offers its exponentials and their sum for reuse. Only the ratio is shift-invariant, so reading Σ exp(u − sg(max)) directly makes stop-gradient's gradient differ from the finite difference by definition (seed 100607; JAX agrees with ddx).

🤖 Generated with Claude Code

alxmrs added a commit that referenced this pull request Sep 30, 2026
… collide

From the v2 soak (#89): a region is rebuilt keeping every column, so a rank
filter over a rank filter gave DataFusion a window over rows that already
carried an identically printed window column, and the step failed with
DuplicateUnqualifiedField after the program was accepted. Each window
column is now replaced in its place by `CASE WHEN true THEN w ELSE w END`
as soon as it is computed: the same values and column number, under a
name no later window can take.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 30, 2026
…o coalesce

From the v2 soak (#89, #90, #92):

- SUM skips a row whose argument is NULL, but the reduce rule still gave
  that row the group's cotangent, so its other inputs got gradient: the p
  of SUM(p + q) with q NULL, or SUM(d - p) with a NULL in constant data.
  The seed is now NULL where the argument is.
- A gradient step was named after its table case and all (`…_grad_0_W`);
  DataFusion folds the unquoted name it is registered under, so it could
  not be found again. Step names are lower case.
- The gradient step used coalesce, which DataFusion 54 runs only after
  SimplifyExpressions has rewritten it; it is now a CASE, so a program runs
  on a context without that rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 30, 2026
From the v2 soak (#89, #91):

- MAX and MIN sent the cotangent to the recomputed rows equal to the
  saved extreme. A recomputation need not match the forward pass to the
  bit (DataFusion's grouped SUM over several partitions adds in arrival
  order), so about half the time no row attained the saved maximum: a
  silent zero gradient, or NaN/inf from dividing by zero attaining rows.
  The extreme, the count of rows attaining it, and AVG's count are now
  windows over the recomputed rows themselves, partitioned by the group's
  keys, so the comparison is always with the values being compared. The
  saved aggregate is no longer read back, and neither is a second copy of
  the region.
- AVG, MAX and MIN also get the NULL-row mask SUM got in #74 (applied in
  the rebase of the rules commit): no gradient through a row the
  aggregate skips.
- Tests for both, and for a rank filter over a rank filter (fixed in #73).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 30, 2026
From the v2 soak (#89, #90): grad_plan takes any LogicalPlan, a
DataFrame's unanalyzed one included. There a CASE between a BIGINT and a
DOUBLE branch is typed before type coercion, so the step's logical schema
disagreed with the batches it produced and run failed with "Mismatch
between schema and batches" after the program was accepted. run_step now
registers each step under the schema of the physical plan that ran.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 30, 2026
From the v2 soak (#89, #92), fixed in ddx-ad (#74): pinned in Python too,
since ddxdb.ad reads gradient steps back by name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
@alxmrs

alxmrs commented Sep 30, 2026

Copy link
Copy Markdown
Member Author

Composability review: the oracles are an asset, make them reusable

Praise: the finite-difference oracle, the simulation generator shared with ddx_core::test_utils, and the mutation-testing pass (#93) are the real evidence that this AD is correct. They are also the most valuable thing a second adapter could inherit, because "matches jax.grad/finite difference on random plans" is the conformance test for any engine.

Suggestion: ad_simulation.rs (about 5,000 lines) is coupled to SessionContext. If the harness took a small engine trait (run SQL, produce Substrait, run a program, return a table) the same suite could run against the DuckDB adapter and any other, and would double as the "is your Substrait producer good enough for ddx?" test that other projects would want. Cheap to do now, expensive once a second engine has its own copy.

@alxmrs

alxmrs commented Oct 2, 2026

Copy link
Copy Markdown
Member Author

Re: reusable oracles

Agreed it's valuable, but I'd do it once a second adapter is in the repo, not now. The harness's engine surface is wider than run-SQL / produce-Substrait / run-a-program: optimizer variants, Substrait mutations, partition and row-order rewrites. The right trait is easier to cut against a real second engine than to guess at. The first half exists now: ddx_ad::Backend (#79) is the 'run a program' part, and tests/backend.rs already drives DataFusion through it. Recorded in design.md S14 (#82). One small change here: the refusal classifier treats an AdError kind it doesn't know as a bug (14f7d87), since AdError is now non_exhaustive.

🤖 Generated with Claude Code

@alxmrs
alxmrs force-pushed the m4/v2-simulation branch 2 times, most recently from 2bf3667 to 97936ff Compare October 3, 2026 19:05
@alxmrs
alxmrs force-pushed the m4/docs branch 2 times, most recently from 844cee5 to dd8eec9 Compare October 3, 2026 19:07
@alxmrs
alxmrs force-pushed the m4/v2-simulation branch 2 times, most recently from b8968f0 to d9c4db1 Compare October 3, 2026 20:20
@alxmrs
alxmrs force-pushed the m4/docs branch 2 times, most recently from ae2506b to 7f7d446 Compare October 3, 2026 20:27
@alxmrs
alxmrs force-pushed the m4/v2-simulation branch 2 times, most recently from c14cc2d to 567e6c9 Compare October 3, 2026 20:40

@alxmrs alxmrs left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Base automatically changed from m4/docs to main October 3, 2026 21:22
alxmrs and others added 5 commits October 3, 2026 14:22
ad_simulation.rs generates loss queries from the relational primitives
(map, joins on shared dims, grouped SUM/AVG/MAX/MIN, filters, semi-joins,
a rank filter, ORDER BY/LIMIT, UNION ALL, DISTINCT, softmax with and
without its shift stopped, NULLs, repeated keys in data, three dim
storage types) and checks every gradient against a finite difference of
the query DataFusion computes, screened at kinks and
Richardson-extrapolated.

Around that oracle: the calculus (scale, shift, chain, a loss read twice,
stop-gradient products), grad = vjp seeded with 1, vjp(R, c) = grad(Σ R·c)
and vjp's linearity, invariance to CTE inlining, the unoptimized plan,
the wrt list, row order and partition count, the program contract
(reuse on new values, the dims check, concurrent programs, the catalog
run and release leave), and grad(loss, t.col) in SQL, spelled several
ways, with an SGD step as a join.

Six bounded suites run in cargo test; soak_v2_query_ad runs nightly
beside ddx-core's soak, with the same log format. The report script now
takes each soak's own repro line, and counts a top-level log once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu
Each is reduced to a small query with a gradient worked by hand, and
ignored as a known bug until its fix lands:

- A row an aggregate skips as NULL still sends its cotangent through its
  other inputs: SUM(p + q) with q NULL at row 1 gives ∂/∂p(1) = 1, not 0.
  Silently wrong, from a NULL in a parameter or in constant data.
- Two rank filters in one recomputed region: accepted, then the backward
  step has two window columns of one name and DataFusion refuses it.
- An unoptimized plan (grad_plan takes any LogicalPlan) with a constant
  join condition under a CASE: accepted, then its backward step does not
  match its schema.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu
…gnore them

The NULL-row leak (fixed in #74 and #76), the stacked rank filters (#73)
and the unoptimized CASE (#79) now pass, so their repros in ad_findings.rs
run as ordinary tests, each entry naming its fix. DDX_V2_NULL_PCT's doc no
longer calls the NULL leak a known bug.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
…o itself

A soak on the fixed stack (seed 100607) reported a gradient 0.07 off the
finite difference. ddx was right: the softmax's shift was stopped, and a
later relation reused its unnormalized sum S = Σ exp(u - sg(max u))
directly, inside greatest(...). Only the ratio is shift-invariant, so
there stop-gradient makes the gradient differ from the loss's derivative
by definition (JAX gives ddx's value). Those two relations are now not
offered for reuse when the shift is stopped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
AdError is non_exhaustive now (#70).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
@alxmrs
alxmrs merged commit 61fa808 into main Oct 3, 2026
9 checks passed
@alxmrs
alxmrs deleted the m4/v2-simulation branch October 3, 2026 21:34
@github-actions github-actions Bot mentioned this pull request Oct 3, 2026
alxmrs added a commit that referenced this pull request Oct 3, 2026
…s (M4 6) (#90)

Stacked on #89. Design §4.2 names plan shapes as v2's recurring coverage
risk, and two of #89's three bugs were plan shapes ddx accepted and then
mis-emitted. This PR varies the plan without changing the query.

## What it adds

A `shapes` property group in `ad_simulation.rs`:

- **Optimizer invariance.** Each case is re-planned with one rule
removed (twice), with a random ~70% of the rules, with the rules
shuffled, and with none. The loss must first compute the same value on
that context. After that the gradient must match, or ddx must refuse.
- **Substrait rewrites** (`tests/ad_simulation/mutate.rs`). Random nodes
are rewritten into equivalent shapes: an identity projection, a
permutation and its inverse, `WHERE true`, a computed-then-dropped
column, swapped inner-join sides (fields remapped, order restored by
emit), a `FILTER (WHERE true)` copy of a measure, and a sort not under a
fetch. These cover the optimized and the unoptimized plan. DataFusion
consumes every rewrite and must compute the same loss before ddx sees
it, so a bad rewrite is skipped as the harness's mistake, not reported
as ddx's.

The soak summary now tallies which rewrites were compared, refused (and
why), or not equivalent. It also tallies engine faults.
`DDX_V2_PROPS=shapes` spends a soak on one group.

## What it found

About 5,000 cases with the `shapes` group. Every new failure traced to
DataFusion, not ddx:

- **Upstream DataFusion 54 wrong result**, pinned in `ad_findings.rs`
with ddx not involved. Without `push_down_limit`, a sort beneath a limit
is dropped under a join once the projection above it removes the sort
key: `SELECT t.i FROM (SELECT i FROM u ORDER BY val DESC LIMIT 1) t
CROSS JOIN one` returns `i = 1` instead of `0`. ddx's recomputed regions
have exactly this shape, so on such a context its gradient lands on the
wrong rows. DataFusion's default rules fuse the limit into the sort
first, which hides it.
- **Physical-planning faults** on steps whose logical plan DataFusion
accepted. These are tallied, not failed: DataFusion can't execute
`COALESCE` unless `SimplifyExpressions` rewrote it (ddx's gradient steps
use it, so the backward program depends on that rule),
`ProjectionPushdown` rejects a step with duplicate unqualified column
names, and the sanity check rejects a step with a sort in a join input.
I couldn't reduce the last two to plain SQL, so their root cause is
unconfirmed.

Two refinements to #89's findings:

- Bug 3's trigger is a **CASE choosing between integer columns on a
varied condition**; the join condition was incidental. The repro is
sharpened and renamed.
- ddx refuses an aggregate measure with a `FILTER`. That's an allowed
refusal, now visible in the coverage tally.

A related note for bug 2: DataFusion's Substrait consumer refuses two
identical expressions in one relation (they share a name), so whatever
ddx emits must never repeat one.

Next in the stack: NULL-heavy generation, deliberate ties, extreme
values, and larger multi-batch tables.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant