test: a simulation soak for query-level AD, and three bugs it found (M4 5) - #89
Conversation
… collide From the v2 soak (#89): a region is rebuilt keeping every column, so a rank filter over a rank filter gave DataFusion a window over rows that already carried an identically printed window column, and the step failed with DuplicateUnqualifiedField after the program was accepted. Each window column is now replaced in its place by `CASE WHEN true THEN w ELSE w END` as soon as it is computed: the same values and column number, under a name no later window can take. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
…o coalesce From the v2 soak (#89, #90, #92): - SUM skips a row whose argument is NULL, but the reduce rule still gave that row the group's cotangent, so its other inputs got gradient: the p of SUM(p + q) with q NULL, or SUM(d - p) with a NULL in constant data. The seed is now NULL where the argument is. - A gradient step was named after its table case and all (`…_grad_0_W`); DataFusion folds the unquoted name it is registered under, so it could not be found again. Step names are lower case. - The gradient step used coalesce, which DataFusion 54 runs only after SimplifyExpressions has rewritten it; it is now a CASE, so a program runs on a context without that rule. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
From the v2 soak (#89, #91): - MAX and MIN sent the cotangent to the recomputed rows equal to the saved extreme. A recomputation need not match the forward pass to the bit (DataFusion's grouped SUM over several partitions adds in arrival order), so about half the time no row attained the saved maximum: a silent zero gradient, or NaN/inf from dividing by zero attaining rows. The extreme, the count of rows attaining it, and AVG's count are now windows over the recomputed rows themselves, partitioned by the group's keys, so the comparison is always with the values being compared. The saved aggregate is no longer read back, and neither is a second copy of the region. - AVG, MAX and MIN also get the NULL-row mask SUM got in #74 (applied in the rebase of the rules commit): no gradient through a row the aggregate skips. - Tests for both, and for a rank filter over a rank filter (fixed in #73). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
From the v2 soak (#89, #90): grad_plan takes any LogicalPlan, a DataFrame's unanalyzed one included. There a CASE between a BIGINT and a DOUBLE branch is typed before type coercion, so the step's logical schema disagreed with the batches it produced and run failed with "Mismatch between schema and batches" after the program was accepted. run_step now registers each step under the schema of the physical plan that ran. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
From the v2 soak (#89, #92), fixed in ddx-ad (#74): pinned in Python too, since ddxdb.ad reads gradient steps back by name. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
4887387 to
d1146e4
Compare
|
Thanks: all three findings were real ddx bugs, and they're fixed in the stack below. This branch is rebased onto it, and
My commits on this branch:
🤖 Generated with Claude Code |
… collide From the v2 soak (#89): a region is rebuilt keeping every column, so a rank filter over a rank filter gave DataFusion a window over rows that already carried an identically printed window column, and the step failed with DuplicateUnqualifiedField after the program was accepted. Each window column is now replaced in its place by `CASE WHEN true THEN w ELSE w END` as soon as it is computed: the same values and column number, under a name no later window can take. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
…o coalesce From the v2 soak (#89, #90, #92): - SUM skips a row whose argument is NULL, but the reduce rule still gave that row the group's cotangent, so its other inputs got gradient: the p of SUM(p + q) with q NULL, or SUM(d - p) with a NULL in constant data. The seed is now NULL where the argument is. - A gradient step was named after its table case and all (`…_grad_0_W`); DataFusion folds the unquoted name it is registered under, so it could not be found again. Step names are lower case. - The gradient step used coalesce, which DataFusion 54 runs only after SimplifyExpressions has rewritten it; it is now a CASE, so a program runs on a context without that rule. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
From the v2 soak (#89, #91): - MAX and MIN sent the cotangent to the recomputed rows equal to the saved extreme. A recomputation need not match the forward pass to the bit (DataFusion's grouped SUM over several partitions adds in arrival order), so about half the time no row attained the saved maximum: a silent zero gradient, or NaN/inf from dividing by zero attaining rows. The extreme, the count of rows attaining it, and AVG's count are now windows over the recomputed rows themselves, partitioned by the group's keys, so the comparison is always with the values being compared. The saved aggregate is no longer read back, and neither is a second copy of the region. - AVG, MAX and MIN also get the NULL-row mask SUM got in #74 (applied in the rebase of the rules commit): no gradient through a row the aggregate skips. - Tests for both, and for a rank filter over a rank filter (fixed in #73). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
From the v2 soak (#89, #90): grad_plan takes any LogicalPlan, a DataFrame's unanalyzed one included. There a CASE between a BIGINT and a DOUBLE branch is typed before type coercion, so the step's logical schema disagreed with the batches it produced and run failed with "Mismatch between schema and batches" after the program was accepted. run_step now registers each step under the schema of the physical plan that ran. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
From the v2 soak (#89, #92), fixed in ddx-ad (#74): pinned in Python too, since ddxdb.ad reads gradient steps back by name. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
d1146e4 to
6c4f96e
Compare
6c4f96e to
f464728
Compare
|
Composability review: the oracles are an asset, make them reusable Praise: the finite-difference oracle, the simulation generator shared with Suggestion: |
f464728 to
0a77c58
Compare
0a77c58 to
14f7d87
Compare
|
Re: reusable oracles Agreed it's valuable, but I'd do it once a second adapter is in the repo, not now. The harness's engine surface is wider than run-SQL / produce-Substrait / run-a-program: optimizer variants, Substrait mutations, partition and row-order rewrites. The right trait is easier to cut against a real second engine than to guess at. The first half exists now: 🤖 Generated with Claude Code |
2bf3667 to
97936ff
Compare
844cee5 to
dd8eec9
Compare
97936ff to
993b14d
Compare
993b14d to
40470f9
Compare
b8968f0 to
d9c4db1
Compare
ae2506b to
7f7d446
Compare
d9c4db1 to
31be9c0
Compare
c14cc2d to
567e6c9
Compare
567e6c9 to
fb086a9
Compare
ad_simulation.rs generates loss queries from the relational primitives (map, joins on shared dims, grouped SUM/AVG/MAX/MIN, filters, semi-joins, a rank filter, ORDER BY/LIMIT, UNION ALL, DISTINCT, softmax with and without its shift stopped, NULLs, repeated keys in data, three dim storage types) and checks every gradient against a finite difference of the query DataFusion computes, screened at kinks and Richardson-extrapolated. Around that oracle: the calculus (scale, shift, chain, a loss read twice, stop-gradient products), grad = vjp seeded with 1, vjp(R, c) = grad(Σ R·c) and vjp's linearity, invariance to CTE inlining, the unoptimized plan, the wrt list, row order and partition count, the program contract (reuse on new values, the dims check, concurrent programs, the catalog run and release leave), and grad(loss, t.col) in SQL, spelled several ways, with an SGD step as a join. Six bounded suites run in cargo test; soak_v2_query_ad runs nightly beside ddx-core's soak, with the same log format. The report script now takes each soak's own repro line, and counts a top-level log once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu
Each is reduced to a small query with a gradient worked by hand, and ignored as a known bug until its fix lands: - A row an aggregate skips as NULL still sends its cotangent through its other inputs: SUM(p + q) with q NULL at row 1 gives ∂/∂p(1) = 1, not 0. Silently wrong, from a NULL in a parameter or in constant data. - Two rank filters in one recomputed region: accepted, then the backward step has two window columns of one name and DataFusion refuses it. - An unoptimized plan (grad_plan takes any LogicalPlan) with a constant join condition under a CASE: accepted, then its backward step does not match its schema. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu
…gnore them The NULL-row leak (fixed in #74 and #76), the stacked rank filters (#73) and the unoptimized CASE (#79) now pass, so their repros in ad_findings.rs run as ordinary tests, each entry naming its fix. DDX_V2_NULL_PCT's doc no longer calls the NULL leak a known bug. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
…o itself A soak on the fixed stack (seed 100607) reported a gradient 0.07 off the finite difference. ddx was right: the softmax's shift was stopped, and a later relation reused its unnormalized sum S = Σ exp(u - sg(max u)) directly, inside greatest(...). Only the ratio is shift-invariant, so there stop-gradient makes the gradient differ from the loss's derivative by definition (JAX gives ddx's value). Those two relations are now not offered for reuse when the shift is stopped. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
AdError is non_exhaustive now (#70). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
fb086a9 to
b998e80
Compare
…s (M4 6) (#90) Stacked on #89. Design §4.2 names plan shapes as v2's recurring coverage risk, and two of #89's three bugs were plan shapes ddx accepted and then mis-emitted. This PR varies the plan without changing the query. ## What it adds A `shapes` property group in `ad_simulation.rs`: - **Optimizer invariance.** Each case is re-planned with one rule removed (twice), with a random ~70% of the rules, with the rules shuffled, and with none. The loss must first compute the same value on that context. After that the gradient must match, or ddx must refuse. - **Substrait rewrites** (`tests/ad_simulation/mutate.rs`). Random nodes are rewritten into equivalent shapes: an identity projection, a permutation and its inverse, `WHERE true`, a computed-then-dropped column, swapped inner-join sides (fields remapped, order restored by emit), a `FILTER (WHERE true)` copy of a measure, and a sort not under a fetch. These cover the optimized and the unoptimized plan. DataFusion consumes every rewrite and must compute the same loss before ddx sees it, so a bad rewrite is skipped as the harness's mistake, not reported as ddx's. The soak summary now tallies which rewrites were compared, refused (and why), or not equivalent. It also tallies engine faults. `DDX_V2_PROPS=shapes` spends a soak on one group. ## What it found About 5,000 cases with the `shapes` group. Every new failure traced to DataFusion, not ddx: - **Upstream DataFusion 54 wrong result**, pinned in `ad_findings.rs` with ddx not involved. Without `push_down_limit`, a sort beneath a limit is dropped under a join once the projection above it removes the sort key: `SELECT t.i FROM (SELECT i FROM u ORDER BY val DESC LIMIT 1) t CROSS JOIN one` returns `i = 1` instead of `0`. ddx's recomputed regions have exactly this shape, so on such a context its gradient lands on the wrong rows. DataFusion's default rules fuse the limit into the sort first, which hides it. - **Physical-planning faults** on steps whose logical plan DataFusion accepted. These are tallied, not failed: DataFusion can't execute `COALESCE` unless `SimplifyExpressions` rewrote it (ddx's gradient steps use it, so the backward program depends on that rule), `ProjectionPushdown` rejects a step with duplicate unqualified column names, and the sanity check rejects a step with a sort in a join input. I couldn't reduce the last two to plain SQL, so their root cause is unconfirmed. Two refinements to #89's findings: - Bug 3's trigger is a **CASE choosing between integer columns on a varied condition**; the join condition was incidental. The repro is sharpened and renamed. - ddx refuses an aggregate measure with a `FILTER`. That's an allowed refusal, now visible in the coverage tally. A related note for bug 2: DataFusion's Substrait consumer refuses two identical expressions in one relation (they share a name), so whatever ddx emits must never repeat one. Next in the stack: NULL-heavy generation, deliberate ties, extreme values, and larger multi-batch tables. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Adversarial property/fuzz testing for the M3/M4 stack (#69–#83): a generator of loss queries, a finite-difference oracle that shares nothing with ddx, a set of metamorphic relations, and a nightly soak. Stacked on #82.
What it checks
crates/ddx-datafusion/tests/ad_simulation.rsgenerates loss queries as DAGs of CTEs over small random tables. It uses every primitive in design.md §4.3 and the shapes around them: maps (including about 10% drawn from ddx-core's owngen_expr), joins on shared dims and cross joins, LEFT JOIN with COALESCE, CASE on data, grouped SUM/AVG/MAX/MIN, repeated identical aggregates, HAVING, dim and value filters, IN/NOT IN/EXISTS semi-joins, rank filters, ORDER BY … LIMIT, UNION ALL, DISTINCT, softmax with and withoutddx_stop_gradienton the shift, and scalar subqueries over data. Tables get NULLs, orphan rows, repeated keys in data, shuffled rows, and BIGINT/INT/VARCHAR dims.Oracle. The loss is a query DataFusion can run without ddx, so each gradient is checked against
(L(θ+hd) − L(θ−hd))/2h. It uses random directions and single entries. The loss is screened at kinks (the second difference must shrink like h²) and Richardson-extrapolated. A screened point is counted, and the bounded suites assert how many cases were accepted.Metamorphic relations, where ddx is compared with itself:
∇(cL),∇(L+c),∇(L²),∇sin L, a loss CTE read twice,L·sg(L),L + sg(L);grad=vjpseeded with 1 (and 2.5);vjp(R, c)=grad(Σ R·c);vjpis linear inc;wrtorder, case and splitting, row order, and partitions 1/4/7;runonly the value and gradients remain, afterreleasenothing does, and user tables are untouched; three programs spawned concurrently on one context don't interfere (S11);grad(loss, t.val)equals the program, and an SGD join equalsθ − 0.1·∇L. The statement is spelled withGRAD/Grad (, comments holdinggrad(and multibyte text, a quoted"Loss Fn"CTE,W.VAL, and extra CTEs around the loss;A refusal is always allowed and is tallied by kind. A panic, an
InternalorInvalidPlanerror, a program that is accepted but fails to run, or a wrong number is a failure.Running it. Six bounded suites run under
cargo test(24 seeds each, about 16s in debug).soak_v2_query_adis#[ignore]d. It uses the sameDDX_SOAK_*knobs and log format as ddx-core's soak, and now runs nightly as a second job..github/scripts/report_fuzz_findings.shnow reads each soak's ownREPROline, and no longer counts a top-level log twice.DDX_V2_SEED=<n>withreplay_one_seedprints a case's SQL and tables.DDX_V2_NULL_PCT=0turns off NULL injection, to hunt past the NULL bug; it replays the same case minus its NULLs, which is how every NULL failure below was attributed.What it found
Soaks totalled about 6,500 cases, of which about 84% were accepted. That gave about 17k finite-difference comparisons (fewer than 0.1% screened) and about 100k metamorphic comparisons. Every failure maps to one of three bugs, pinned in
tests/ad_findings.rsas#[ignore = "known bug: …"]repros with hand-worked gradients (CONTRIBUTING's failing-test-first convention;-- --ignoredshows all four failing):SUM(p.val + q.val)withqNULL at row 1, ddx gives ∂/∂p(1) = 1 where the answer is 0. This happens whether the NULL is in a parameter or in constant data, and it accounts for every finite-difference failure (all pass once the case's NULLs are removed).grad_plan(which takes anyLogicalPlan, a DataFrame's included), and then its backward step fails DataFusion's schema check.No failures came from the calculus identities, vjp, wrt handling, the SQL rewriter, dim types, program reuse, concurrency or catalog hygiene.
Also worth knowing, though not a failure: ddx refuses some rankings the generator makes total over the CTE's dims ("a window function whose PARTITION BY and ORDER BY…", "over rows joined to…"). These are conservative refusals, not wrong answers, but they are coverage the rank-select rule could have.
Not covered
The Python
ddxdb.adlayer (it wraps the same Rust calls), Float32 value columns (the finite-difference oracle would need f32 tolerances), andWITH RECURSIVE, which ddx refuses.🤖 Generated with Claude Code
https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu