Conversation
3d9c008 to
96dbc6a
Compare
b5769a7 to
ac8a97f
Compare
|
Docs: write down what ddx needs from a Substrait producer The second-engine spike learned each of these the hard way, and none is documented. Suggest a short section in
Also, §4.3 currently says the window idiom "round-trips through DuckDB". In the spike only DuckDB's producer side was confirmed (and it rewrote the window into an aggregate and join); its consumer rejects the expression. Worth stating which direction was verified. |
1bcedf8 to
b81f9e2
Compare
85e51bf to
b289afc
Compare
844cee5 to
dd8eec9
Compare
aa903be to
fd40adb
Compare
79e11a6 to
40d3d3a
Compare
ae2506b to
7f7d446
Compare
alxmrs
left a comment
There was a problem hiding this comment.
Quick note of feedback.
| rankings), and nothing in it is labelled. The backward pass is a sequence of | ||
| plain Substrait plans the engine runs. | ||
| [`examples/nn`](crates/ddx-datafusion/examples/nn) trains nn.py's MLP this way, | ||
| with nn.py's own loss query; its gradients equal nn.py's hand-written backward |
There was a problem hiding this comment.
No one but me know what nn.py is.
There was a problem hiding this comment.
Fair point. Fixed in 8d966be. The READMEs (top-level, ddx-datafusion and tests/) and the example's own docs now say what it is: a two-layer MLP over 6×6 images, written entirely in SQL, adapted from a pure-SQL neural network in xarray-sql, whose hand-written backward queries ddx's gradients match to 1e-12. docs/design.md already introduces nn.py with that link at its first mention, so I left it.
🤖 Generated with Claude Code
The top-level README shows training with grad(loss, table.column) in SQL and updates the status table; ddx-datafusion's README documents ad::sql and the substrait pin; CONTRIBUTING drops the under-construction note; design.md records M4 as built, and §9 adds query-level jvp and the save-or-recompute question. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
…'s own
From the adversarial review: design.md still named the fixed tables
(`__ddx_cotangent`, `__ddx_saved_{n}`) and `ddx_ad::emit::bind_reads`,
and did not say that the SQL surface refuses a loss in a WITH RECURSIVE
clause independently of DataFusion 54's recursive-CTE bug. The M4 entry now records
both, and that nn.py's network is checked against jax.grad.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
…S14) From the composability review and its DuckDB spike (#82): a table of producer and consumer behaviors and what ddx does about each; §4.3 now says which direction of the DuckDB window round trip held (DuckDB's consumer rejects window_function); §4.6 lists the portable form of the reduce rules' windows as open; S14 records what moved into ddx-ad. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
A program is built before a runner exists, so whether to emit windows has to reach grad (composability re-review on #76). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
From the review of #82: nobody but its author knows nn.py. The READMEs and the example's docs now say what it is, a two-layer MLP over 6x6 images written entirely in SQL, adapted from a pure-SQL demo in xarray-sql (linked), whose hand-written backward queries ddx's gradients match. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
…M4 5) (#89) Adversarial property/fuzz testing for the M3/M4 stack (#69–#83): a generator of loss queries, a finite-difference oracle that shares nothing with ddx, a set of metamorphic relations, and a nightly soak. Stacked on #82. ## What it checks `crates/ddx-datafusion/tests/ad_simulation.rs` generates loss queries as DAGs of CTEs over small random tables. It uses every primitive in design.md §4.3 and the shapes around them: maps (including about 10% drawn from ddx-core's own `gen_expr`), joins on shared dims and cross joins, LEFT JOIN with COALESCE, CASE on data, grouped SUM/AVG/MAX/MIN, repeated identical aggregates, HAVING, dim and value filters, IN/NOT IN/EXISTS semi-joins, rank filters, ORDER BY … LIMIT, UNION ALL, DISTINCT, softmax with and without `ddx_stop_gradient` on the shift, and scalar subqueries over data. Tables get NULLs, orphan rows, repeated keys in data, shuffled rows, and BIGINT/INT/VARCHAR dims. **Oracle.** The loss is a query DataFusion can run without ddx, so each gradient is checked against `(L(θ+hd) − L(θ−hd))/2h`. It uses random directions and single entries. The loss is screened at kinks (the second difference must shrink like h²) and Richardson-extrapolated. A screened point is counted, and the bounded suites assert how many cases were accepted. **Metamorphic relations, where ddx is compared with itself:** - the calculus: `∇(cL)`, `∇(L+c)`, `∇(L²)`, `∇sin L`, a loss CTE read twice, `L·sg(L)`, `L + sg(L)`; - `grad` = `vjp` seeded with 1 (and 2.5); `vjp(R, c)` = `grad(Σ R·c)`; `vjp` is linear in `c`; - invariance to CTE vs inline subqueries, the unoptimized plan, `wrt` order, case and splitting, row order, and partitions 1/4/7; - the program contract: a program built at θ₀ and run at θ₁ (sometimes with extra rows) equals a fresh one; repeated dims are refused by the checks; after `run` only the value and gradients remain, after `release` nothing does, and user tables are untouched; three programs spawned concurrently on one context don't interfere (S11); - the SQL surface: `grad(loss, t.val)` equals the program, and an SGD join equals `θ − 0.1·∇L`. The statement is spelled with `GRAD`/`Grad (`, comments holding `grad(` and multibyte text, a quoted `"Loss Fn"` CTE, `W.VAL`, and extra CTEs around the loss; - shape: one gradient row per table row, Float64, NULL exactly where the value is NULL. A refusal is always allowed and is tallied by kind. A panic, an `Internal` or `InvalidPlan` error, a program that is accepted but fails to run, or a wrong number is a failure. **Running it.** Six bounded suites run under `cargo test` (24 seeds each, about 16s in debug). `soak_v2_query_ad` is `#[ignore]`d. It uses the same `DDX_SOAK_*` knobs and log format as ddx-core's soak, and now runs nightly as a second job. `.github/scripts/report_fuzz_findings.sh` now reads each soak's own `REPRO` line, and no longer counts a top-level log twice. `DDX_V2_SEED=<n>` with `replay_one_seed` prints a case's SQL and tables. `DDX_V2_NULL_PCT=0` turns off NULL injection, to hunt past the NULL bug; it replays the same case minus its NULLs, which is how every NULL failure below was attributed. ## What it found Soaks totalled about 6,500 cases, of which about 84% were accepted. That gave about 17k finite-difference comparisons (fewer than 0.1% screened) and about 100k metamorphic comparisons. Every failure maps to one of three bugs, pinned in `tests/ad_findings.rs` as `#[ignore = "known bug: …"]` repros with hand-worked gradients (CONTRIBUTING's failing-test-first convention; `-- --ignored` shows all four failing): 1. **Silently wrong gradient at NULL rows.** SUM, AVG, MAX and MIN skip a row whose argument is NULL, but ddx still sends the group's cotangent through that row's other inputs. In `SUM(p.val + q.val)` with `q` NULL at row 1, ddx gives ∂/∂p(1) = 1 where the answer is 0. This happens whether the NULL is in a parameter or in constant data, and it accounts for every finite-difference failure (all pass once the case's NULLs are removed). 2. **Two rank filters in one recomputed region are accepted, then fail to run.** The optimizer gives both window columns the same name, and the rebuilt region keeps both. 3. **An unoptimized plan with a constant join condition under a CASE** is accepted by `grad_plan` (which takes any `LogicalPlan`, a DataFrame's included), and then its backward step fails DataFusion's schema check. No failures came from the calculus identities, vjp, wrt handling, the SQL rewriter, dim types, program reuse, concurrency or catalog hygiene. Also worth knowing, though not a failure: ddx refuses some rankings the generator makes total over the CTE's dims ("a window function whose PARTITION BY and ORDER BY…", "over rows joined to…"). These are conservative refusals, not wrong answers, but they are coverage the rank-select rule could have. ## Not covered The Python `ddxdb.ad` layer (it wraps the same Rust calls), Float32 value columns (the finite-difference oracle would need f32 tolerances), and `WITH RECURSIVE`, which ddx refuses. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Stacked on #81. Docs only; the last PR of the series.
grad(loss, w.val), and what backs it (the nn example, 1e-12 agreement with nn.py's hand-written backward pass and withjax.grad). The status section says M3 and M4 are built and ship with the next release, and "next" is now M5.ad::sqlwith an example,sql_all, the programs underneath, the nn example, and thesubstraitpin added to the version table.ddx-adis no longer "under construction", and its integration tests live inddx-datafusion/tests/.jvp, which the review asked for alongsidegradandvjpand which isn't built; and whether some recomputed regions should be saved instead, a performance question for the fused-contraction work.substrait_ad_marker_spike.pyusesddx_contract_markfrom an earlier draft, and stands as evidence that a ddx-claimed function survives the Substrait trip.🤖 Generated with Claude Code
https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2