Skip to content

docs: M4 in the READMEs and the design (M4 4) - #82

Merged
alxmrs merged 6 commits into
mainfrom
m4/docs
Oct 3, 2026
Merged

alxmrs merged 6 commits into
mainfrom
m4/docs

Conversation

@alxmrs

@alxmrs alxmrs commented Sep 23, 2026 •

Copy link
Copy Markdown
Member

Stacked on #81. Docs only; the last PR of the series.

Rewritten after the ontology review on #80: the docs lead with grad(loss, table.column) in SQL, and no example carries a label.

  • README.md: a "Training: gradients of whole queries" section showing an SGD step as a join with grad(loss, w.val), and what backs it (the nn example, 1e-12 agreement with nn.py's hand-written backward pass and with jax.grad). The status section says M3 and M4 are built and ship with the next release, and "next" is now M5.
  • ddx-datafusion's README: ad::sql with an example, sql_all, the programs underneath, the nn example, and the substrait pin added to the version table.
  • CONTRIBUTING.md: ddx-ad is no longer "under construction", and its integration tests live in ddx-datafusion/tests/.
  • design.md: §8 records M4 as built. §9 updates the higher-order question and adds two new questions: query-level jvp, which the review asked for alongside grad and vjp and which isn't built; and whether some recomputed regions should be saved instead, a performance question for the fused-contraction work.
  • docs/spikes/README.md: a note that substrait_ad_marker_spike.py uses ddx_contract_mark from an earlier draft, and stands as evidence that a ddx-claimed function survives the Substrait trip.

🤖 Generated with Claude Code

https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2

@alxmrs
alxmrs force-pushed the m4/docs branch 2 times, most recently from 3d9c008 to 96dbc6a Compare September 24, 2026 04:51
@alxmrs
alxmrs force-pushed the m4/docs branch 2 times, most recently from b5769a7 to ac8a97f Compare September 30, 2026 09:26
@alxmrs

alxmrs commented Sep 30, 2026

Copy link
Copy Markdown
Member Author

Docs: write down what ddx needs from a Substrait producer

The second-engine spike learned each of these the hard way, and none is documented. Suggest a short section in design.md and the ddx-ad crate docs:

Producer behavior Observed What ddx needs
Optimized or not DataFusion (into_optimized_plan) folds constants; DuckDB only with its optimizer on fold literal casts itself (see #71), or say it wants optimized plans
Shared subtrees DuckDB: ReferenceRel and multi-root plans. DataFusion: inlined tree follow ReferenceRel (#73)
Function names DataFusion: bare. DuckDB: compound (sum:fp64), bare for its own (%) already handled by name normalization; say so
Extension declarations DataFusion 54 and 55: u32::MAX (apache/datafusion#11545). DuckDB: real extension_urns emit valid ones (see #72)
Table names some producers drop the schema qualifier how wrt names match, and an error on ambiguity (#74)
Consumer subsets DuckDB's consumer rejects window_function emit the portable shape (#76)

Also, §4.3 currently says the window idiom "round-trips through DuckDB". In the spike only DuckDB's producer side was confirmed (and it rewrote the window into an aggregate and join); its consumer rejects the expression. Worth stating which direction was verified.

@alxmrs
alxmrs force-pushed the m4/docs branch 2 times, most recently from 1bcedf8 to b81f9e2 Compare October 3, 2026 18:49
@alxmrs
alxmrs force-pushed the m4/python branch 2 times, most recently from 85e51bf to b289afc Compare October 3, 2026 19:05
@alxmrs
alxmrs force-pushed the m4/docs branch 2 times, most recently from 844cee5 to dd8eec9 Compare October 3, 2026 19:07
@alxmrs
alxmrs force-pushed the m4/python branch 2 times, most recently from aa903be to fd40adb Compare October 3, 2026 19:19
@alxmrs
alxmrs force-pushed the m4/docs branch 2 times, most recently from 79e11a6 to 40d3d3a Compare October 3, 2026 19:57
@alxmrs
alxmrs force-pushed the m4/docs branch 2 times, most recently from ae2506b to 7f7d446 Compare October 3, 2026 20:27

@alxmrs alxmrs left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Quick note of feedback.

Comment thread README.md Outdated
rankings), and nothing in it is labelled. The backward pass is a sequence of
plain Substrait plans the engine runs.
[`examples/nn`](crates/ddx-datafusion/examples/nn) trains nn.py's MLP this way,
with nn.py's own loss query; its gradients equal nn.py's hand-written backward

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No one but me know what nn.py is.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair point. Fixed in 8d966be. The READMEs (top-level, ddx-datafusion and tests/) and the example's own docs now say what it is: a two-layer MLP over 6×6 images, written entirely in SQL, adapted from a pure-SQL neural network in xarray-sql, whose hand-written backward queries ddx's gradients match to 1e-12. docs/design.md already introduces nn.py with that link at its first mention, so I left it.

🤖 Generated with Claude Code

Base automatically changed from m4/python to main October 3, 2026 20:57
alxmrs and others added 6 commits October 3, 2026 13:57
The top-level README shows training with grad(loss, table.column) in SQL and
updates the status table; ddx-datafusion's README documents ad::sql and the
substrait pin; CONTRIBUTING drops the under-construction note; design.md
records M4 as built, and §9 adds query-level jvp and the save-or-recompute
question.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
…'s own

From the adversarial review: design.md still named the fixed tables
(`__ddx_cotangent`, `__ddx_saved_{n}`) and `ddx_ad::emit::bind_reads`,
and did not say that the SQL surface refuses a loss in a WITH RECURSIVE
clause independently of DataFusion 54's recursive-CTE bug. The M4 entry now records
both, and that nn.py's network is checked against jax.grad.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
…S14)

From the composability review and its DuckDB spike (#82): a table of
producer and consumer behaviors and what ddx does about each; §4.3 now
says which direction of the DuckDB window round trip held (DuckDB's
consumer rejects window_function); §4.6 lists the portable form of the
reduce rules' windows as open; S14 records what moved into ddx-ad.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
A program is built before a runner exists, so whether to emit windows
has to reach grad (composability re-review on #76).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
From the review of #82: nobody but its author knows nn.py. The READMEs
and the example's docs now say what it is, a two-layer MLP over 6x6
images written entirely in SQL, adapted from a pure-SQL demo in
xarray-sql (linked), whose hand-written backward queries ddx's
gradients match.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2

@alxmrs alxmrs left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@alxmrs
alxmrs merged commit 7215892 into main Oct 3, 2026
9 checks passed
@alxmrs
alxmrs deleted the m4/docs branch October 3, 2026 21:22
@github-actions github-actions Bot mentioned this pull request Oct 3, 2026
alxmrs added a commit that referenced this pull request Oct 3, 2026
…M4 5) (#89)

Adversarial property/fuzz testing for the M3/M4 stack (#69–#83): a
generator of loss queries, a finite-difference oracle that shares
nothing with ddx, a set of metamorphic relations, and a nightly soak.
Stacked on #82.

## What it checks

`crates/ddx-datafusion/tests/ad_simulation.rs` generates loss queries as
DAGs of CTEs over small random tables. It uses every primitive in
design.md §4.3 and the shapes around them: maps (including about 10%
drawn from ddx-core's own `gen_expr`), joins on shared dims and cross
joins, LEFT JOIN with COALESCE, CASE on data, grouped SUM/AVG/MAX/MIN,
repeated identical aggregates, HAVING, dim and value filters, IN/NOT
IN/EXISTS semi-joins, rank filters, ORDER BY … LIMIT, UNION ALL,
DISTINCT, softmax with and without `ddx_stop_gradient` on the shift, and
scalar subqueries over data. Tables get NULLs, orphan rows, repeated
keys in data, shuffled rows, and BIGINT/INT/VARCHAR dims.

**Oracle.** The loss is a query DataFusion can run without ddx, so each
gradient is checked against `(L(θ+hd) − L(θ−hd))/2h`. It uses random
directions and single entries. The loss is screened at kinks (the second
difference must shrink like h²) and Richardson-extrapolated. A screened
point is counted, and the bounded suites assert how many cases were
accepted.

**Metamorphic relations, where ddx is compared with itself:**
- the calculus: `∇(cL)`, `∇(L+c)`, `∇(L²)`, `∇sin L`, a loss CTE read
twice, `L·sg(L)`, `L + sg(L)`;
- `grad` = `vjp` seeded with 1 (and 2.5); `vjp(R, c)` = `grad(Σ R·c)`;
`vjp` is linear in `c`;
- invariance to CTE vs inline subqueries, the unoptimized plan, `wrt`
order, case and splitting, row order, and partitions 1/4/7;
- the program contract: a program built at θ₀ and run at θ₁ (sometimes
with extra rows) equals a fresh one; repeated dims are refused by the
checks; after `run` only the value and gradients remain, after `release`
nothing does, and user tables are untouched; three programs spawned
concurrently on one context don't interfere (S11);
- the SQL surface: `grad(loss, t.val)` equals the program, and an SGD
join equals `θ − 0.1·∇L`. The statement is spelled with `GRAD`/`Grad (`,
comments holding `grad(` and multibyte text, a quoted `"Loss Fn"` CTE,
`W.VAL`, and extra CTEs around the loss;
- shape: one gradient row per table row, Float64, NULL exactly where the
value is NULL.

A refusal is always allowed and is tallied by kind. A panic, an
`Internal` or `InvalidPlan` error, a program that is accepted but fails
to run, or a wrong number is a failure.

**Running it.** Six bounded suites run under `cargo test` (24 seeds
each, about 16s in debug). `soak_v2_query_ad` is `#[ignore]`d. It uses
the same `DDX_SOAK_*` knobs and log format as ddx-core's soak, and now
runs nightly as a second job. `.github/scripts/report_fuzz_findings.sh`
now reads each soak's own `REPRO` line, and no longer counts a top-level
log twice. `DDX_V2_SEED=<n>` with `replay_one_seed` prints a case's SQL
and tables. `DDX_V2_NULL_PCT=0` turns off NULL injection, to hunt past
the NULL bug; it replays the same case minus its NULLs, which is how
every NULL failure below was attributed.

## What it found

Soaks totalled about 6,500 cases, of which about 84% were accepted. That
gave about 17k finite-difference comparisons (fewer than 0.1% screened)
and about 100k metamorphic comparisons. Every failure maps to one of
three bugs, pinned in `tests/ad_findings.rs` as `#[ignore = "known bug:
…"]` repros with hand-worked gradients (CONTRIBUTING's
failing-test-first convention; `-- --ignored` shows all four failing):

1. **Silently wrong gradient at NULL rows.** SUM, AVG, MAX and MIN skip
a row whose argument is NULL, but ddx still sends the group's cotangent
through that row's other inputs. In `SUM(p.val + q.val)` with `q` NULL
at row 1, ddx gives ∂/∂p(1) = 1 where the answer is 0. This happens
whether the NULL is in a parameter or in constant data, and it accounts
for every finite-difference failure (all pass once the case's NULLs are
removed).
2. **Two rank filters in one recomputed region are accepted, then fail
to run.** The optimizer gives both window columns the same name, and the
rebuilt region keeps both.
3. **An unoptimized plan with a constant join condition under a CASE**
is accepted by `grad_plan` (which takes any `LogicalPlan`, a DataFrame's
included), and then its backward step fails DataFusion's schema check.

No failures came from the calculus identities, vjp, wrt handling, the
SQL rewriter, dim types, program reuse, concurrency or catalog hygiene.

Also worth knowing, though not a failure: ddx refuses some rankings the
generator makes total over the CTE's dims ("a window function whose
PARTITION BY and ORDER BY…", "over rows joined to…"). These are
conservative refusals, not wrong answers, but they are coverage the
rank-select rule could have.

## Not covered

The Python `ddxdb.ad` layer (it wraps the same Rust calls), Float32
value columns (the finite-difference oracle would need f32 tolerances),
and `WITH RECURSIVE`, which ddx refuses.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant