Skip to content

test: a forward-mode oracle, a cost bound and a SQL-text fuzz; what they found (M4 10) - #98

Merged
alxmrs merged 3 commits into
mainfrom
m4/v2-round2
Oct 4, 2026
Merged

alxmrs merged 3 commits into
mainfrom
m4/v2-round2

Conversation

@alxmrs

@alxmrs alxmrs commented Sep 30, 2026

Copy link
Copy Markdown
Member

Stacked on #93. This round tests the fixed stack: every soak and every earlier finding passes on it. The goal was to find what the existing oracles can't see.

New techniques

  • Forward-mode oracle (exact). Every generated relation gets a hand-written twin: the same SQL with a tangent column dv beside each value. It uses JAX's rule for MAX/MIN (the mean of the tangents of the rows that attain the extreme exactly) and ddx's documented conventions at kinks. DataFusion runs it, and reverse mode's ⟨∇L, d⟩ must match it to 1e-8. It shares no code with the transposes, so it's exact to rounding rather than to a finite difference's step. That lets it see conventions at points a step would cross, and errors a step would hide.
  • Cost bound (cost). The plans DataFusion builds for a program's steps must stay within 200× the forward query's. A superlinear blowup then fails as a finding instead of killing the process. A new fan node (one relation read by several columns) feeds it.
  • SQL-text fuzz (tests/ad_sql_text.rs). Valid grad(…) statements with whitespace, CRLF, comments (some holding (, ), quotes or multibyte text), and changes of case and quoting between tokens. Each spelling must give the plain statement's rows or fail loudly.
  • Targeted probes past the generator: function semantics (log, atan2, cbrt, round, …), filters pushed into a custom provider's scan (Exact pushdown), window aggregates as values, DISTINCT aggregates, GROUP BY expressions, CASE forms, decimals, computed and inequality join keys, deep and wide plans, and vjp cotangent edge cases.

Found

Each is pinned in ad_findings.rs as an ignored known bug:

finding severity notes
Fan-in is exponential. A value read by N columns gets N cotangent terms, folded to skip NULLs (80ca10f). Each fold names the accumulator three times, and DataFusion's Substrait consumer names a column by its expression. high 10 columns: 117 MB of plan; 12 columns: over 4 GB, on a two-row table. Introduced by a NULL fix.
Deep row-wise chains are exponential. Maps between two aggregates are rebuilt as unnamed appended columns, and names double per layer where a map reads its input twice. high 14 layers: 61 MB; 20 layers exhaust 13 GB. GELU/LayerNorm-style layers have this shape.
Stack overflow in ddx_ad::grad. Transposer::region clones a rebuilt region recursively, hundreds of relations deep. high On a 2 MB stack (a tokio worker's), a 66 KB plan aborts the whole process; nothing can catch it.
Panics and internal errors in grad(…) in SQL. call_span counts parentheses inside comments. medium Of 3,000 valid spellings: 953 raised ddx's Internal error, 18 panicked in GradCalls::rewrite, and none returned wrong rows. grad /* c */ (…) isn't recognised either.
MAX/MIN near-ties. The 8-ulp window added against jitter also shares the cotangent between exact values a few ulps apart, where MAX is differentiable. low MAX(1, 1+2ulp) gives (0.5, 0.5); jax.grad gives (0, 1). Invisible to any finite difference.
vjp adds up repeated cotangent keys. low wrt tables get a uniqueness check; cotangents get none, and a row with a NULL key is dropped silently.
A simple CASE x WHEN … is reported as an invalid plan, which the harness treats as ddx misreading its producer. low

Not bugs, for the record: filters inside a scan are honoured; unsupported functions are refused rather than mis-derived; NULL-skipping functions (greatest/least with NULL, COALESCE, NVL) match the finite difference. Coverage gaps (refusals of queries ddx could support): LIMIT … OFFSET when the ORDER BY dim isn't selected, and power(v, e) with e built from dims.

Running it

The bounded exact suite runs in CI. cost and the text fuzz are ignored until their bugs are fixed; the soak runs cost every night. A soak can now use a machine's memory, so CONTRIBUTING.md says to run it under systemd-run --user --scope -p MemoryMax=8G. Big mode and the fan node are sized to stay within memory meanwhile.

🤖 Generated with Claude Code

https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu

alxmrs added a commit that referenced this pull request Sep 30, 2026
From the v2 soak's round two (#98): `CASE i WHEN 0 THEN … END` was
refused as if the plan were malformed. DataFusion's producer writes it as
an IfThen whose first clause is the operand with no result (its consumer
reads it back so), and a Substrait switch is the standard form; the map
rule read neither. Both are now a piecewise hole like a CASE: the CASE
whose conditions are `x = v`, spelled out with `equal` when its partials
are emitted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 30, 2026
… keys

From the v2 soak's round two (#98):

- Fan-in was exponential. A value read by N columns gets N cotangent
  terms, which the NULL-skipping fold summed by nesting the running sum,
  naming it three times a step. DataFusion's Substrait consumer names a
  column by its expression, so ten readers made a 117 MB plan and twelve
  did not fit in 4 GB. The terms are now one flat sum: NULL only if every
  term is, each term in it a fixed number of times.
- Deep plans overflowed the stack. The backward step appended one
  projection per cotangent column, so a wide region (a CTE read twice per
  layer, nine layers) nested hundreds of relations deep, and cloning it on
  a 2 MB stack (a tokio worker's) aborted the process. Cotangents are now
  projected in batches, a new one only where a term reads a column still
  in the current batch, so the step is as deep as the chain of columns
  reading each other, not as the region is wide.
- vjp joined a cotangent whose keys repeat as it was, doubling that row's
  gradient. Its program now carries a check that the cotangent has one row
  per key, like a wrt table's dims check; the check reads the cotangent
  table unbound, so the harness binds check reads as it binds a step's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 30, 2026
From the v2 soak's round two (#98): the 8-ulp attainment window, added so
a MAX/MIN tie that rounding breaks is shared the same way every run, also
shared between exact values a few ulps apart that do not tie: MAX(1,
1 + 2 ulps) gave (0.5, 0.5), where the function is differentiable with
gradient (0, 1), as jax.grad gives. Only an argument that can differ in its
last bits between recomputations now gets the tolerance: one that reads a
saved aggregate or constant data (which can hold one). Table values, and
elementwise functions of them, are the same every run and compare exactly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 30, 2026
…k reads

From the v2 soak's round two (#98): DataFusion's Substrait consumer names
a computed column by its whole expression, and ddx's plans compute each
column from earlier ones, so a map layer that reads its input twice
(sin(v) + 0.1 * v, a GELU- or LayerNorm-style layer) doubled every name
after it: 60 MB of plan at 14 layers, and 20 layers exhausted 13 GB. The
names mean nothing to ddx, whose plans refer to columns by position. The
new ad::logical_plan binds a plan's reads and consumes it with a consumer
that names each projection's computed columns `__ddx_c{n}`; run_step and
run_checks use it, so a check's reads are now bound too (a vjp's
cotangent-keys check reads the cotangent table, whose types ddx does not
know).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 30, 2026
From the v2 soak's round two (#98, a SQL-text fuzz): a call's end was
found by counting parentheses in the text, skipping quotes but not
comments. `/* ) */` cut the call short and `/* ( */` never closed it (an
Internal error, 953 of 3,000 valid spellings), and a `(` in one call's
comment ran its span into the next call's, so rewriting sliced the
statement backwards and panicked (18). A comment between `grad` and `(`
also hid the call. GradCalls::find now tokenizes the statement with the
dialect's tokenizer, finds `grad` then `(` across whitespace and
comments, and counts parenthesis tokens; overlapping spans are refused
rather than sliced.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 30, 2026
A vjp program's checks now include one that the cotangent's keys do not
repeat (#98), and it reads the cotangent table, whose types ddx does not
know. ddxdb.ad binds a check's reads as it binds a step's, through one
_consume helper. The module docs note that datafusion-python's Substrait
consumer names computed columns by their expressions, so a deep chain of
row-wise maps makes a large plan from Python, which the Rust adapter
avoids.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
@alxmrs

alxmrs commented Sep 30, 2026

Copy link
Copy Markdown
Member Author

Thanks: all seven were real, and they're fixed in the stack. This branch is rebased onto it. The findings run as ordinary tests, and cost and the SQL-text fuzz are back in the PR gate.

Finding Fix
Fan-in exponential #74: a column's cotangent terms are one flat NULL-skipping sum, with no nested running sum
Deep chains exponential #79: ad::logical_plan consumes a step with short computed-column names (__ddx_c{n}). The names are DataFusion's, and ddx's plans refer to columns by position. 20 layers now plan small, and the gradient matches the chain rule to 1e-12.
Stack overflow #74: cotangents are projected in batches, so a step is as deep as the chain of columns reading each other, not as the region is wide. Your 9-layer repro runs on a 2 MB stack.
Comments in grad(…) #83: GradCalls::find tokenizes with the dialect's tokenizer. It finds grad, then (, across comments, counts parenthesis tokens, and refuses overlapping spans rather than slicing
MAX/MIN near-ties #76: the tie tolerance applies only to arguments that can jitter (ones that read a saved aggregate or constant data). Table values compare exactly, so MAX(1, 1+2ulp) gives (0, 1)
vjp repeated cotangent keys #74: a check that the cotangent has one row per key, like a wrt table's dims check. The adapters (#79, #81) now bind a check's reads.
Simple CASE x WHEN … #71: DataFusion writes it as an IfThen whose first clause is the operand with no result, and a Substrait switch is the standard form. Both are now read as the CASE on x = v.

Changes to your harness, in my commit on this branch:

  • consumed_plan_bytes and cost_checks measure plans through ad::logical_plan, which is what ad::run builds, not DataFusion's default consumer.
  • The stack-overflow repro's child process no longer passes --ignored; now that the test is un-ignored, that flag would filter it out.

One limit remains: datafusion-python's Substrait consumer can't be given short names, so a very deep chain still makes a large plan from Python. It's documented in ddxdb.ad and design S13.

Verified here:

  • All workspace tests pass, including the exact and cost bounded suites.
  • ddxdb (68) and the JAX oracle (137) pass.
  • Capped soaks found nothing: 1,143 general cases from base 20,000,000 and 540 extreme-value cases from 21,000,000.

🤖 Generated with Claude Code

@alxmrs
alxmrs force-pushed the m4/v2-round2 branch 2 times, most recently from 3587c9d to a8ff8d8 Compare October 3, 2026 20:27
@alxmrs
alxmrs force-pushed the m4/v2-mutation branch 2 times, most recently from 0cb8a8b to 59222bd Compare October 3, 2026 20:57

@alxmrs alxmrs left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@alxmrs
alxmrs force-pushed the m4/v2-round2 branch 2 times, most recently from 339a3e9 to 920328e Compare October 3, 2026 22:16
Base automatically changed from m4/v2-mutation to main October 3, 2026 22:27
alxmrs and others added 2 commits October 3, 2026 15:27
…hey found

Round two, against the fixed stack.

- exact: every generated relation gets a forward-mode twin (a tangent
  column beside each value; rules written by hand, JAX's for MAX and MIN
  and ddx's documented ones at a kink), and reverse mode's ⟨∇L, d⟩ must
  equal the twin's directional derivative to 1e-8. It shares no code with
  the transposes, and alone kills every value mutant, several faster than
  the finite difference. 11,000+ comparisons on the fixed stack: none
  disagree.
- cost: the plans DataFusion builds for a program's steps must stay within
  200x of the forward query's, so a blowup is a finding, not an OOM kill.
  A `fan` node (one relation read by several columns) feeds it.
- ad_sql_text.rs fuzzes grad(…) in SQL at the text level.

Found, pinned in ad_findings.rs as ignored known bugs:

- MAX/MIN's 8-ulp attainment window shares the cotangent between exact
  values a few ulps apart, where MAX is differentiable: (0.5, 0.5), not
  (0, 1). A finite difference cannot see it.
- Fan-in's NULL-skipping fold (80ca10f) names its accumulator three times,
  and DataFusion names a column by its expression: a value read by N
  columns plans to 3^N. 10 columns: 117 MB of plan; 12: over 4 GB, on a
  two-row table.
- A chain of row-wise maps doubles the same names per layer: 14 layers,
  61 MB; 20 layers exhaust 13 GB.
- ddx_ad::grad clones a rebuilt region recursively, hundreds of relations
  deep: on a 2 MB stack (a tokio worker's) a 66 KB plan aborts the process.
- grad(…) in SQL counts parentheses inside comments: 953 of 3000 valid
  spellings raise ddx's Internal error and 18 panic in GradCalls::rewrite.
  None returned wrong rows.
- vjp adds a cotangent's repeated keys; a simple CASE is read as an
  invalid plan.

The soak's big-mode and fan-node settings keep a case within memory; run
soaks under a memory cap (CONTRIBUTING.md).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu
All seven now pass on the stack below: the MAX/MIN near-tie (#76); fan-in,
the stack overflow and vjp's cotangent keys (#74); the deep chain (#79);
comments in grad(…) (#83); the simple CASE (#71). So:

- ad_findings.rs runs them as ordinary tests; the stack-overflow repro's
  child process no longer passes --ignored, which would now filter it out.
- The plan-size measures (consumed_plan_bytes, cost_checks) consume a step
  as ad::run does, through ad::logical_plan, which names computed columns
  briefly; DataFusion's default consumer names them by their expressions.
- The cost bound and the SQL-text fuzz run in the PR gate again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
Rust 1.88 inferred pick's element type from the expected &str argument
(T = str) rather than from the slice; a turbofish pins it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
@alxmrs
alxmrs merged commit 9b4c208 into main Oct 4, 2026
9 checks passed
@alxmrs
alxmrs deleted the m4/v2-round2 branch October 4, 2026 00:26
@github-actions github-actions Bot mentioned this pull request Oct 3, 2026
alxmrs added a commit that referenced this pull request Oct 4, 2026
…gs (M4 11) (#102)

Stacked on #98. A soak aimed at the three changes that landed today (#74
batched cotangent projections, #76 the "can this value jitter?" rule for
ties, #79 the ShortNames consumer), and a benchmark of how a program's
cost scales.

## The soak

Three regions, three hours each, run in parallel under a memory cap.
About 55,800 cases in total.

| region | aimed at | cases | exact comparisons | failures |
|---|---|---|---|---|
| general (all groups and modes) | #79, everything | 19,116 | 44,115 |
1: the anti-join filter bug below |
| ties (`ulps` mode, every table differentiated) | #76 | 18,795 | 45,432
| 59 seeds, all the constant-data near-tie bug below |
| scale (big mode: thousands of rows, fan-in, cost bound) | #74, #79 at
scale | 17,867 | 43,590 | 0 in ddx (3 seeds: two harness limits, one
upstream NaN bug) |

**ShortNames and the batched cotangents produced no failure** in about
37,000 cases and 88,000 exact forward-vs-reverse comparisons.

Harness additions:
- The forward-mode oracle now mirrors #76's rule. Each relation records
whether its value can jitter (it reads an aggregate's output); MAX/MIN
tie tests use the 8-ulp window only there. A disagreement is a
`[tie-rule]` failure, and agreeing with exact equality (jax.grad's
convention) always passes.
- A new `ulps` data mode puts parameters a few ulps apart: near-ties
that aren't ties. It's soak-only while the finding below is open.
`DDX_V2_WRT_ALL` differentiates every parameter table a loss reads.
- A twin whose loss varies run to run is tallied as ill-conditioned
(`sin` of 3×10⁸⁰ is decided by a parallel sum's last bits), and
`DDX_V2_DEBUG` compares every twin with its relation.

## Found

Pinned in `ad_findings.rs` as ignored known bugs:

1. **#76's rule is incomplete.** It gives "constant data" the tolerance,
but a table outside `wrt` is exactly as repeatable as a `wrt` table.
`MAX(p·d)` over products 2 ulps apart still gives `(0.5, 0.5)` instead
of `(0, 1)`. As a result **a table's gradient depends on which other
tables are differentiated**, which the ties region caught 59 times
through its metamorphic checks. The rule should key on whether a value
reads an aggregate's output, not on whether it reads a differentiated
table.
2. **A filter above an anti-join is lost when the region is
recomputed.** With `push_down_filter` off, gradient reaches a row the
query excludes (`NOT IN` then `j <> 2`) while the loss is unchanged.
DataFusion's default rules push such filters below the join, so this
needs a non-default configuration or another engine's plans.
3. **Upstream:** in one query, DataFusion's grouped `MAX` skips NaN for
one group (giving 5.0) and returns NaN for another, depending on the
order partial aggregates merge. With NaN in the data, the same loss
comes out finite on some runs and NaN on others, and ddx's gradients
inherit that.

## Cost

`tests/ad_perf.rs` (ignored; run one family per process under a memory
cap). These timings are from a quiet machine; "build + run" is building
the program plus running it, as a multiple of the forward query:

| family | build + run ÷ forward | scaling |
|---|---|---|
| nn.py's MLP, 64 → 4,096 samples | 2.8–2.9× | linear in data; building
the program takes a constant 14 ms |
| matrix product w.r.t. both operands, 1k → 50k rows | 8× → 13× |
linear, large constant (the large input's dense gradient) |
| layers (a map and a grouped sum each), 2 → 16 | 5.7× → 4.1× | linear
in depth across aggregates |
| residual blocks, 1 → 6 | 4.2× → 1.8× | better than the forward pass,
which DataFusion recomputes per CTE read |
| fan-in, 2 → 64 columns | 9× → 15× | near-linear (DataFusion's plans:
26 KB at 64) |
| **depth within one region, 5 → 80 maps** | 6× → 16× (40) → **63×
(80)** | DataFusion's plans grow about quadratically |
| **a CTE reused inside a region, 2 → 6** | 8× → 54× → **236×** | about
the cube of the forward pass (ddx's plans grow 4× per level to the
forward's 2×) |

The last two are the performance limits worth deciding on. A realistic
block between two matrix products (GELU, LayerNorm, a residual) is 10–20
operations, at about 6–7×. Reuse without an aggregate in between is
uncommon; residual networks put a matrix product there.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant