Skip to content

test: grad in SQL under awkward names, and SGD loops in SQL (M4 8) - #92

Merged
alxmrs merged 8 commits into
mainfrom
m4/v2-adapter-surface
Oct 3, 2026
Merged

alxmrs merged 8 commits into
mainfrom
m4/v2-adapter-surface

Conversation

@alxmrs

@alxmrs alxmrs commented Sep 29, 2026

Copy link
Copy Markdown
Member

Stacked on #91. This PR tests the adapter's SQL surface (ad::sql, ad::sql_all) the way users will reach it.

What it adds

  • names: grad(loss, t.col) must equal the program's gradient in each of these setups:
    • the table read as s1.t or "T t", while a decoy under its old name holds different values, so reading the wrong table can't pass by coincidence;
    • a CTE shadowing the table's name;
    • the call inside a subquery or a later CTE;
    • two losses in one statement (grad(2L) − 2·grad(L) = 0);
    • user tables named like ddx's steps (value, saved_0, grad_0_w, …), which must come through untouched.
  • train: three SGD steps written in SQL, one statement per table through sql_all. Each must move θ to exactly θ − lr·∇L as a fresh ad::grad program computes it. Some step along −∇L must lower a smooth loss (kinked losses are exempt). The context must be clean after every step.
  • Stricter oracle: a NaN or infinite gradient where the loss is finite and smooth along the direction is now a failure. It used to be skipped.
  • A refusal only the SQL surface makes (it plans SELECT * FROM loss, a slightly different plan from the program API's) is tallied as coverage, not failed.

What it found

About 4,900 cases:

  • A table whose name has a capital has no gradient in SQL. Pinned in ad_findings.rs. The gradient step is named after the table (…_grad_0_W). ctx.register_table folds that unquoted name to lower case, and ad::sql then reads it back quoted, so it isn't found. Any "Weights"-style table name hits this. Only the soak tries this variant until it's fixed, so the PR gate stays green.
  • The MAX recomputation bug from test: ties, NULLs, extreme values and big tables in the v2 soak (M4 7) #91 also produces NaN and inf, not just zeros: the even split divides by zero attaining rows. The stricter check catches these now.
  • Minor, noted rather than pinned: with NaN or infinite data, ddx-core's quotient-rule expression evaluates ∞/∞ = NaN where the partial's limit is 0. That's a v1 expression-form limit only infinite inputs reach, and those cases are exempt from the NaN check.

Nothing wrong in schema-qualified names, shadowing CTEs, calls in subqueries or later CTEs, two losses per statement, look-alike user tables, or the SGD loop itself.

Next and last in the stack: mutation testing of ddx-ad, to measure how fast this harness catches seeded bugs.

🤖 Generated with Claude Code

https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu

alxmrs added a commit that referenced this pull request Sep 29, 2026
…o coalesce

From the v2 soak (#89, #90, #92):

- SUM skips a row whose argument is NULL, but the reduce rule still gave
  that row the group's cotangent, so its other inputs got gradient: the p
  of SUM(p + q) with q NULL, or SUM(d - p) with a NULL in constant data.
  The seed is now NULL where the argument is.
- A gradient step was named after its table case and all (`…_grad_0_W`);
  DataFusion folds the unquoted name it is registered under, so it could
  not be found again. Step names are lower case.
- The gradient step used coalesce, which DataFusion 54 runs only after
  SimplifyExpressions has rewritten it; it is now a CASE, so a program runs
  on a context without that rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 29, 2026
From the v2 soak (#92). Fixed by #74's lower-case step names; this pins it
at the SQL surface, which reads the step back by name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 29, 2026
From the v2 soak (#89, #92), fixed in ddx-ad (#74): pinned in Python too,
since ddxdb.ad reads gradient steps back by name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch from 76694d0 to 941fcf9 Compare September 29, 2026 17:26
@alxmrs

alxmrs commented Sep 29, 2026

Copy link
Copy Markdown
Member Author

The capitals finding is fixed (#74: gradient step names are lower case). It's un-ignored, and the names group now tries quoted capitals in the PR gate too, not only in the soak.

The ∞/∞ note turned out to be fixable, so the NaN/∞ exemption is narrower now. ddx-core's quotient rule is (du − (u/v)·dv)/v in #96, and a negative power is written as a division in #97. Infinite data is compared again.

Soaks of the fixed stack then showed where a NaN or ∞ gradient is reverse mode's own answer (JAX gives the same), and those points are now screened:

  • the loss doesn't move beyond rounding along the direction: sqrt(0) or tanh(b·∞) gives 0 · ∞ (seeds 811, 500192);
  • NaN over infinite data (600421);
  • a derivative beyond √f64::MAX (2000728);
  • tiny values, where dividing by a parameter near 1e-160 overflows the chain rule (2400607).

Other changes on this branch:

  • Descent check: now counts the head's MAX/MIN as a kink (1900692).
  • Names check: allows 1e-6 relative error in huge mode, where sin(SUM(v)) turns an ulp of a sum near 1e9 into about 1e-7 (1900495).
  • New pin: ad_findings.rs pins power(v, 0.5) at v = 0, which failed a program ddx had accepted until fix(ddx-core): write a negative power as a division #97.

🤖 Generated with Claude Code

alxmrs added a commit that referenced this pull request Sep 30, 2026
…o coalesce

From the v2 soak (#89, #90, #92):

- SUM skips a row whose argument is NULL, but the reduce rule still gave
  that row the group's cotangent, so its other inputs got gradient: the p
  of SUM(p + q) with q NULL, or SUM(d - p) with a NULL in constant data.
  The seed is now NULL where the argument is.
- A gradient step was named after its table case and all (`…_grad_0_W`);
  DataFusion folds the unquoted name it is registered under, so it could
  not be found again. Step names are lower case.
- The gradient step used coalesce, which DataFusion 54 runs only after
  SimplifyExpressions has rewritten it; it is now a CASE, so a program runs
  on a context without that rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 30, 2026
From the v2 soak (#92). Fixed by #74's lower-case step names; this pins it
at the SQL surface, which reads the step back by name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
alxmrs added a commit that referenced this pull request Sep 30, 2026
From the v2 soak (#89, #92), fixed in ddx-ad (#74): pinned in Python too,
since ddxdb.ad reads gradient steps back by name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch from 63e39f1 to 18e3aff Compare September 30, 2026 02:47
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch 2 times, most recently from 15bdf2d to 441764b Compare September 30, 2026 21:58
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch 2 times, most recently from 45aff4b to a87696b Compare October 2, 2026 22:51
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch 2 times, most recently from 3d17ec6 to d29c37b Compare October 3, 2026 00:04
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch 2 times, most recently from fe910eb to 7e54771 Compare October 3, 2026 01:55
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch from d29c37b to 3d908cf Compare October 3, 2026 01:55
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch from 3d908cf to c9fafb2 Compare October 3, 2026 02:16
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from 3f2d42e to 6b6bea4 Compare October 3, 2026 02:22
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch 2 times, most recently from 5ad051e to ff6acd7 Compare October 3, 2026 02:24
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from 6b6bea4 to 5c74898 Compare October 3, 2026 02:24
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from 5dc960b to 0c4d9fe Compare October 3, 2026 20:20
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch from 54cac76 to 9f195a5 Compare October 3, 2026 20:27
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from 0c4d9fe to 3159d21 Compare October 3, 2026 20:27
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch from 9f195a5 to ce9c611 Compare October 3, 2026 20:38
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch 2 times, most recently from 8b33ade to 94e28d3 Compare October 3, 2026 20:40
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch from ce9c611 to 404dc83 Compare October 3, 2026 20:40
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from 94e28d3 to fa044fb Compare October 3, 2026 20:57
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch from 404dc83 to bd34d2b Compare October 3, 2026 20:57

@alxmrs alxmrs left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from fa044fb to b7781f4 Compare October 3, 2026 21:22
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch from bd34d2b to 15c096a Compare October 3, 2026 21:22
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from b7781f4 to 41c8b3b Compare October 3, 2026 21:34
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch from 15c096a to 19de433 Compare October 3, 2026 21:34
@alxmrs
alxmrs force-pushed the m4/v2-data-shapes branch from 41c8b3b to c4ab48b Compare October 3, 2026 21:50
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch from 19de433 to 7bb05c2 Compare October 3, 2026 21:50
Base automatically changed from m4/v2-data-shapes to main October 3, 2026 22:02
alxmrs and others added 8 commits October 3, 2026 15:02
…in SQL

Two property groups for the adapter's SQL surface:

- names: grad(loss, t.col) must equal the program's gradient with the
  table read as s1.t or "T t" while a decoy under its old name holds other
  values, under a CTE shadowing it, with the call in a subquery or a later
  CTE, and with two losses in one statement (grad(2L) − 2·grad(L) = 0);
  user tables named like ddx's steps must come through untouched.
- train: three SGD steps written in SQL, one statement per table through
  sql_all, must each move θ to θ − lr·∇L as a fresh program computes it,
  some step along −∇L must lower a smooth loss, and nothing may be left
  on the context.

The finite difference now also fails a NaN or infinite gradient where the
loss is finite and smooth; it used to skip them. A refusal only the SQL
surface makes (it plans `SELECT * FROM loss`, a slightly different plan)
is tallied as coverage, not failed.

Found, and pinned in ad_findings.rs: a table whose name has a capital has
no gradient in SQL. Its step is named after it, DataFusion lowercases the
unquoted name it is registered under, and ad::sql reads it back quoted.
Only the soak tries that variant until it is fixed. The MAX recomputation
bug also shows as NaN and inf gradients (a share over zero attaining
rows), which the stricter check now catches. With NaN or infinite data,
ddx-core's quotient rule gives ∞/∞ where the limit is 0; noted, and those
cases are exempt.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu
…st it always

grad_in_sql_of_a_table_with_capitals passes since #74 names steps in lower
case, so it runs as an ordinary test, and the `names` group now tries the
quoted-capitals variant in the PR gate too, not only in the soak.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
…er finding

ddx-core now writes a quotient's derivative term by term and a negative
power as a division (the fix chain beneath the stack), so:

- The soak's NaN check skips only NaN data. Infinite data is compared: a
  partial whose limit at ∞ is 0 now evaluates to 0, not ∞/∞.
- ad_findings.rs pins what a soak on the fixed stack found (seed 811): the
  partial of power(v, 0.5) at v = 0 was power(0, -0.5), which DataFusion
  refuses, failing a program ddx had accepted. It is now 0.5 / 0 = inf.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
Soaks on the fixed stack failed the stricter NaN check twice where the
loss does not move along d beyond rounding and the chain rule multiplies
an infinite partial by a zero one: sqrt of ln(softmax) of a single row is
sqrt(0) for every value (seed 811), and tanh(b · ∞) is 1 whatever b is,
so tanh'(∞) · ∂(b · ∞)/∂b = 0 · ∞ (seed 500192, infinite data). jax.grad
gives NaN on both. A NaN gradient where the loss moves by no more than
rounding is now screened; one where the loss moves is still a failure.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
…de's 0 · ∞

From an extreme-values soak on the fixed stack (seed 600421): σ(x · w)
with x = ∞ is 1 whatever w is, so that row does not move the loss, but its
partial is σ'(∞) · ∞ = 0 · ∞ = NaN, summed into w's gradient while other
rows keep the loss moving. jax.grad gives NaN on the same function. In
infinite-data mode a NaN gradient is now screened; a finite one is still
compared, which ddx-core's term-by-term quotient rule made possible.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
… the names check

From an extreme-values soak on the fixed stack:

- The descent check exempts kinked losses, but read only the root
  relation's kinds, not the head's aggregate: MAX(v) over parameters near
  1e-160 kinks within any step tried, so no step lowered the loss (seed
  1900692). Case::kinds now includes the head's MAX or MIN.
- grad in SQL plans the loss a little differently from the program API,
  so the two add in different orders; with huge values sin(SUM(v))
  turned an ulp of a sum near 1e9 into a relative difference near 1e-7
  (seed 1900495). The names check allows 1e-6 in huge mode.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
… screen it

From an extreme-values soak (seed 2000728, tiny values): d/dv ln(3/v) at
v = 1e-160 is -1e160, finite, but the chain rule passes through 3/v²,
which overflows, so ddx gives -inf (as jax.grad does). A non-finite
gradient where the loss moves faster than √f64::MAX is now screened.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
…t can't be judged

From an extreme-values soak on the fixed stack (seed 2400607, tiny
values): 1.5 / (0.7 / v) has derivative 1.5 / 0.7, but the chain rule
passes through -0.7 / v², which at v near 1e-160 overflows, so the
gradient is inf (jax.grad gives 0 · ∞ = NaN). In tiny mode a non-finite
gradient is now screened, as in infinite-data mode; a finite one is still
compared. The names check's grad(2L) - 2·grad(L) = 0 likewise expects
nothing where the gradient itself is not finite (∞ - ∞).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvszJhH8pn8H9G69wEPMU2
@alxmrs
alxmrs force-pushed the m4/v2-adapter-surface branch from 7bb05c2 to 74781b1 Compare October 3, 2026 22:02
@alxmrs
alxmrs merged commit 8fa6710 into main Oct 3, 2026
9 checks passed
@alxmrs
alxmrs deleted the m4/v2-adapter-surface branch October 3, 2026 22:15
@github-actions github-actions Bot mentioned this pull request Oct 3, 2026
alxmrs added a commit that referenced this pull request Oct 3, 2026
…(M4 9) (#93)

Stacked on #92. A clean soak proves only as much as the soak can see.
This PR measures that: it seeds one bug at a time into `ddx-ad` and the
DataFusion adapter and times how long the v2 soak takes to catch each.

## What it adds

- **`.github/scripts/mutation_test.py`**: 15 mutants, one per rule or
contract: each reduce rule, tie sharing, the map rule, fan-in, the
unreached-0 and NULL-row conventions, stop-gradient, the dims and
ranking checks, step naming, `release`, and `grad` in SQL's dims/values
split. Each is an exact text replacement that must match its file
exactly once, so a stale mutant fails loudly instead of silently testing
nothing. The source is restored afterwards, even on an interrupt.
- **A clean baseline.** The script first runs the soak unmutated over
the same seeds and skips every seed that fails there. The variants that
always meet a known, pinned bug are turned off (`DDX_V2_SKIP_KNOWN`). So
a kill is the mutant's doing. (The first attempt didn't have this; a
known bug "killed" a mutant.)
- **Kills by refusal.** If a mutant makes ddx accept 5 points fewer of
the generated cases than the baseline, it counts as killed. A
`ddx_stop_gradient` ddx no longer understands gets *refused*, not
answered wrongly, so the only symptom is the refusal rate.
- **`.github/workflows/mutation.yml`**: weekly and manually triggerable.
The table goes to the job summary; a survivor fails the job but files no
issue, since it's a gap in the tests, not a bug in ddx.
- `DDX_SOAK_STOP_ON_FAIL` for the runner.

## Results

The first sweep killed 10 of 15. Each survivor was a blind spot, now
closed:

| survivor | why | fix |
|---|---|---|
| `null-row-zero` | the clean baseline had turned NULLs off | keep NULLs
on; the baseline skips the known NULL leak |
| `stop-gradient-ignored` | the damage shows as refusals, not wrong
answers | refusal-rate kill |
| `ranking-check-off` | the generator only wrote total rankings |
non-total rankings over what a `wrt` table feeds; ddx must refuse them |
| `gradient-names-collide` | `w`, `b`, `u` never shared a name | `wrt`
tables `t` and `s1.t` |
| `sql-case-sensitive` | one value column per table hid the wrong
dims/values split | a two-value table, `grad(loss, t2.val, t2.VAL2)`
read with `SELECT *` |

Now all 15 die:

| mutant | time to kill | first failure |
|---|---|---|
| sum-halved | 1 case | vjp-vs-weighted-grad |
| mean-undivided, mean-over-sum, partial-is-one | 1 case | finite-diff |
| fan-in-first | 1 case | read-twice |
| dims-check-off | 1 case | dims-check |
| release-keeps-value | 1 case | catalog |
| extreme-everyone | 5 cases, 2s | finite-diff |
| sql-case-sensitive | 8 cases, 2s | names |
| gradient-names-collide | 14 cases, 6s | names |
| tie-unshared (wrong only at ties) | 16 cases, 8s | finite-diff |
| null-row-zero | 19 cases, 8s | null |
| unreached-one | 91 cases, 43s | finite-diff |
| ranking-check-off | 82 cases, 44s | contract |
| stop-gradient-ignored | 120s | accepts 73% of cases vs 80% unmutated |

The slowest three are where the soak is thinnest: rows no gradient
reaches, non-total rankings, and stop-gradient.

**A finding along the way.** The non-total ranking check first fired on
a ranking over *constant* data, which ddx doesn't need to refuse. But it
shows a latent risk in the same family as #91's MAX bug: constant
subtrees are recomputed in each backward step, unread and unchecked. A
tie-breaking-free ranking over constant data could keep a different row
on the way back, and a `CASE` on it would then take a different branch.
Recorded in `ad_findings.rs` under the recomputation finding.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_012NE18ox5ivwTc7ZSbHZUGu

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant