Skip to content

perf: avoid unnecessary deep copies from direct replacement function calls - #875

Closed
davidbudzynski wants to merge 49 commits into
fastverse:developmentfrom
davidbudzynski:perf/no-deep-copy-replacement-calls
Closed

davidbudzynski wants to merge 49 commits into
fastverse:developmentfrom
davidbudzynski:perf/no-deep-copy-replacement-calls

Conversation

@davidbudzynski

@davidbudzynski davidbudzynski commented Aug 25, 2026 •

Copy link
Copy Markdown

Description

Fixes #311.

As described in the issue, calling base R replacement functions directly (e.g. `oldClass<-`(x, "cls")) entails a deep copy in some cases, while the equivalent assignment syntax (oldClass(x) <- "cls") does not. This PR rewrites all remaining direct replacement-function call sites in R/ to use assignment syntax. Semantics are fully preserved — only unnecessary allocations are removed.

Main Changes

Replacement calls converted to assignment syntax

All active call sites (~175 across 35 files in R/) of the following base replacement primitives were converted:

Primitive Example before Example after
`oldClass<-` / `class<-` return(\oldClass<-`(l, "GRP"))` oldClass(l) <- "GRP"; return(l)
`attributes<-` \`attributes<-\`(x, ax) attributes(x) <- ax; return(x)
`attr<-` setRnDF <- function(df, nm) \`attr<-\`(df, "row.names", nm) { attr(df, "row.names") <- nm; df }
`names<-` / `dimnames<-` return(\names<-`(res, lev))` names(res) <- lev; return(res)
`dim<-` \`dim<-\`(outer(...), NULL) res <- outer(...); dim(res) <- NULL
inline `[<-` \`[<-\`(logical(n), ind, TRUE) out <- logical(n); out[ind] <- TRUE

Files covered include the core grouping machinery (GRP.R, BY.R, qG/qF conversions), all grouped statistical functions (fsum, fmean, fsd/fvar, fmin/fmax, fprod, fnobs, ffirst/flast, fmode, fnth/fmedian, varying), data manipulation (fsubset/ftransform/fmutate/across, fselect/get_vars*, recode_replace, roworder/colorder, join, rsplit), quick conversions (qDF/qDT/qM/unattrib), and summary/printing methods (descr, qsu, psmat, pwcor/pwcov, psacf, unlist2d, flm, indexing, collap, dapply, fcount, B/W/HDB/HDW/STD).

Related clean-ups found during the sweep

  • Removed redundant attribute stripping before lapply()/vapply() in fdapply(), colsubset() and date_vars() — these iterate elements and ignore object-level attributes, so stripping only caused an extra copy
  • setup_across() now sets the class on a shallow copy (dsub <- d; oldClass(dsub) <- pe$cld) so the unclassed data returned in the result stays untouched — a comment explains why mutating d directly would be observable
  • Where a rewrite needed care to stay observable-behavior identical, comments were added or preserved (e.g. the GRP order attribute copy-semantics note, flm()'s dimnames handling for the qr method)

Intentionally left untouched

  • Commented-out legacy code and misc/legacy/
  • collapse's own user-defined replacement functions (`ftransform<-`, `get_vars_ind<-`, `add_vars<-`, ...): these are R-defined closures, not base primitives, and are not subject to the forced-copy behavior described in Optimize class assignments #311
  • src/ C code: attribute handling there uses different mechanisms entirely

Verification

  • Package builds locally; full testthat suite produces results identical to the unmodified development branch: the same 8 pre-existing environment-related errors on both, zero new failures (verified by building development into a separate library and diffing test outcomes)
  • All .Call() argument vectors in modified files were diffed byte-for-byte against development (this caught and fixed an accidentally flipped stable.algo flag during development — see commit history)
  • Spot-checked converted paths against the baseline build for identical outputs: fgroup_by(mtcars, cyl, vs2 = vs, am) renaming, findex_by() renaming, qtab() auto dimension naming, colorder(pos = "exchange"), na_omit(na.attr = TRUE)

Checklist

  • I have performed a self-review of my code.
  • I have commented on my code, particularly in hard-to-understand areas.
  • I have updated the documentation where applicable.

Notes on the checklist:

  • Self-review: full hunk-by-hunk review of the diff after the initial sweep; it surfaced 4 remaining inline [<- sites (in fgroup_by, findex_by, qtab and posord) which were converted in follow-up commits
  • Comments: added where the transformation is non-obvious (setup_across, GRP order attribute, flm, redundant-strip removals); mechanical one-to-one rewrites are left uncommented to keep the diff readable
  • Documentation: all touched helpers are internal (not exported); exported functions keep identical signatures and behavior, so no roxygen/.Rd updates were applicable

Additional Context

Benchmark evidence

Allocations per call measured with Rprofmem() (R 4.6.1, macOS arm64), comparing the direct-call form with the equivalent assignment form used in this PR:

Case direct-call form assignment form
names<- on freshly computed grouped result 21.2 KB/call 0.8 KB/call
dimnames<- on grouped matrix result 2.4 KB/call 1.1 KB/call
attributes<- strip before lapply() 0.4 KB/call 0.3 KB/call
oldClass<- on freshly built list 781.6 KB/call 781.5 KB/call
attributes<- factor conversion 1.7 KB/call 1.7 KB/call

The largest wins are on hot paths that name small grouped results (use.g.names = TRUE), where the direct-call form duplicated the entire result vector.

  • 47 commits, one conventional commit per file / function group for easy review
  • No API or behavioral changes intended; purely an internal allocation optimization

Dawid Budzyński added 30 commits August 25, 2026 14:44
Replace direct calls to base replacement functions (attr<-, attributes<-,
dim<-, dimnames<-) with equivalent assignment syntax, which avoids the
forced deep copy that occurs when replacement functions are invoked
directly. See fastverse#311.
- na_omit(): set the omit class with oldClass(x) <- before attaching
- ffka(), setRnDF(), setRownames(): assignment syntax instead of direct
  attributes<-/attr<- calls
- fdapply(), colsubset(): drop redundant attribute stripping before
  lapply()/vapply(), which ignore object-level attributes anyway

See fastverse#311.
Build the GRP object into a local variable and set its class with
oldClass(g) <- "GRP" instead of wrapping the list literal in a direct
oldClass<- call, and use names(x) <- / attributes(x) <- assignment
syntax for the group names and retained order attributes.

See fastverse#311.
Construct the GRP list in a local variable, assign group names with
names(x) <- and set the class with oldClass(x) <- "GRP" instead of
direct oldClass<-/names<- replacement calls.

See fastverse#311.
- fgroup_by: assign names on the name-lookup list with names(x) <-
- names<-.GRP_df: unclass, rename, then restore the saved class
- fgroup_vars: build named indices/logicals with assignment syntax
  instead of direct names<-/[<- replacement calls
- as_factor_qG/qG: set attributes with attributes(x) <- instead of
  direct attributes<- calls

See fastverse#311.
Replace direct names<-/dimnames<- replacement calls on freshly computed
results with the equivalent assignment form, avoiding forced deep
copies of the result vectors/matrices. See fastverse#311.
Replace direct names<-/dimnames<- replacement calls on freshly computed
results with the equivalent assignment form, avoiding forced deep
copies. See fastverse#311.
- get_vars_ind/get_vars_indl/fselect: build named indices and logical
  vectors with assignment syntax instead of direct names<-/[<-/attr<-
  replacement calls
- get_vars_ind<-: restore the class with oldClass(x) <- before the
  conditional data.table allocation check
- date_vars: drop redundant attribute stripping before vapply()

See fastverse#311.
- fsubset/ftransform/fmutate helpers: use names(x) <- and oldClass(x) <-
  assignment syntax instead of direct names<-/oldClass<-[<- calls
- setup_across: set the class on a shallow copy so the unclassed data
  returned in the result stays untouched
- fFUN internals: name results with assignment syntax

See fastverse#311.
Use oldClass(x) <- and attributes(x) <- assignment syntax instead of
direct oldClass<-/attributes<- replacement calls in replace_na,
replace_outliers and pad. See fastverse#311.
Build named statistics and the attrib attribute with assignment syntax
instead of direct names<-/attr<- replacement calls. See fastverse#311.
Dawid Budzyński added 17 commits August 25, 2026 14:59
The fsd.default/fvar.default refactors accidentally swapped the
stable.algo flag in the grouped vector paths (fsd must use TRUE,
fvar FALSE). Restores the exact original .Call arguments.
Re-adds the explanatory notes about shallow/deep copy behavior on the
GRP order attribute and documents why flm strips dimnames for the qr
method only.
@davidbudzynski
davidbudzynski marked this pull request as ready for review August 25, 2026 13:48
@davidbudzynski
davidbudzynski marked this pull request as draft August 25, 2026 13:48
@davidbudzynski
davidbudzynski marked this pull request as ready for review August 25, 2026 13:48
@SebKrantz

Copy link
Copy Markdown
Member

Thanks @davidbudzynski! @claude can we also check whether

lapply(res, function(l) {
    names(l) <- rn
    l
  })

is really an improvement over

res <- lapply(res, `names<-`, rn)

Please test using different data sizes.

@claude

claude Bot commented Aug 26, 2026 •

Copy link
Copy Markdown

Claude finished @SebKrantz's task in 3m 34s —— View job


Task: benchmark lapply + assignment-syntax vs lapply(res, \names<-`, rn)`

  • Read trigger comment and understand request
  • Locate the relevant call site in the PR diff
  • Attempt to run benchmark across different data sizes
  • Report results and recommendation

Location

The exact snippet quoted is t_list() in R/rsplit.R:9-18 (the only active, non-commented call site matching this pattern — used by unlist2d/matrix-to-list conversion to attach row names):

t_list <- function(l) {
  lmat <- do.call(rbind, l)
  dn <- dimnames(lmat)
  res <- .Call(Cpp_mctl, lmat, !is.null(dn[[2L]]), 0L)
  if(length(rn <- dn[[1L]])) res <- lapply(res, function(l) {
    names(l) <- rn
    l
  })
  .Call(C_copyMostAttrib, res, l)
}

⚠️ Unable to execute the benchmark in this environment

This job's sandbox has no R/Rscript installed (type R → "not found"), and package-manager commands are blocked, so I could not actually run Rprofmem()/timing comparisons here as requested. I don't want to report fabricated numbers, so instead here is (a) the mechanistic reasoning for what to expect, and (b) a ready-to-run script for a maintainer to execute locally across sizes, mirroring the methodology already used in the PR description.

Expected behavior (mechanistic reasoning)

The two forms differ in two ways that pull in opposite directions:

  1. Per-call closure overhead: lapply(res, names<-, rn) calls the primitive names<- directly — one R-level call per element. The closure form function(l) { names(l) <- rn; l } also does one call per element, but that call itself evaluates a complex assignment (*tmp* <- l; l <- \names<-`(tmp, rn); rm(tmp)) plus a final l` lookup — strictly more interpreter work per element.
  2. Duplication avoided: complex assignment (names(l) <- rn) is specifically optimized in R's evaluator (applydefine in eval.c) to avoid bumping the refcount of l before calling names<-, so if l's only reference at that point is the local binding, the primitive can set the attribute in place. Calling names<- as an ordinary function argument (FUN(X[[i]], ...) inside lapply) does not get this treatment, so the primitive typically sees MAYBE_REFERENCED(l) == TRUE and must duplicate() the whole vector before attaching rn.

So the expected crossover is size-dependent, exactly like the other cases already measured in the PR body (e.g. attributes<- on a small factor showed ~no difference, while names<- on a large grouped vector showed 21.2 KB → 0.8 KB): for very small elements (a handful of values per list entry) the fixed closure-call overhead may make the two forms roughly a wash or even slightly favor the direct-call form; for anything beyond trivial size, avoiding the full-vector duplicate() should dominate and favor the assignment form, scaling with length(l).

Suggested benchmark (for a maintainer to run locally)

library(collapse)
bench <- function(n_elem, len) {
  rn <- paste0("r", seq_len(len))
  mk <- function() replicate(n_elem, stats::rnorm(len), simplify = FALSE)

  # allocations
  res1 <- mk(); res2 <- mk()
  m1 <- sum(gc(reset = TRUE)["Vcells", 2]); a1 <- {
    invisible(lapply(res1, `names<-`, rn))
  }
  m_direct <- utils::capture.output(Rprofmem(tf <- tempfile()))
  Rprofmem(tf1 <- tempfile()); invisible(lapply(res1, `names<-`, rn)); Rprofmem(NULL)
  Rprofmem(tf2 <- tempfile()); invisible(lapply(res2, function(l) { names(l) <- rn; l })); Rprofmem(NULL)
  mem <- function(f) sum(as.numeric(gsub("^([0-9]+).*", "\\1", readLines(f)[grepl("^[0-9]", readLines(f))])))
  bytes_direct <- mem(tf1); bytes_assign <- mem(tf2)

  # timing
  tm <- bench::mark(
    direct = lapply(res1, `names<-`, rn),
    assign = lapply(res2, function(l) { names(l) <- rn; l }),
    iterations = 200, check = FALSE
  )
  list(n_elem = n_elem, len = len,
       bytes_direct = bytes_direct, bytes_assign = bytes_assign,
       time_direct = tm$median[1], time_assign = tm$median[2])
}

grid <- expand.grid(n_elem = c(5, 50, 500), len = c(10, 1e3, 1e5))
res <- do.call(rbind, Map(function(n, l) as.data.frame(bench(n, l)), grid$n_elem, grid$len))
print(res)

(Vary n_elem = number of list elements / matrix columns, len = length of each element / number of rows, to cover both "many small columns" and "few huge columns" cases relevant to t_list()'s actual usage from matrix-to-list conversion.)

Recommendation

Given I can't produce real numbers here, I'd suggest @davidbudzynski or @SebKrantz run the script above (or a simplified Rprofmem-only version matching the PR's existing table) at small (len ~10), medium (len ~1e3–1e4) and large (len ~1e5–1e6) element sizes before merging this particular hunk. If small-len cases show the closure form is neutral-to-worse there but large-len cases show a clear win (which is the expected pattern based on how the rest of the PR's benchmarks behave), it's probably still a net positive since rsplit/unlist2d matrix conversions are more often used on non-trivial data — but it's worth confirming empirically rather than assuming, since this is the one hunk in the PR where the two forms are closest in cost (unlike e.g. the oldClass<- cases where the closure form is unambiguously better).

If it turns out the direct-call form is faster/equal for realistic sizes, this hunk (R/rsplit.R:13-16) should probably be reverted to res <- lapply(res, \names<-`, rn)` while keeping the rest of the PR's changes.

Benchmarks requested in fastverse#875 (R 4.6.1, macOS arm64) show
lapply(res, `names<-`, rn) is ~2x faster than the closure form at all
tested sizes (5-500 elems x 10-1e6 rows), and tracemem/profmem confirm
that neither form duplicates element data through lapply on R >= 4.x,
so the complex-assignment rewrite only added interpreter overhead
(~150 ns/elem) with no memory benefit.
…rk audit

Systematic benchmarks on R 4.6.1 (macOS arm64) show that direct calls to
base replacement primitives are performance-equivalent to assignment
syntax in every context occurring in this sweep: neither form duplicates
data through lapply/apply-family callbacks, function arguments, or
inline-computed temporaries (collapse requires R >= 4.1, so reference
counting semantics apply throughout). Closure rewrites inside iteration
add interpreter overhead (~150 ns/element), and removing the attribute
strip before vapply() in date_vars() silently added column names to two
documented return modes.

Whole-function dev-vs-PR comparisons (fmean/fsum/collap/B/W/fmutate over
nrow 1e4-1e6 x ncol 10-50 x groups 10-500) were statistically
indistinguishable, and all 10 functional snapshots matched byte-for-byte.
Full test suite: 12,683 passes / 0 failures / 10 pre-existing skips.

Kept (verified behaviorally identical, strictly less work):
- fdapply(): drop attribute strip before lapply()
- colsubset(): drop attribute strip before vapply()

See fastverse#875 and fastverse#311 for the full audit.
@davidbudzynski

davidbudzynski commented Aug 26, 2026 •

Copy link
Copy Markdown
Author

@SebKrantz Benchmark results for your question (R 4.6.1, macOS arm64; min of 5 auto-sized batches per size, exact allocations via profmem(), duplication detection via tracemem(), all forms gated on identical() output before timing):

elems × len lapply(res, \names<-`, rn)` closure form for-loop
50 × 10 7.4 µs 16.1 µs 5.5 µs
500 × 1e3 67.0 µs 154.3 µs 443.8 µs
100 × 1e6 15.5 µs 33.3 µs 43,000 µs

the direct call is consistently ~2× faster at every size tested (5–500 elements × 10–10⁶ rows), and tracemem() shows neither form duplicates any element data through lapply on R ≥ 4.x (only ~O(100 B)/call attribute-node bookkeeping differs). The closure form just adds interpreter overhead (~150 ns/element). (The for-loop is the variant that does deep-copy: its allocations scale with element length.)

I've reverted t_list() back to lapply(res, `names<-`, rn) in e1d30e0, and then went on to benchmark-audit every rewrite pattern in this PR — summary in my comment below; full scripts: https://gist.github.com/davidbudzynski/9af71984f7b06b4ee459f25c69833e88.

@davidbudzynski

Copy link
Copy Markdown
Author

Following up on the review discussion above: after systematically benchmarking every rewrite pattern in this PR, I've rolled the sweep back to just two changes that demonstrably help (04333c6).

What exactly was benchmarked

Each old form was compared head-to-head with its PR rewrite:

# Pattern (development → PR) Sizes
0 lapply(res, `names<-`, rn) → closure {names(l) <- rn; l} (t_list), + for-loop variant 5–500 elems × 10–10⁵ rows; 10–100 × 10⁶
1 y <- `names<-`(fresh_local, v) → bind-then-assign len 10, 10³, 10⁵, 10⁶
2 same primitive applied to a function argument (return path) same grid
3 primitive fed an inline-computed temp (`names<-`(rnorm(n), v)) → bind-first rewrite same grid
4 df reclassing `oldClass<-`(df, cls) → oldClass(df) <- cls; `attr<-` row.names on fresh frames 10³–10⁶ rows × 5–10 cols (+ tracemem dup check)
5 attribute-strip removals before lapply()/vapply() (fdapply, colsubset, date_vars) behavioral equivalence suites vs a development build
E2E whole calls, dev-lib vs PR-lib: fmean (vector & df), fsum (matrix), collap, B, W, fmutate nrow {10⁴, 10⁶} × ncol {10, 50} × groups {10, 500}; 10 functional snapshots diffed with identical()

How

  • R 4.6.1, macOS arm64, collapse 2.1.7; both branches exported via git archive and installed with R CMD INSTALL into separate library trees (no dev-mode loading for any measurement)
  • Timing: ≥2 warm-up passes; iteration count auto-tuned to ~0.25 s/batch; reported statistic = fastest of 5 batches; gc(reset = TRUE) before every batch
  • Memory: exact allocation bytes per call via profmem(); real duplication independently verified with tracemem()
  • Correctness gates: identical(values + attributes) required between forms before any number was recorded; RNG state seeded or inputs deep-copied so both forms always receive identical data
  • Behavioral hunks verified by running both versions side-by-side from the isolated builds on realistic inputs incl. edge cases

Findings

  • On R ≥ 4.1 (collapse's minimum), direct replacement calls are performance-equivalent to assignment syntax in all tested contexts: nothing is duplicated through lapply, function arguments, or computed temporaries — only ~O(100 B)/call of attribute-node bookkeeping differs
  • Closure rewrites inside iteration add pure interpreter overhead (~150 ns/element)
  • Two hunks changed behavior, not just performance: dropping the attribute-strip in date_vars() silently added column names to the documented "indices"/"logical" return modes, and the setup_across shallow-copy change fixed no actual mutation (refcounting already prevented it)
  • End-to-end timings are statistically indistinguishable; all 10 functional snapshots byte-identical between branches

What was kept

Verified behaviorally identical and strictly less work: removing the redundant attributes<-(x, NULL) strips before lapply()/vapply() in fdapply() and colsubset().

Net branch diff vs development is now these 2 lines; full test suite: 12,683 passed / 0 failed / 10 pre-existing skips. Unless you'd prefer folding these into #311, my plan is to keep this PR open as this minimal change.

Scripts & raw data: https://gist.github.com/davidbudzynski/9af71984f7b06b4ee459f25c69833e88

Full archetype results (µs/call · bytes/call · tracemem copies)
archetype elem_len form us/call bytes/call tracemem copies
A1_local 10 direct 0.92 NA 1
A1_local 10 assign 1 NA 1
A1_local 1,000 direct 22.88 20,144 0
A1_local 1,000 assign 23.36 20,144 0
A1_local 100,000 direct 2236.36 2,000,144 0
A1_local 100,000 assign 2207.21 2,000,144 0
A1_local 1,000,000 direct 23300 20,000,144 0
A1_local 1,000,000 assign 23300 20,000,144 0
A2_arg 10 direct 1.04 NA 1
A2_arg 10 assign 1.06 NA 1
A2_arg 1,000 direct 22.56 20,144 0
A2_arg 1,000 assign 22.32 20,144 0
A2_arg 100,000 direct 2093.46 2,000,144 0
A2_arg 100,000 assign 2076.92 2,000,144 0
A2_arg 1,000,000 direct 22600 20,000,144 0
A2_arg 1,000,000 assign 23200 20,000,144 0
A3_computed 10 direct 0.82 NA 0
A3_computed 10 assign 0.9 NA 0
A3_computed 1,000 direct 21.92 20,144 0
A3_computed 1,000 assign 21.76 20,144 0
A3_computed 100,000 direct 2043.1 2,000,144 0
A3_computed 100,000 assign 2068.38 2,000,144 0
A3_computed 1,000,000 direct 23000 20,000,144 0
A3_computed 1,000,000 assign 22200 20,000,144 0
A5_dfclass 1,000 direct 129.62 140,576 1
A5_dfclass 1,000 assign 128.18 140,576 1
A5_dfclass 100,000 direct 9280 14,000,576 1
A5_dfclass 100,000 assign 9240 14,000,576 1
A5_dfclass 1,000,000 direct 86600 140,000,576 1
A5_dfclass 1,000,000 assign 88400 140,000,576 1

A5 shows both forms shallow-copy the frame identically when reclassing through an argument — assignment syntax buys nothing there either.

End-to-end dev-vs-PR timing deviations (%)
nrow ncol groups fmean_v fsum_m fmean_df B collap ftransform (PR vs dev)
10,000 10 10 +3.6% +2.9% +1.9% +4.3% +6.1% +4.7%
1,000,000 10 10 -0.5% +4.3% -0.5% +4.9% +2.5% +1.7%
10,000 50 10 +3.6% +4.3% -0.5% +5.8% +2.1% +2.2%
1,000,000 50 10 +0.3% +1.5% +0.0% +0.5% -1.2% +1.1%
10,000 10 500 +0.0% +1.9% +2.8% +3.1% +8.7% +3.1%
1,000,000 10 500 +0.0% +3.5% +1.0% +4.6% +4.0% +4.0%
10,000 50 500 +0.0% +2.9% +1.8% +6.8% +1.6% +2.7%
1,000,000 50 500 +1.5% +1.8% +3.1% +0.0% +4.2% +0.4%

Bold cells exceed ±5%; all are small-data cases where absolute deltas are single-digit microseconds, and functions untouched by the PR show the same scatter (run noise).

@SebKrantz

Copy link
Copy Markdown
Member

Thanks @davidbudzynski! I just tested this though and find:

library(collapse)
#> collapse 2.1.7, see ?`collapse-package` or ?`collapse-documentation`
#> 
#> Attaching package: 'collapse'
#> The following object is masked from 'package:stats':
#> 
#>     D
library(bench)

duplAttributes <- collapse:::duplAttributes
fdapply <- function(X, FUN, ...) duplAttributes(lapply(`attributes<-`(X, NULL), FUN, ...), X)
fdapply2 <- function(X, FUN, ...) duplAttributes(lapply(X, FUN, ...), X)

df <- qDF(lapply(rep(10000,1000), rnorm))
mark(fdapply(df, identity), fdapply2(df, identity))
#> # A tibble: 2 × 6
#>   expression                  min   median `itr/sec` mem_alloc `gc/sec`
#>   <bch:expr>             <bch:tm> <bch:tm>     <dbl> <bch:byt>    <dbl>
#> 1 fdapply(df, identity)     141µs    164µs     5919.    7.86KB     62.0
#> 2 fdapply2(df, identity)    147µs    167µs     5934.    7.86KB     65.1
df <- qDF(lapply(rep(100000,100), rnorm))
mark(fdapply(df, identity), fdapply2(df, identity))
#> # A tibble: 2 × 6
#>   expression                  min   median `itr/sec` mem_alloc `gc/sec`
#>   <bch:expr>             <bch:tm> <bch:tm>     <dbl> <bch:byt>    <dbl>
#> 1 fdapply(df, identity)    14.7µs   17.2µs    57083.      848B     68.6
#> 2 fdapply2(df, identity)   15.9µs   18.3µs    52384.      848B     62.9
df <- qDF(lapply(rep(1000,10000), rnorm))
mark(fdapply(df, identity), fdapply2(df, identity))
#> # A tibble: 2 × 6
#>   expression                  min   median `itr/sec` mem_alloc `gc/sec`
#>   <bch:expr>             <bch:tm> <bch:tm>     <dbl> <bch:byt>    <dbl>
#> 1 fdapply(df, identity)    1.41ms   1.57ms      632.    78.2KB     82.8
#> 2 fdapply2(df, identity)   1.48ms   1.64ms      607.    78.2KB     79.3

Created on 2026-08-27 with reprex v2.1.1

The reasons appears to be that lapply(), while ignoring attributes, still does internal checks on them which can be eliminated by removing them beforehand. That is just one of those things I guess which only the nerdiest of microbenchmark fanatics, of which I was one when I wrote that code some 5 years ago, get to discover.

So yeah, really sorry if that leaves this PR void, but I highly appreciate the effort and believe we can close #311 now – it is very good to know that all the memory issues have been addressed and we are really not making any mistakes as far as based R programming in collapse is concerned. Perhaps also #311 was resolved by imrovements to the R source code itself. R has become a lot more memory efficient over the years.

@SebKrantz

Copy link
Copy Markdown
Member

Closed for now...

@SebKrantz SebKrantz closed this Aug 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants