Skip to content

bench: reproducible 50-question recall comparison with Engram - #39

Merged
renezander030 merged 1 commit into
mainfrom
bench/agent-recall-engram
Oct 1, 2026
Merged

renezander030 merged 1 commit into
mainfrom
bench/agent-recall-engram

Conversation

@renezander030

Copy link
Copy Markdown
Owner

Adds a frozen 50-question synthetic recall benchmark against Engram v2.2.1 using identical title/body memories, both Engram match modes, and the real ATS Core hybrid RRF path with pinned local MiniLM embeddings. No vector database or hosted model key is used. The runner isolates the Engram store, saves per-question results and provenance, fails on degraded retrieval, and trips if ATS hybrid recall falls below the stronger Engram mode.

Measured Recall@5: ATS keyword 12/50, ATS hybrid 47/50, Engram default all 20/50, Engram any 32/50. README/report publish the keyword-only result, bucket breakdown, missed questions, raw output, reproduction steps, and limitations. Engram and ATS hybrid tie on exact-title and terse queries. This is ten distinct target decisions with five correlated questions each, not a production-scale product-superiority claim or an RRF ablation.

Validation: full repository gate passed all 454 tests and required proofs; check:publish passed. Added scorer and dataset integrity checks. The frozen dataset was not tuned after the measured run. Runtime/model caches and ordinary run output are ignored; only reviewed synthetic output is committed.

@renezander030
renezander030 merged commit a56e0d9 into main Oct 1, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant