Repository navigation
OpenHuman-shaped simulation, hardening & benchmark suite for tinymemory on local CortexDB #251
Description
Activity
- addedpriority: p3Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect.Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect.
on Oct 10, 2026 Large proposal to build an OpenHuman-shaped benchmark/hardening suite for tinymemory: host simulator, high-complexity scenarios, KPI upgrades, leakage audit, fuzzing, and docs. It's an enhancement request with substantial new functionality, not a broken core path; no matching prior art among candidates.
Labelled
priority: p3.Benchmark update after PR #252
PR #252 is merged. It added the first OpenHuman host eval profile, deadline-aware pre-turn probes, p99 and timeout reporting, and a JSON-scripted 500-turn loop guard. This is the first slice of #251; the live model step, memory tool, import/source paths, broader JSON scenarios, leakage audit, fuzzing, and other-engine guide remain open.
What ran
All CortexDB runs used a throwaway local CortexDB v0.10.4 on the direct wire. The full mock run used deterministic mock inference across the existing 12 scenarios (53 probes, 52 with an expected pack answer). The live runs used OpenRouter models, the optional answer model, and the tool_heavy scenario (7 scripted exchanges; 6 probes scored before and after synthesis). I repeated the live run three times with fresh databases and an enrichment wait cap of 90 seconds. The 500-turn guard used the reference engine. These are saved in the eval JSON reports and summarized in docs/evals/openhuman-host.md.
Run Pack hit before / after synthesis Model answer after synthesis Scripted pre-turn p95 Timed-out pre-turn calls CortexDB model cost Full mock, 12 scenarios 41/52 / 40/52 not run 358 ms 0/175 $0.000 OpenRouter tool_heavy, repeat 1 0/6 / 5/6 4/6 1502 ms 10/19 $0.050 OpenRouter tool_heavy, repeat 2 0/6 / 5/6 5/6 1502 ms 11/19 $0.012 OpenRouter tool_heavy, repeat 3 0/6 / 5/6 4/6 1508 ms 12/19 $0.030 The live timeout denominator is 7 scripted turns plus 12 pre-turn probes per run. Across repeats, 16/21 scripted turns and 17/36 probes timed out: 33/57 combined (58%). Seventeen of the eighteen first-phase probes timed out; the remaining probe also missed its expected fact. After enrichment and belief builds, no synthesis-phase probe timed out and 15/18 packs contained the expected answer. The answer model got 13/18 synthesis probes right. These are six-probe runs, so that accuracy is not a full-suite estimate.
The mock run had 28/32 lexical and 13/20 paraphrase pack hits before synthesis; 28/32 and 12/20 after it. Team handoff was 0/3 before and 1/3 after; the needle scenario was 1/2 in both phases; both context.md probes missed. No measured leak probe fired (0/6 checked probe phases). The mock extractor builds no facts or beliefs, so its 0/52 captured score says nothing about live extraction. Its p50/p95/p99 scripted pre-turn times were 49/358/386 ms.
The live scenario consistently missed who-to-ask after synthesis (3/3 repeats). The synthetic git_blame result names jmiller, and the derived-layer captured check passed for all six facts, but the pack never included that name. This narrows the next investigation to retrieval, ranking, and pack inclusion of tool-result details; it does not show that the write was lost. ToolCallRef itself retains only a name and ID, while this profile preserves a bounded result line in the assistant reply.
Live cost also varied substantially for the same fixture: 80–181 CortexDB model calls, 46,619–146,933 tokens, and $0.012–$0.050 per run, plus about $0.002/run for the optional answer model. Enrichment accounted for about 83% of the three-run CortexDB cost. Built beliefs ranged from 10 to 56 and held facts from 30 to 70. Those ranges need a controlled enrichment/queue completion check before attributing the variation to a particular model or setting.
The reference-engine loop guard completed 500 turns with 500 nonempty packs, zero timeouts, and zero stored literal <memory-context> tags. Mean pack size moved from 114 to 130 tokens (first versus last 50 turns); repeated bullet-line rate moved from 82.4% to 85.7%. The repeated acknowledgement in the script causes most of those duplicate lines. This guard checks the literal wrapper in exported stored items; it does not yet prove that unwrapped pack prose or derived layers cannot echo.
Interpretation and next fixes
- Correct host parity before using the timeout figures as a production estimate. The merged profile always invokes pre_turn_dated with a ready None hint. OpenHuman's default date_hint is false and normally invokes pre_turn. TinyMemory treats any late hint as dated and reads up to 3× deeper in fetch sections. This may inflate the measured timeouts. Make the eval call pre_turn by default, use pre_turn_dated only when date hints are enabled, and exercise pre_turn_resumed after compaction. Add a parity assertion against the host defaults, then rerun at least three live repeats. The current 53–63% rates describe the merged eval profile, not verified default OpenHuman behavior.
- Instrument the hot path before optimizing it. Record scope discovery, per-section and per-scope fetch duration, OpenRouter embedding time, and the eventual completion time of calls that exceeded the 1500 ms host deadline. The current p95 near 1500 ms is censored by the timeout and cannot reveal backend p95. CortexEngine fetch can issue one query embedding per scope with four scope packs in flight; measure that cost and then consider reuse/caching or fewer scope reads. Keep timeout rate as a separate release gate.
- Give tool results a retrievable unit and probe. Test storing a bounded tool result with source/provenance as its own turn or item, then verify the jmiller probe across direct and hosted wires. Compare retrieval rank and pack budget before changing the ranking policy. Keep the existing logged-reply line as the host-facing fallback.
- Improve mock fidelity and coverage. Add deterministic semantic-ish embeddings and a minimal extractor so offline fact, belief, paraphrase, and conflict metrics are meaningful. Extend JSON turn scripts to coding, long tasks, compaction, deletion, tenant isolation, and needle-at-scale cases. For the loop guard, compare pack text hashes/excerpts against stored turns and derived layers, not just the wrapper tag.
- Run the safety and isolation audit as separate regression slices. Confirm each suspected leak or deletion residual on a live engine, add a reproducing test, and track only confirmed defects in focused fix issues. Measure leakage by channel and deletion immediately and after enrichment.
The next immediate benchmark action is the parity correction in (1), followed by repeat live runs with completion-stage timings. That will tell us how much of the observed timeout gap is the harness's dated overfetch and how much remains on OpenHuman's default path.
Benchmark follow-up: PR #253
PR #253 implements the next harness and measurement slice from the analysis above. The prior live runs always used dated recall, while OpenHuman defaults to plain recall. This PR mirrors the plain path by default, exercises resumed recall after compaction, records when timed-out work actually finishes, removes probe questions from the measured corpus, and measures scope discovery and recall-pack stages. It adds seeded scale sweeps and a deletion/isolation audit. The detailed methods and commands are in the updated eval report.
Corrected host run
Three fresh OpenRouter/CortexDB v0.10.4
tool_heavyruns used the plain OpenHuman hook and the same six questions as before:Repeat Initial pack Synthesis pack Synthesis model answer Timed-out host calls CortexDB model cost 1 3/6 6/6 5/6 7/19 $0.031 2 1/6 6/6 6/6 10/19 $0.025 3 4/6 6/6 6/6 6/19 $0.017 The previously consistent
jmillerwho-to-ask miss is present in all three synthesis packs. Initial retrieval and the 1.5-second host deadline remain poor: 6–10 of 19 calls timed out per run. The 19 calls are seven scripted turns plus six probes in each phase. In the third repeat, 57 timing samples had p95 scope discovery of 13 ms, recall-pack time of 2,233 ms, and CortexDB fetch total of 2,235 ms. Pack timing includes server embedding/ranking; the wire does not expose them separately. Timed-out tasks continue and their eventual completion time is now reported, so the deadline-censored p95 is no longer the only latency view.Coverage and scale
The final mock run covered 14 scenarios and 58 scored probes per phase. Expected text appeared in 48/58 packs and the extractive answerer got 28/58 correct in each phase. It had zero timeouts across 195 host calls; pre-turn p50/p95/p99 were 30/76/144 ms. New coding and task-drift cases each had 3/3 pack hits but 1/3 extractive answers. Mock inference created no derived facts or beliefs, so its derived-layer metrics are not a live accuracy estimate. An opt-in 200-turn reference compaction run preserved 3/3 pack hits in both phases; deduplicating identical pending builds changed synthesis from 41 jobs/84.7 seconds to 2 jobs/4.1 seconds.
The seeded reference-engine sweep placed one answer among 100, 1,000, or 10,000 near-duplicate documents at early, middle, and late positions. Lexical lookup hit rank 1 in all nine cells; paraphrase missed in all nine. This is a toy scorer baseline. The 10,000-item batch setup took 6.8–7.5 seconds after removing quadratic duplicate checks, and probes took 457–1,079 ms.
Live scale results show three separate failure modes:
Fresh collections List settlement Direct ranked readiness Host pack result Cost per run 100 documents, 3 repeats 1.3–2.1 s yes, after 15.1–19.8 s 0/6 hits; 6/6 timed out $0.006–$0.018 1,000 documents, 3 repeats 0.4–31.7 s 2/3 ready within 30 s 1/6 lexical, 0/6 paraphrase; 4/6 timed out $0.097–$0.103 An earlier 100-document controlled run became ranked-ready after 5.5 seconds and returned the lexical answer but missed the paraphrase. An earlier 1,000-document attempt had only 645 listable writes at the old 90-second cap; the final repeats used a 300-second cap and larger pages. Each repeat uses a fresh database. The readiness query names the planted owner without using either scored question. Listability, ranked readiness, and the host deadline therefore need separate gates; a ready direct fetch did not imply a timely host pack.
Safety and next fixes
A controlled live audit passed 20/20 sibling isolation, forget, and erase checks across packs, direct answers,
context.md, export, and derived layers. It established a derived record before each deletion and confirmed the restore created a new record. One earlier immediate-forget run showed a derived residual after forget while public item channels were clear; two immediate/controlled follow-ups and this stronger audit passed. That race is not yet reproduced deterministically and should have its own focused regression before a production fix is claimed.The next engineering targets are (1) server recall-pack latency, especially embedding/ranking inside the 1.5-second host budget; (2) index readiness and its variance after bulk accepted writes; (3) semantic paraphrase retrieval; and (4) a focused extraction-versus-forget race test. The current PR makes each measurable but does not claim those live performance gates are green. OpenRouter spending for this work stayed under the $5 run budget.
Accuracy follow-up after the harness work:
- Complex agent-history case: The
coding_sessionpack contained both the earlytest_retries_refused_connectionresult and the latertest_does_not_repeat_unknown_writeresult. The answer model chose the early result when it appeared first. Reversing their presentation in a controlled call corrected the answer. PR #254 presents already-selected agent-history turns newest first within each thread, with regression tests. On the same mock CortexDB fixture andopenai/gpt-4.1-minianswerer, thecoding_session,task_driftsix-question slice held 6/6 pack hits in each phase; model answers moved from 5/6 to 6/6 in recall and from 5/6 to 6/6 in synthesis. This is a targeted six-question result, not a full-suite gain estimate. - Host deadline: OpenHuman #7346 merged the default pre-turn timeout change from 1.5 to 5 seconds; explicit configurations still override it. PR Improve answers from recent agent history #254 mirrors that value in the eval and records the deadline in JSON. In a fresh live CortexDB v0.10.4 100-document run with OpenRouter models, the lexical probe hit and returned the correct model answer in 2.08 seconds; the paraphrase probe missed and returned
unknownin 1.78 seconds. Neither timed out. This collection is not paired with the older 1.5-second runs, so it cannot establish a timeout-rate delta. - Remaining accuracy gap: A direct recall over the same 100-document index ranked the planted owner document feat: stand up the TinyMemory workspace — contract, driver registry, shared mandatory families, TinyCortex adapter #1 for the lexical question and Let configuration select the memory engine, gated by the registry (#18 §A5) #23 for the paraphrase. The agent pack admitted only the top slice, so the latter answer was absent before the answer model ran. Extending the deadline does not fix that semantic ranking. The next accuracy experiment should test CortexDB query variants or a reranker with repeated fresh collections and model-scored pack/answer rates, including cost and pack-size effects. Simply raising a global document limit would bring many decoys into the context and can hurt answer quality.
Full methodology and prior 1.5-second runs:
docs/evals/openhuman-host.mdin PR #254. TinyMemory's four local contract checks and the 29 eval-example tests passed; PR CI is running.- Complex agent-history case: The
Why
Today
memory_evalruns 12 hand-written scenarios, about 53 probes, one after another, with substring scoring. It drives the library's idea of the lifecycle, not OpenHuman's. OpenHuman calls a different surface:pre_turn_dated/pre_turn_resumedunder a 1500 ms timeout,logged_replytool lines, its ownmemorytool (recall|fetch|learn|forget) instead ofMemoryTools, a scrub guard before every store,Brain::ingest_with(Accepted),collect_items → store_many, and the v1 import loop with its own retries. Much of that surface has never run end-to-end against a live engine. We want a suite that:CortexDB only for now. Section H documents how any other engine is benchmarked the same way.
What already exists (build on it, don't duplicate)
crates/tinymemory-integrations/examples/memory_eval/:scenarios.rs); KPIs (kpi.rs) covering pack hit, MRR, extractive and LLM answers, captured, leaks, cost viav1/admin/usage, and p50/p95;comparewith noise bands, and JSON and markdown output.scripts/memory-eval.shruns the eval withMODELS=openrouteron a throwaway compose stack.scripts/memory-flag-sweep.shruns 17 flag profiles and checks the key's budget first.integration/cortexdb/holds the compose file,cortex.toml,mock_inference.py(hash embeddings, no extraction) andflags/*.env.tests/live_cortexdb.rsandtests/live_cortex_lifecycle.rsare the live contract and lifecycle tests.src/cortex/testing/is the in-process HTTP double, with fault knobs.crates/tinymemory-api/tests/conformance_reference.rsinjects faults into the conformance suite.docs/evals/cortex-flags.md): pack hit 96%, model answer 87%, planted conflicts flagged 0%, fresh-first 57%, synthesis gain +0 pp, $0.43/run,pre_turnp50/p95 1.1/1.5 s, which is already at OpenHuman's 1500 ms timeout.A. OpenHuman host simulator (
--host openhumanprofile in memory_eval)The dependency direction is one-way. OpenHuman imports tinymemory (vendored at
vendor/tinymemory, pinned at v1.27.1 against the current v1.27.2). tinymemory must never depend onopenhuman-core, not even as a dev-dependency or from an example. A cycle would break both builds and the vendoring.<memory-context>pack injected into a prompt must never be written back as memory. That would be a recall → store → recall echo that amplifies itself turn after turn. A probe asserts that stored turns and ingested items never contain pack text, and that pack size and duplicate rate stay flat over a 500-turn run.A mocked OpenHuman turn is the unit of simulation:
user message →
pre_turn_*(with the 1500 ms timeout) → model step → zero or more tool calls, including thememorytool and a fake tool with canned results → reply →post_turn(logged_reply)→ enqueue the returned jobs →BackgroundRunner.The model step is pluggable:
memory.The same scenario runs in both modes. Comparing them separates retrieval quality from model behaviour.
AgentMemory::with_policywith OpenHuman's defaults: budget 1200, learnings 8, brain 6, history 6, team 0, beliefs every 10 turns,build_delay_secs300.pre_turn_dated/pre_turn_resumedwrapped intokio::time::timeout(1500ms)(a timeout counts as an empty pack and is recorded).post_turnuseslogged_reply(a "tool → result" line appended to the reply).2n/2n+1.<memory-context>.recall_for_compactionruns under an 8000 ms timeout.safety::scrub_item_withguard, as in OpenHuman'sguard.rs.memorytool with recall/fetch/learn/forget/erase/list, calling the engine directly asops.rsdoes.BackgroundRunnerwith a persisted queue, plus hand-builtBuildBeliefsafter a source sync.Brain::ingest_with(WriteOptions::accepted());reader_for_request+collect_items→store_many;LegacyWorkspace::open(..).skip_connector_syncs(true)+items_fromwith a retry loop.cortexdband hostedtinyhumanswhen its env is set. OpenHuman defaults to hosted, and the lifecycle has never run live on it.B. Higher-complexity OpenHuman scenarios
Each scenario gets probes tagged lexical/paraphrase/temporal/multi-hop, with expect/accept/stale/forbidden. A generator (seeded, deterministic) can scale any scenario from S to L (tens → thousands of events).
ToolCallRefkeeps only name/id).pre_turn_resumed.in_prompt_from.in_prompt_from, and never the logged turn.observed_actor/subjectattribution.refers_to, plus trigger triage over noisy events.ConverterChain/OfficeConverterand the folder reader.context.md, immediately and again after enrichment and belief builds.needle_in_noiseis one rule among 24 threads. Extend it with:KPIs per (size × depth): pack hit, rank/MRR, answer accuracy and
pre_turnlatency. Report them as a heatmap table so recall decay with scale is visible.13. Multi-tenant org. Org root, several users and shared agents. Cross-user and cross-org probes must never leak.
C. KPIs and harness upgrades
--max-usd) that reads OpenRouterlimit_remainingbefore and during a run.--judge) with recorded rationale;in_prompt_from.context.md), with forbidden strings checked across all of them.pre_turnracingpost_turnon the same thread.REPEAT ≥ 3to measure the actual noise band (never measured so far).D. Leakage and isolation audit — simulate and confirm
Code reading surfaced these. Each needs a live simulation, then a regression test, then a fix issue if it is confirmed.
parse_belowlocates the root withrposition(cortex/envelope/layout.rs:378-381). A namespace such asws:main/user:42under rootuser:42can parse as the root.engine/scopes.rs:156,184).engine/recall.rs:16-28,52,71).reach(engine/beliefs.rs:79,145-160).cascade: "redact_events"is trusted but unverified against real CortexDB, especially for beliefs that cite several events (log/forget.rs:28).memory_forgetwith no reach deletes ids globally, service sandboxes included (tools/write/mod.rs:102);ToolScope::default()andContextSpec.reach: Noneare unscoped;Facet::narrow(Namespace)replaces a confined reach (explore/mod.rs:158).##headings or---frontmatter in the pack (recall/render.rs:159).engine/forget.rs:71-80,engine/erase.rs:47-52).transport/failure.rs:161,192-215, logged atstore.rs:212andrecall/gather.rs:112).InvalidRequeston a dated recall turns date hints off for the life of the process (engine/fetch.rs:185-188).log/read.rs:141-150).Segment::sanitizedcollisions (a valid id spelled like a sanitized one; ids over 119 chars rely on a 32-bit FNV), with PII left in scope paths (namespace/mod.rs:126-145). Follows up Namespace addressing: non-injective sanitization, PII-redacted section prefixes, and unspecified recall(None) #116.envelope/labels.rs:33).The conformance suite should gain recall-answer and beliefs isolation checks across sibling namespaces, with new fault variants in
conformance_reference.rs.E. Fuzzing and property tests
There are none today. Add
proptest(dev-dependency) for:Segment::sanitizedinjectivity;render/parseround-trip, including an interior root repeat;MetaFilter/Reach::admits;host_fixed_keyrecursion has no depth bound);TimeHint(no span limit).Add
cargo-fuzztargets for envelope decode/rebuild and tool-call JSON. Add chaos runs that use the HTTP double's fault knobs (429/503, apply-then-fail, hidden listings), and on live, a proxy that drops or delays requests.F. Mock fidelity (so CI measures more than wiring)
mock_inference.py: deterministic semantic-ish embeddings (for example a hashed bag of words or char n-grams) and a minimal rule-based extractor, so facts and beliefs appear offline.v1/facts,v1/conflicts,v1/admin/usageand/admin/ready, so the inspector runs against it.memory_eval --host openhumanon the mocks. Real-model runs stay manual or nightly behindOPENROUTER_API_KEY.G. Running it locally
MODELS=openrouter OPENROUTER_API_KEY=… scripts/memory-eval.shbrings upcortexdb:v0.10.4on :3145 with real models.limit_remainingfirst (the sweep already does).target/memory-eval/<label>.{md,json}, andmemory_eval comparegives the deltas.H. Docs: benchmarking another engine's integration
Add
docs/evals/engine-integration.md:tinymemory_api::conformance::run(engine), including the new isolation checks.registrysomemory_eval --engine <id>and--host openhumancan drive it.REPEAT ≥ 3and compare against the CortexDB baseline.Also refresh
docs/evals/agent-memory.md(it predates the 12-scenario harness) and add the missingintegration/cortexdb/flags/README.md.Proposed gates (initial, to be tuned after the first real-model baseline)
pre_turnp95 (OpenHuman profile)Suggested order (sub-issues)
Out of scope