Skip to content

OpenHuman-shaped simulation, hardening & benchmark suite for tinymemory on local CortexDB #251

Description

@senamakel

Why

Today memory_eval runs 12 hand-written scenarios, about 53 probes, one after another, with substring scoring. It drives the library's idea of the lifecycle, not OpenHuman's. OpenHuman calls a different surface: pre_turn_dated/pre_turn_resumed under a 1500 ms timeout, logged_reply tool lines, its own memory tool (recall|fetch|learn|forget) instead of MemoryTools, a scrub guard before every store, Brain::ingest_with(Accepted), collect_items → store_many, and the v1 import loop with its own retries. Much of that surface has never run end-to-end against a live engine. We want a suite that:

  1. replays realistic OpenHuman workloads through that exact surface, against a local CortexDB on real OpenRouter models;
  2. scores recall, accuracy, latency, cost and leakage with noise-aware comparisons;
  3. turns gaps into tracked findings and regression tests.

CortexDB only for now. Section H documents how any other engine is benchmarked the same way.

What already exists (build on it, don't duplicate)

  • crates/tinymemory-integrations/examples/memory_eval/:
    • 12 scenarios (scenarios.rs); KPIs (kpi.rs) covering pack hit, MRR, extractive and LLM answers, captured, leaks, cost via v1/admin/usage, and p50/p95;
    • compare with noise bands, and JSON and markdown output.
  • scripts/memory-eval.sh runs the eval with MODELS=openrouter on a throwaway compose stack.
  • scripts/memory-flag-sweep.sh runs 17 flag profiles and checks the key's budget first.
  • integration/cortexdb/ holds the compose file, cortex.toml, mock_inference.py (hash embeddings, no extraction) and flags/*.env.
  • tests/live_cortexdb.rs and tests/live_cortex_lifecycle.rs are the live contract and lifecycle tests.
  • src/cortex/testing/ is the in-process HTTP double, with fault knobs.
  • crates/tinymemory-api/tests/conformance_reference.rs injects faults into the conformance suite.
  • Last recorded real-model baseline (docs/evals/cortex-flags.md): pack hit 96%, model answer 87%, planted conflicts flagged 0%, fresh-first 57%, synthesis gain +0 pp, $0.43/run, pre_turn p50/p95 1.1/1.5 s, which is already at OpenHuman's 1500 ms timeout.

A. OpenHuman host simulator (--host openhuman profile in memory_eval)

The dependency direction is one-way. OpenHuman imports tinymemory (vendored at vendor/tinymemory, pinned at v1.27.1 against the current v1.27.2). tinymemory must never depend on openhuman-core, not even as a dev-dependency or from an example. A cycle would break both builds and the vendoring.

  • The simulator therefore mirrors OpenHuman's memory shim inside the tinymemory eval. It re-implements the few hooks below, with OpenHuman's defaults as named constants and a parity table in the docs that points at each OpenHuman source line.
  • Scenarios and the turn script are data files (JSON) under the eval. OpenHuman may later load them in its own tests, which keeps the arrow pointing OpenHuman → tinymemory and lets it check parity from its side.
  • Runtime loop guard. The <memory-context> pack injected into a prompt must never be written back as memory. That would be a recall → store → recall echo that amplifies itself turn after turn. A probe asserts that stored turns and ingested items never contain pack text, and that pack size and duplicate rate stay flat over a 500-turn run.

A mocked OpenHuman turn is the unit of simulation:

user message → pre_turn_* (with the 1500 ms timeout) → model step → zero or more tool calls, including the memory tool and a fake tool with canned results → reply → post_turn(logged_reply) → enqueue the returned jobs → BackgroundRunner.

The model step is pluggable:

  • mock (default, CI): deterministic and scripted. It chooses tool calls and replies from the scenario script and answers extractively from the pack.
  • live: an OpenRouter chat model given OpenHuman-style system and memory-context prompts. It records tokens and cost and lets the model decide when to call memory.

The same scenario runs in both modes. Comparing them separates retrieval quality from model behaviour.

  • Policy: AgentMemory::with_policy with OpenHuman's defaults: budget 1200, learnings 8, brain 6, history 6, team 0, beliefs every 10 turns, build_delay_secs 300.
  • Turn hooks: pre_turn_dated/pre_turn_resumed wrapped in tokio::time::timeout(1500ms) (a timeout counts as an empty pack and is recorded).
    • post_turn uses logged_reply (a "tool → result" line appended to the reply).
    • Turn indices are 2n/2n+1.
    • The pack goes inside <memory-context>.
    • recall_for_compaction runs under an 8000 ms timeout.
  • Writes: every store goes through a safety::scrub_item_with guard, as in OpenHuman's guard.rs.
  • Model tool: a memory tool with recall/fetch/learn/forget/erase/list, calling the engine directly as ops.rs does.
  • Jobs: BackgroundRunner with a persisted queue, plus hand-built BuildBeliefs after a source sync.
  • Ingestion:
    • Brain::ingest_with(WriteOptions::accepted());
    • reader_for_request + collect_items → store_many;
    • LegacyWorkspace::open(..).skip_connector_syncs(true) + items_from with a retry loop.
  • Both wires: direct cortexdb and hosted tinyhumans when its env is set. OpenHuman defaults to hosted, and the lifecycle has never run live on it.

B. Higher-complexity OpenHuman scenarios

Each scenario gets probes tagged lexical/paraphrase/temporal/multi-hop, with expect/accept/stale/forbidden. A generator (seeded, deterministic) can scale any scenario from S to L (tens → thousands of events).

  1. Coding session. A multi-hour session in a repo: file reads, failing test output, stack traces, diffs and decisions inside tool results.
    • Probes: "why did we revert X", "which file holds the retry logic", "what was the flaky test".
    • Covers tool-result retention (known gap: ToolCallRef keeps only name/id).
  2. Long-running task (task_manager/goals). A goal worked over a simulated 14 days with checkpoints, blockers, a changed plan, and resumes through pre_turn_resumed.
    • Probes: current status, why the plan changed, what is still blocked.
    • Tests "newest state wins" and multi-day aging.
  3. Recall after compaction. Threads of 200, 500 and 1000 turns, compacted repeatedly with a sliding in_prompt_from.
    • Facts planted only in the compacted region must come back.
    • Invariant: nothing at or after in_prompt_from, and never the logged turn.
  4. Orchestrator → subagents. The planner, critic and summarizer fan out.
    • Workflow and service-sandbox writes must stay out of chat packs.
    • The team section must not echo the agent's own turns (known follow-up).
  5. Multi-channel inbox. Slack, Telegram and email threads with several senders and observed_actor/subject attribution.
    • Probes: "what did Alice ask for", and cross-thread disambiguation.
  6. Morning briefing / cron triggers. Dated recall such as "what happened yesterday" and "last Tuesday's meeting" via refers_to, plus trigger triage over noisy events.
  7. Document-heavy brain.
    • Real PDF, DOCX, PPTX, XLSX, Markdown and code fixtures through ConverterChain/OfficeConverter and the folder reader.
    • Web and RSS readers against a local fixture HTTP server.
    • Score extraction fidelity: tables, multi-page, long docs split into chunks.
  8. Legacy v1 import → recall. Import a synthetic v1 workspace, then ask questions only the imported data answers.
  9. Preference drift and corrections. The user changes preferences and corrects the agent; the latest must win, and the stale value must never come first.
  10. Multi-hop and temporal "as of". "Who owned billing when the outage happened?" Add a temporal KPI.
  11. Right to be forgotten. A user asks the agent to forget X through the tool, or the host erases a source. X must be absent from the pack, recall answer, beliefs, facts, export and context.md, immediately and again after enrichment and belief builds.
  12. Needle in a haystack, at scale. Today's needle_in_noise is one rule among 24 threads. Extend it with:
  • a single needle in 100, 1k and 10k events, swept by depth (early, middle, late) and age (same day vs 30 days old);
  • multi-needle (3–5 facts that must be combined);
  • needles hidden in a document page, a tool result, an imported v1 item and a compacted thread region;
  • near-duplicate decoys and paraphrased needles.

KPIs per (size × depth): pack hit, rank/MRR, answer accuracy and pre_turn latency. Report them as a heatmap table so recall decay with scale is visible.
13. Multi-tenant org. Org root, several users and shared agents. Cross-user and cross-org probes must never leak.

C. KPIs and harness upgrades

  • Latency:
    • p50/p95/p99 per call (pre_turn, post_turn, compaction, ingest, recall tool);
    • cold vs warm;
    • timeout rate against OpenHuman's 1500 ms and 8000 ms budgets;
    • write→visible lag and write→extracted lag per item.
  • Cost:
    • USD and tokens per turn, per write, per ingested MB and per correct answer, split by role (enrichment, extraction, answer, verifier);
    • an in-eval budget guard (--max-usd) that reads OpenRouter limit_remaining before and during a run.
  • Accuracy:
    • keep substring scoring;
    • add an optional LLM-judge (--judge) with recorded rationale;
    • a paraphrase/lexical/temporal/multi-hop breakdown;
    • planted conflicts across more than the single "refund" case.
  • Pack invariants: budget adherence (with a real tokenizer as well as 4 chars/token), dedupe, no logged turn, nothing at or after in_prompt_from.
  • Leakage: leak rate per channel (pack, answer text, beliefs, facts, team section, context.md), with forbidden strings checked across all of them.
  • Deletion: residual rate after forget or erase, measured immediately and after enrichment.
  • Concurrency: N parallel agents and tenants, with throughput and tail latency while enrichment runs; pre_turn racing post_turn on the same thread.
  • Repeats: a real-model REPEAT ≥ 3 to measure the actual noise band (never measured so far).

D. Leakage and isolation audit — simulate and confirm

Code reading surfaced these. Each needs a live simulation, then a regression test, then a fix issue if it is confirmed.

  1. Interior root repeat. Hosted parse_below locates the root with rposition (cortex/envelope/layout.rs:378-381). A namespace such as ws:main/user:42 under root user:42 can parse as the root.
    • Risk: root-inherited reads/beliefs, and erase-at-root reaching that scope (engine/scopes.rs:156,184).
  2. Recall answer text ignores metadata filters. Only citations are filtered; the answer is grounded on a whole pack plus facts/beliefs (engine/recall.rs:16-28,52,71).
  3. Belief hits are not re-checked against reach (engine/beliefs.rs:79,145-160).
  4. Derived layers survive forget/erase. cascade: "redact_events" is trusted but unverified against real CortexDB, especially for beliefs that cite several events (log/forget.rs:28).
  5. Fail-open defaults:
    • memory_forget with no reach deletes ids globally, service sandboxes included (tools/write/mod.rs:102);
    • ToolScope::default() and ContextSpec.reach: None are unscoped;
    • Facet::narrow(Namespace) replaces a confined reach (explore/mod.rs:158).
  6. Unscrubbed writes and prompt injection.
    • The safety scrub runs only in OpenHuman's guard, never in tinymemory itself.
    • It skips tags, tool_calls, thread/agent ids, repo and source.id.
    • Recalled prose is rendered raw and can forge ## headings or --- frontmatter in the pack (recall/render.rs:159).
  7. Partial forget/erase on transport failure returns an error with no report of what was already deleted (engine/forget.rs:71-80, engine/erase.rs:47-52).
  8. Backend error bodies (up to 300 chars) reach logs. If a 4xx echoes content, user text is logged (transport/failure.rs:161,192-215, logged at store.rs:212 and recall/gather.rs:112).
  9. Sticky downgrades.
    • One InvalidRequest on a dated recall turns date hints off for the life of the process (engine/fetch.rs:185-188).
    • Attribution is turned off after one refusal.
  10. Silent truncation beyond 1000 scopes on lenient reads (log/read.rs:141-150).
  11. Segment::sanitized collisions (a valid id spelled like a sanitized one; ids over 119 chars rely on a 32-bit FNV), with PII left in scope paths (namespace/mod.rs:126-145). Follows up Namespace addressing: non-injective sanitization, PII-redacted section prefixes, and unspecified recall(None) #116.
  12. Path digests are unsalted SHA-256, so they can be reversed by dictionary (envelope/labels.rs:33).

The conformance suite should gain recall-answer and beliefs isolation checks across sibling namespaces, with new fault variants in conformance_reference.rs.

E. Fuzzing and property tests

There are none today. Add proptest (dev-dependency) for:

  • Segment::sanitized injectivity;
  • layout render/parse round-trip, including an interior root repeat;
  • MetaFilter/Reach::admits;
  • tool-args parsing (host_fixed_key recursion has no depth bound);
  • TimeHint (no span limit).

Add cargo-fuzz targets for envelope decode/rebuild and tool-call JSON. Add chaos runs that use the HTTP double's fault knobs (429/503, apply-then-fail, hidden listings), and on live, a proxy that drops or delays requests.

F. Mock fidelity (so CI measures more than wiring)

  • mock_inference.py: deterministic semantic-ish embeddings (for example a hashed bag of words or char n-grams) and a minimal rule-based extractor, so facts and beliefs appear offline.
  • The Rust double gains v1/facts, v1/conflicts, v1/admin/usage and /admin/ready, so the inspector runs against it.
  • A CI smoke job runs memory_eval --host openhuman on the mocks. Real-model runs stay manual or nightly behind OPENROUTER_API_KEY.

G. Running it locally

  • MODELS=openrouter OPENROUTER_API_KEY=… scripts/memory-eval.sh brings up cortexdb:v0.10.4 on :3145 with real models.
  • Mind the shared key's daily cap: check limit_remaining first (the sweep already does).
  • Results go to target/memory-eval/<label>.{md,json}, and memory_eval compare gives the deltas.

H. Docs: benchmarking another engine's integration

Add docs/evals/engine-integration.md:

  1. Pass tinymemory_api::conformance::run(engine), including the new isolation checks.
  2. Register the engine in registry so memory_eval --engine <id> and --host openhuman can drive it.
  3. Run the scenario suite with REPEAT ≥ 3 and compare against the CortexDB baseline.
  4. Meet the release gates.
  5. Run the leakage, deletion and fuzz suites.

Also refresh docs/evals/agent-memory.md (it predates the 12-scenario harness) and add the missing integration/cortexdb/flags/README.md.

Proposed gates (initial, to be tuned after the first real-model baseline)

KPI Gate
Leaks (all channels) 0
Deletion residual 0
Pack-invariant violations 0
pre_turn p95 (OpenHuman profile) < 1500 ms; timeout rate < 1%
Pack hit / paraphrase pack hit ≥ 95% / ≥ 90%
Model answer accuracy ≥ 90%
Fresh-first on contradictions ≥ 90% (today 57%)
Planted conflicts flagged > 0 (today 0%)
Cost per full run tracked, with no regression beyond the noise band

Suggested order (sub-issues)

  1. A: the host simulator and the mocked OpenHuman turn (mock model first), plus the C latency/timeout/cost upgrades and the loop guard.
  2. D: confirmation runs (fix issues per confirmed risk).
  3. B scenarios 1–3: coding, long-running, compaction.
  4. B scenarios 4–13, plus the generator for scale.
  5. C concurrency and the LLM judge.
  6. E: fuzz and proptest.
  7. F: mock fidelity and the CI smoke job.
  8. H: docs.

Out of scope

  • Implementing other engines.
  • Changes in the OpenHuman repo, beyond noting the vendored-pin gap.

Activity

  1. added
    priority: p3Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect.
    on Oct 10, 2026
  2. tinysweeper commented on Oct 10, 2026

    @tinysweeper

    Large proposal to build an OpenHuman-shaped benchmark/hardening suite for tinymemory: host simulator, high-complexity scenarios, KPI upgrades, leakage audit, fuzzing, and docs. It's an enhancement request with substantial new functionality, not a broken core path; no matching prior art among candidates.

    Labelled priority: p3.

  3. senamakel commented on Oct 10, 2026

    @senamakel
    MemberAuthor

    Benchmark update after PR #252

    PR #252 is merged. It added the first OpenHuman host eval profile, deadline-aware pre-turn probes, p99 and timeout reporting, and a JSON-scripted 500-turn loop guard. This is the first slice of #251; the live model step, memory tool, import/source paths, broader JSON scenarios, leakage audit, fuzzing, and other-engine guide remain open.

    What ran

    All CortexDB runs used a throwaway local CortexDB v0.10.4 on the direct wire. The full mock run used deterministic mock inference across the existing 12 scenarios (53 probes, 52 with an expected pack answer). The live runs used OpenRouter models, the optional answer model, and the tool_heavy scenario (7 scripted exchanges; 6 probes scored before and after synthesis). I repeated the live run three times with fresh databases and an enrichment wait cap of 90 seconds. The 500-turn guard used the reference engine. These are saved in the eval JSON reports and summarized in docs/evals/openhuman-host.md.

    Run Pack hit before / after synthesis Model answer after synthesis Scripted pre-turn p95 Timed-out pre-turn calls CortexDB model cost
    Full mock, 12 scenarios 41/52 / 40/52 not run 358 ms 0/175 $0.000
    OpenRouter tool_heavy, repeat 1 0/6 / 5/6 4/6 1502 ms 10/19 $0.050
    OpenRouter tool_heavy, repeat 2 0/6 / 5/6 5/6 1502 ms 11/19 $0.012
    OpenRouter tool_heavy, repeat 3 0/6 / 5/6 4/6 1508 ms 12/19 $0.030

    The live timeout denominator is 7 scripted turns plus 12 pre-turn probes per run. Across repeats, 16/21 scripted turns and 17/36 probes timed out: 33/57 combined (58%). Seventeen of the eighteen first-phase probes timed out; the remaining probe also missed its expected fact. After enrichment and belief builds, no synthesis-phase probe timed out and 15/18 packs contained the expected answer. The answer model got 13/18 synthesis probes right. These are six-probe runs, so that accuracy is not a full-suite estimate.

    The mock run had 28/32 lexical and 13/20 paraphrase pack hits before synthesis; 28/32 and 12/20 after it. Team handoff was 0/3 before and 1/3 after; the needle scenario was 1/2 in both phases; both context.md probes missed. No measured leak probe fired (0/6 checked probe phases). The mock extractor builds no facts or beliefs, so its 0/52 captured score says nothing about live extraction. Its p50/p95/p99 scripted pre-turn times were 49/358/386 ms.

    The live scenario consistently missed who-to-ask after synthesis (3/3 repeats). The synthetic git_blame result names jmiller, and the derived-layer captured check passed for all six facts, but the pack never included that name. This narrows the next investigation to retrieval, ranking, and pack inclusion of tool-result details; it does not show that the write was lost. ToolCallRef itself retains only a name and ID, while this profile preserves a bounded result line in the assistant reply.

    Live cost also varied substantially for the same fixture: 80–181 CortexDB model calls, 46,619–146,933 tokens, and $0.012–$0.050 per run, plus about $0.002/run for the optional answer model. Enrichment accounted for about 83% of the three-run CortexDB cost. Built beliefs ranged from 10 to 56 and held facts from 30 to 70. Those ranges need a controlled enrichment/queue completion check before attributing the variation to a particular model or setting.

    The reference-engine loop guard completed 500 turns with 500 nonempty packs, zero timeouts, and zero stored literal <memory-context> tags. Mean pack size moved from 114 to 130 tokens (first versus last 50 turns); repeated bullet-line rate moved from 82.4% to 85.7%. The repeated acknowledgement in the script causes most of those duplicate lines. This guard checks the literal wrapper in exported stored items; it does not yet prove that unwrapped pack prose or derived layers cannot echo.

    Interpretation and next fixes

    1. Correct host parity before using the timeout figures as a production estimate. The merged profile always invokes pre_turn_dated with a ready None hint. OpenHuman's default date_hint is false and normally invokes pre_turn. TinyMemory treats any late hint as dated and reads up to 3× deeper in fetch sections. This may inflate the measured timeouts. Make the eval call pre_turn by default, use pre_turn_dated only when date hints are enabled, and exercise pre_turn_resumed after compaction. Add a parity assertion against the host defaults, then rerun at least three live repeats. The current 53–63% rates describe the merged eval profile, not verified default OpenHuman behavior.
    2. Instrument the hot path before optimizing it. Record scope discovery, per-section and per-scope fetch duration, OpenRouter embedding time, and the eventual completion time of calls that exceeded the 1500 ms host deadline. The current p95 near 1500 ms is censored by the timeout and cannot reveal backend p95. CortexEngine fetch can issue one query embedding per scope with four scope packs in flight; measure that cost and then consider reuse/caching or fewer scope reads. Keep timeout rate as a separate release gate.
    3. Give tool results a retrievable unit and probe. Test storing a bounded tool result with source/provenance as its own turn or item, then verify the jmiller probe across direct and hosted wires. Compare retrieval rank and pack budget before changing the ranking policy. Keep the existing logged-reply line as the host-facing fallback.
    4. Improve mock fidelity and coverage. Add deterministic semantic-ish embeddings and a minimal extractor so offline fact, belief, paraphrase, and conflict metrics are meaningful. Extend JSON turn scripts to coding, long tasks, compaction, deletion, tenant isolation, and needle-at-scale cases. For the loop guard, compare pack text hashes/excerpts against stored turns and derived layers, not just the wrapper tag.
    5. Run the safety and isolation audit as separate regression slices. Confirm each suspected leak or deletion residual on a live engine, add a reproducing test, and track only confirmed defects in focused fix issues. Measure leakage by channel and deletion immediately and after enrichment.

    The next immediate benchmark action is the parity correction in (1), followed by repeat live runs with completion-stage timings. That will tell us how much of the observed timeout gap is the harness's dated overfetch and how much remains on OpenHuman's default path.

  4. senamakel commented on Oct 10, 2026

    @senamakel
    MemberAuthor

    Benchmark follow-up: PR #253

    PR #253 implements the next harness and measurement slice from the analysis above. The prior live runs always used dated recall, while OpenHuman defaults to plain recall. This PR mirrors the plain path by default, exercises resumed recall after compaction, records when timed-out work actually finishes, removes probe questions from the measured corpus, and measures scope discovery and recall-pack stages. It adds seeded scale sweeps and a deletion/isolation audit. The detailed methods and commands are in the updated eval report.

    Corrected host run

    Three fresh OpenRouter/CortexDB v0.10.4 tool_heavy runs used the plain OpenHuman hook and the same six questions as before:

    Repeat Initial pack Synthesis pack Synthesis model answer Timed-out host calls CortexDB model cost
    1 3/6 6/6 5/6 7/19 $0.031
    2 1/6 6/6 6/6 10/19 $0.025
    3 4/6 6/6 6/6 6/19 $0.017

    The previously consistent jmiller who-to-ask miss is present in all three synthesis packs. Initial retrieval and the 1.5-second host deadline remain poor: 6–10 of 19 calls timed out per run. The 19 calls are seven scripted turns plus six probes in each phase. In the third repeat, 57 timing samples had p95 scope discovery of 13 ms, recall-pack time of 2,233 ms, and CortexDB fetch total of 2,235 ms. Pack timing includes server embedding/ranking; the wire does not expose them separately. Timed-out tasks continue and their eventual completion time is now reported, so the deadline-censored p95 is no longer the only latency view.

    Coverage and scale

    The final mock run covered 14 scenarios and 58 scored probes per phase. Expected text appeared in 48/58 packs and the extractive answerer got 28/58 correct in each phase. It had zero timeouts across 195 host calls; pre-turn p50/p95/p99 were 30/76/144 ms. New coding and task-drift cases each had 3/3 pack hits but 1/3 extractive answers. Mock inference created no derived facts or beliefs, so its derived-layer metrics are not a live accuracy estimate. An opt-in 200-turn reference compaction run preserved 3/3 pack hits in both phases; deduplicating identical pending builds changed synthesis from 41 jobs/84.7 seconds to 2 jobs/4.1 seconds.

    The seeded reference-engine sweep placed one answer among 100, 1,000, or 10,000 near-duplicate documents at early, middle, and late positions. Lexical lookup hit rank 1 in all nine cells; paraphrase missed in all nine. This is a toy scorer baseline. The 10,000-item batch setup took 6.8–7.5 seconds after removing quadratic duplicate checks, and probes took 457–1,079 ms.

    Live scale results show three separate failure modes:

    Fresh collections List settlement Direct ranked readiness Host pack result Cost per run
    100 documents, 3 repeats 1.3–2.1 s yes, after 15.1–19.8 s 0/6 hits; 6/6 timed out $0.006–$0.018
    1,000 documents, 3 repeats 0.4–31.7 s 2/3 ready within 30 s 1/6 lexical, 0/6 paraphrase; 4/6 timed out $0.097–$0.103

    An earlier 100-document controlled run became ranked-ready after 5.5 seconds and returned the lexical answer but missed the paraphrase. An earlier 1,000-document attempt had only 645 listable writes at the old 90-second cap; the final repeats used a 300-second cap and larger pages. Each repeat uses a fresh database. The readiness query names the planted owner without using either scored question. Listability, ranked readiness, and the host deadline therefore need separate gates; a ready direct fetch did not imply a timely host pack.

    Safety and next fixes

    A controlled live audit passed 20/20 sibling isolation, forget, and erase checks across packs, direct answers, context.md, export, and derived layers. It established a derived record before each deletion and confirmed the restore created a new record. One earlier immediate-forget run showed a derived residual after forget while public item channels were clear; two immediate/controlled follow-ups and this stronger audit passed. That race is not yet reproduced deterministically and should have its own focused regression before a production fix is claimed.

    The next engineering targets are (1) server recall-pack latency, especially embedding/ranking inside the 1.5-second host budget; (2) index readiness and its variance after bulk accepted writes; (3) semantic paraphrase retrieval; and (4) a focused extraction-versus-forget race test. The current PR makes each measurable but does not claim those live performance gates are green. OpenRouter spending for this work stayed under the $5 run budget.

  5. senamakel commented on Oct 10, 2026

    @senamakel
    MemberAuthor

    Accuracy follow-up after the harness work:

    • Complex agent-history case: The coding_session pack contained both the early test_retries_refused_connection result and the later test_does_not_repeat_unknown_write result. The answer model chose the early result when it appeared first. Reversing their presentation in a controlled call corrected the answer. PR #254 presents already-selected agent-history turns newest first within each thread, with regression tests. On the same mock CortexDB fixture and openai/gpt-4.1-mini answerer, the coding_session,task_drift six-question slice held 6/6 pack hits in each phase; model answers moved from 5/6 to 6/6 in recall and from 5/6 to 6/6 in synthesis. This is a targeted six-question result, not a full-suite gain estimate.
    • Host deadline: OpenHuman #7346 merged the default pre-turn timeout change from 1.5 to 5 seconds; explicit configurations still override it. PR Improve answers from recent agent history #254 mirrors that value in the eval and records the deadline in JSON. In a fresh live CortexDB v0.10.4 100-document run with OpenRouter models, the lexical probe hit and returned the correct model answer in 2.08 seconds; the paraphrase probe missed and returned unknown in 1.78 seconds. Neither timed out. This collection is not paired with the older 1.5-second runs, so it cannot establish a timeout-rate delta.
    • Remaining accuracy gap: A direct recall over the same 100-document index ranked the planted owner document feat: stand up the TinyMemory workspace — contract, driver registry, shared mandatory families, TinyCortex adapter #1 for the lexical question and Let configuration select the memory engine, gated by the registry (#18 §A5) #23 for the paraphrase. The agent pack admitted only the top slice, so the latter answer was absent before the answer model ran. Extending the deadline does not fix that semantic ranking. The next accuracy experiment should test CortexDB query variants or a reranker with repeated fresh collections and model-scored pack/answer rates, including cost and pack-size effects. Simply raising a global document limit would bring many decoys into the context and can hurt answer quality.

    Full methodology and prior 1.5-second runs: docs/evals/openhuman-host.md in PR #254. TinyMemory's four local contract checks and the 29 eval-example tests passed; PR CI is running.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestpriority: p3Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect.

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions