Memory Benchmarks  ·  STATUS · UPDATED 2026-09-18

What We Measured.What We Have Not.

Agent memory is full of scores. Most are self-authored, unscoped, or measured on query sets derived from the expected answers. This page is the opposite: mechanism evidence, four measured baselines with their flaws attached, and a named list of the benchmarks we have not run.

0.19mean injection utility, live baseline (≈81% unused)
0.0%citation rate in the claim-annotation A/B (both arms)
AbsentLongMemEval · LoCoMo · longitudinal compounding
Phase 6the measurement program for all of the above

State of play

Mechanism Evidence Is Not Outcome Evidence.

A retrieval pipeline can be architecturally sound and still inject mostly noise. The honest sequence is: build the mechanism, measure whether it helps, publish both. We are between steps two and three, and this page says so.

Measured · with caveats

Four Baselines, Flaws Attached.

These are the readings that exist. Each one is scoped, dated and labeled with what it does not prove, because the flaw is part of the finding.

MeasurementResultn / scopeStatus
Synthetic ranking baselinenDCG@5 0.648–0.671 · MRR 0.834–0.875100 seeded nodes → 53 queries · fixed seed · queries derived from expected nodes, so the result is circular by constructionCircular by design
Observed ranking baselineInsufficient positives: 185 feedback rows, 2 sessions, 4 query hashes, 0 positive observationsLive store · 2026-08-31Inconclusive
Claim annotation A/BNull result: citation rate 0.0% in both arms · treatment added +32% injection-block tokensn=57 observations · 3 paired prompts · 2026-08-31Negative, published
Injection utility baselineMean 0.189 → 0.186 (≈81% of injected memory unused)918 → 999 samples · live store · 2026-08-31Open finding
Mechanism evidenceSix-stage retrieval, submodular packing under a 4,096-token budget, per-type decay (Decision half-life ≈347 d → Approach ≈23 d), git-anchored drift demotion, Beta-Bernoulli utility updatesShipped paths, default-on unless noted; the learning stack is default-offMechanism, not outcome

Absent · named

The Benchmarks We Do Not Have.

An absent measurement is absent, not zero. These are the runs analysts will ask for first, listed before anyone has to ask.

∅LongMemEvalNo run exists. Long-horizon memory QA would be the right fixture, and we have not executed it, so we do not cite a score.
∅LoCoMoNo run exists. Same rule: absent, not zero.
∅Live shadow comparisonThe shadow evaluation path is implemented and default-off; it has not yet run over organic sessions at a sample size worth publishing.
∅Longitudinal compoundingThe claim that session N+1 costs less and knows more than session N has no longitudinal study behind it yet, it is the thesis Phase 6 exists to test.

Phase 6 · the plan

How The Compounding Claim Gets Tested.

The thesis (session N+1 costs less and knows more than session N) is falsifiable, and the fixtures are specified so the result is not negotiable after the fact.

01Session N+1 vs NSame repository, same task class, repeated sessions. Measure uncached tokens, turns to completion and verified outcome, the compounding curve, not a single retrieval score.
02Injection precision and regretOf the memory injected, how much was used; of what should have been injected, how much was missed. Knowledge regret is the metric that punishes selective memory.
03Fixture protocol: pre-registeredFail→fix headless, denial handling, tool-only exploration, producer-off baselines. An uncomputable metric reports UNKNOWN, never zero.
04Negatives stay publishedThe retired compression levers and the null claim-annotation result set the precedent: anything that loses gets a row and a date, not a deletion.

For evaluators

Four Questions For Any Memory Claim.

Including ours. These are the questions that separated our publishable results from the ones that stayed in the drawer.

01What is n?A single session, a hundred nodes, or 999 sampled injections are three different kinds of claim. If the number is missing, nothing else matters.
02Who verified the outcome?Model self-reports and retrieval scores are not outcomes. A verifier that fails on the pristine baseline and passes on a known-good solution is the minimum bar.
03Are the queries circular?If the test set is generated from the expected answers, the benchmark validates plumbing, not quality. Ours does exactly that, and we label it as circular.
04Mechanism or outcome?'We store typed nodes and decay them' is a mechanism. 'Session 50 costs 40% less than session 1' is an outcome. Vendors blur these constantly; ask which one a sentence is.

Questions

Asked About Memory Benchmarks.

Q

Does Anvaya have a memory benchmark score?

No public score. We have mechanism evidence (retrieval, packing, decay, drift, utility updates) and a set of measured baselines with shortcomings we publish. No LongMemEval or LoCoMo run exists, so there is no score to quote.

Q

What is the injection-utility number?

The mean utility of memories actually injected into live sessions: 0.189 and 0.186 in two readings (918 and 999 samples). Roughly 81% of injected memory went unused. We publish it as an open problem (the reason the experience-driven work exists) not as a win.

Q

Why publish a circular baseline and a null result?

Because the alternative is quoting them as if they were evidence. The synthetic benchmark generates queries from the expected answer, so it validates plumbing, not quality. The A/B tested claim annotation and found no effect, with extra tokens. Both stay in the record.

Q

When will real memory benchmarks land?

Phase 6 of the benchmark program: session N+1 versus N cost and knowledge, injection precision, knowledge regret, and the compounding curve, run against the Experience Benchmark Protocol's fixtures (fail→fix headless, denial handling, tool-only exploration, producer-off baselines), where an uncomputable metric reports UNKNOWN rather than zero.

Q

How should I evaluate a memory claim from anyone?

Ask four questions: what is n, who verified the outcome, is the query set generated from the expected answers (circular), and is the claim about mechanism or outcome? A vendor quoting a benchmark it authored without n and a verifier is quoting marketing.

Run Agents That Fit On Your Laptop.

25.6 MB median RSS. 25 agents ran in parallel on a Core 2 Duo with 4 GB RAM. Hundreds on your machine. Zero cloud required on the Ollama path.

Requires Rust/cargo to build from source. Linux and macOS today, Windows not yet supported. Pre-1.0, public beta. Pricing TBD.