Open Trials  ·  EXPECTED · Q3 2026

Benchmarks On Paper.Measured Numbers Coming Q3.

The competitor numbers below are observable from public documentation, these tools are stateless, so their behavior is predictable. Anvaya’s numbers are projected from the compression ratios measured in our product repos (12-25x per node type), not from a benchmark run. Full reproducibility package arriving end of Q3 2026.

Methodology

How We’ll Measure.

Most benchmarks measure throughput. We measure compounding , whether a tool gets better at your specific codebase the more you use it.

01What we measureNot tokens per second or latency. We measure how much context each tool needs per query over N sessions, and whether that number improves. No other tool in this comparison improves with use.
02What’s measured vs projectedCompetitor numbers are observable from public documentation and code, these tools are stateless, so their behavior is predictable. Anvaya numbers below are projected from the compression ratios measured in our product repos (12-25x per node type, verified in CONTENT.md), not from a benchmark run.
03ReproducibilityMeasurement scripts will be published alongside final results. The methodology is documented. Run them yourself against the same codebase and verify. We believe benchmarks are only useful if you can check them.

The Industry

Every Tool Is The Same On Day 1 And Day 100.

Claude Code, Aider, Copilot, OpenCode, jcode, and Codex CLI are stateless. They don’t improve with use. The context they need on session 50 is the same as session 1. These numbers are observable, no benchmark required.

ToolS1 contextS10 contextMemorySavings at S10
Claude Code~50K~48KNone4%
Aider~45K~42KRepo-map (static)7%
Copilot CLI~48K~47KWorkspace index2%
Opencode~50K~50KNone0%
jcode~50K~49KNone2%
Codex CLI~50K~50KNone0%
VerdictMaximum savings of ~7% (Aider). No tool crosses 10%. Stateless tools don’t compound.

* Competitor figures are estimated from public docs and code inspection. These tools have no persistent memory layer, so their behavior across sessions is predictable and stable.

Expected To Improve With Use.

These numbers are derived from the compression ratios measured in our product repos (12-25x per node type, verified against the Mind README) and projected across session milestones. They are on-paper estimates, not benchmark results. Final measured numbers arrive Q3 2026.

StagePacked contextMemory stateProjected savingsBasis
S1 · Blank slate~50KEmpty graph0%Same as any tool, no memory yet
S10 · Early memory~20K (est.)First nodes captured~60% (est.)Based on 12-25x compression per node type
S20 · Pattern recognition~12K (est.)Causal threads forming~76% (est.)Threads + semantic overlap pull related context
S50+ · Senior engineer~4K (est.)Calibrated utility scoring~92% (est.)Utility scoring + decay keeps only what earned its place
12-25xMeasured compression per node type (Mind README §1)
2-4KPacked context vs ~50K raw (Mind README §1)
500+Tests across the CLI workspace
0Other tools that improve with use

Best Case vs Anvaya At Session 50.

The best any competitor achieves is Aider’s ~7% savings at session 10, from a static repo-map, with no learning loop. Anvaya at session 50 is projected to reach ~92% savings from a calibrated, experience-driven graph. These are on-paper estimates based on measured compression ratios.

Best Competitor (Aider, S10)
~42K tokens per query
Static repo-map, no learning
7% savings, flat from here on
7%
max token savings, plateaued
✓ Anvaya (S50+, projected)
~4K tokens per query
Calibrated, experience-driven graph
~92% savings, still improving
~92%
projected token savings

* Anvaya projected from 12-25x compression per node type (verified) applied across session milestones. Actual benchmark results pending.

When Do The Real Numbers Land?

We’d rather show you measured numbers than promise them. Here’s exactly what’s coming and when.

01Q3 2026, Test harnessBenchmark harness finalized: test projects selected, prompts standardized, measurement scripts written. This is what we run the trials with.
02Q3 2026, Open trialsHead-to-head trials against Claude Code, Aider, Copilot CLI, OpenCode, jcode, and Codex CLI. Same prompts, same codebase, different memory states (S1, S10, S50).
03Q3 2026, Published resultsFull results on this page. Raw data, scripts, and methodology published alongside. Every number reproducible.
04Q4 2026, Public betaSource code available. Run the benchmarks yourself. If the numbers don’t match, we want to know.

Stop Starting From Zero.

One binary. 11+9 Rust crates. 545 tests. Hand-written HNSW index. Three transport modes. Four providers, Ollama, Anthropic, OpenAI, Siemens. Zero API keys required to start. Mind remembers everything after the first session.