Benchmark Records  ·  MEASURED · UPDATED 2026-09-17

Every Record,Including The Losses.

Measured rows only, each with its sample size, configuration and artifact path. Retired levers, null A/B results and an absent benchmark (no LongMemEval or LoCoMo run exists) are published with the same prominence as the wins. Phase 1 of the harness campaign has its own full report.

−11.0%median uncached tokens, micro-fold (n=5)
97.6%dead page tokens in live traces (251 rounds)
7.55MBN=32 supervisor soak, max RSS
3/3SaaS apps, production swarm E2E

Token economics · measured

What Actually Moved The Bill. And What Moved Nothing.

The shipped default is conservative: a context guard, macro compaction, and one promoted prompt lever. Micro-fold is the strongest experimental lever so far and stays gated until the promotion thresholds clear.

MeasurementResultn / scopeSource
Live trace archaeology98–99% cached · 97.6% dead page tokens · 0 artifact re-fetches251 rounds · 2026-09-12 · live sessionsTrace study · 2026-09-12
Micro-fold (artifact offload)−11.0% median uncached tokens · verification 19/19n=5 pairs · mimo-v2.5 · threshold 4,000Paired run · 2026-09-12
Headless fix tasks (real CLI)−11%…−21% uncached · 24/24 verified pairs · RSS flat48 runs · 4 models · --no-tui --yoloMulti-model run · 2026-09-13
Inline tool schema (promoted)−4,725 tokens/round · success 160/160 · pooled validity +2.71pp [−1.53, +6.94]320 live sessions · both armsSession audit · 2026-09
Decision-dense retention8/8 prompt-only decision codes retained through 2–3 forced macro compactionsn=3 · 8K context limit · baseline armRetention run · 2026-09-14
Stage-A/B compaction (shipped default)summarize at 80% · hard-drop at 95% · keep last 5 turnsshipped default · pair-safeShipped behavior

Uncached tokens mean input tokens not served by a provider cache, the quantity a memory or compression layer can actually change. Full configurations and scripts are published with the public Phase 1 dataset where applicable; the rest are internal records, summarized here with scope and dates.

Retired & negative results

We Published The Ones That Lost.

A compression idea that raises uncached tokens or thrashes compaction is a finding, not a footnote. These are kept with their configurations and dates so nobody spends a quarter re-running them.

LeverMeasured resultn / scopeStatus
Meso-fold LRU budgetuncached +7.8%, retires mid-history folding as a class5 pairs (2 valid) · mimo-v2.5Retired · 2026-09-12
Meso-fold dead-bodyuncached +24.5%5 pairs (2 valid) · gate correctRetired · 2026-09-12
WorkingSet synthetic-user placementmedian uncached 35,041 vs 9,645 baseline · 28-compaction thrash3 trials/arm · 12K limitWithdrawn · 2026-09-12
WorkingSet system-suffix placement~10× uncached (92,691 / 107,713 vs 9,982) via prefix-cache invalidation3 trials/armWithdrawn · 2026-09-12
Model-directed context releaseunguided engagement 0 in every pilot · guided apply works; uncached +25%/+53% at 8 KB stale reads; batched variant removes the round tax, economics still open216 runs · 6 pilot rounds · 5 modelsDefault-off · 2026-09-17
Context guard pre-flight overestimatewarning overestimated by 20.7× (1,025,789 vs 49,503 observed)harness campaign · 16K armFixed · 2026-09-14
Live injection utility baseline0.189 → 0.186 mean (≈81% of injected memory unused), the open problem the experience stack exists to fix918→999 samples · 2026-08-31Open finding

Swarm · measured

Supervisor Cost, Fleet Behaviour.

The swarm layer's own overhead is measured separately from model cost: supervisor RSS, spawn throughput and status latency. The 256-agent soak is deferred by product decision and is not claimed.

RunResultn / scopeSource
N=32 supervisor soaksupervisor RSS max 7.55 MB (median 7.41) · spawn wall 396 ms · status p50 4.2 ms / p99 14.9 ms · pass=true32 dummy agents · 15 sSoak record · 2026-09
Production E2E (mvp bundle)3/3 SaaS apps delivered · integration suites 35/43/38 tests green · conflicts [] · parked [] · verdict failures 02026-09-13 · worktree placementE2E record · 2026-09-13
Autonomy baselineall modes 100% first-passn-of-3 × 3 modes · 9 runs · 6-entry corpusInternal baseline · 2026-09
Supervisor design budgetstargets: ≤25 MB RSS @256 idle agents · ≥100 spawns ≤10 s · status p99 ≤5 ms · ≤64 KB/agent bookkeepingdesign targets, not all verifiedDesign targets
Live liveness boundsheartbeat every 10 s from a timer task · suspect 30–60 s · lost >60 s · kill-tree ≤5 slive-verified at N=32Shipped behavior

SWE-bench-style adapter · measured

Offline Corpus. No Vendor Arms Yet.

The evaluation adapter materializes repositories, applies arm configurations, runs the agent and scores the resulting patch with pytest-style parsing. It is benchmark infrastructure, not part of the shipped binary, and these are resolve counts, not a leaderboard position.

RunResultn / scopeSource
swe-mini round 112/12 resolve4-instance corpus · baseline vs nomination · 2 modelsEval record · 2026-09-17
swe-mini round 216/16 resolve · nomination applied 2/2 and 1/1 · micro-fold dominates+19.7 KB bigmodule (micro-offloadable)Eval record · 2026-09-17
swe-mini round 38/8 resolve+8 KB stalewatch (sub-threshold post-edit facts)Eval record · 2026-09-17
Adapter scopeOffline SWE-bench-style corpus, exit-code scoring with pytest-style per-test parsing; container isolation pending; no vendor armsbenchmark infrastructure · never in the shipped binaryScope note

Memory & experience · status

The Honest State Of Memory Measurement.

The memory layer has mechanism evidence, not outcome benchmarks yet. These rows exist so the distinction cannot be missed, and so the missing benchmark is named.

MeasurementResultn / scopeStatus
Synthetic ranking baselinenDCG@5 0.648–0.671 · MRR 0.834–0.875, explicitly circular (queries generated from expected nodes)100 seeded nodes → 53 queries · fixed seedStudy · 2026-08-31
Observed ranking baselineinsufficient positives: 185 feedback rows, 2 sessions, 4 query hashes, 0 positive observationslive store · 2026-08-31Study · 2026-08-31
Claim annotation A/Bnull result: citation rate 0.0% both arms · treatment +32% injection-block tokensn=57 observations · 3 paired promptsNull result · 2026-08-31
LongMemEval / LoCoMoNo run exists. Do not cite one from us.NoneAbsent
Compounding curve (S1→S50)Design thesis, no longitudinal study yet; Phase 6 of the benchmark program is the measurement.NoneNot yet measured

Citing A Record. Or Disputing One.

Cite the measurement, its n, its date and its scope, never a headline number without them. The harness Phase 1 dataset is public at github.com/NaMan6122/Anvaya-Benchmark (MIT code, CC BY 4.0 data). Context, swarm and memory records are internal during beta and summarized here with their dates and sample sizes.

↻Records update with each registry revision; stale rows are corrected, not deleted
nEvery row carries its sample size; n=1 is labeled as such
✕No vendor arms yet; no cross-vendor token claim is made
∅No LongMemEval/LoCoMo run exists, do not cite one from Anvaya

Run Agents That Fit On Your Laptop.

25.6 MB median RSS. 25 agents ran in parallel on a Core 2 Duo with 4 GB RAM. Hundreds on your machine. Zero cloud required on the Ollama path.

Requires Rust/cargo to build from source. Linux and macOS today, Windows not yet supported. Pre-1.0, public beta. Pricing TBD.