Benchmark Records · MEASURED · UPDATED 2026-09-17
Every Record,Including The Losses.
Measured rows only, each with its sample size, configuration and artifact path. Retired levers, null A/B results and an absent benchmark (no LongMemEval or LoCoMo run exists) are published with the same prominence as the wins. Phase 1 of the harness campaign has its own full report.
Token economics · measured
What Actually Moved The Bill. And What Moved Nothing.
The shipped default is conservative: a context guard, macro compaction, and one promoted prompt lever. Micro-fold is the strongest experimental lever so far and stays gated until the promotion thresholds clear.
Uncached tokens mean input tokens not served by a provider cache, the quantity a memory or compression layer can actually change. Full configurations and scripts are published with the public Phase 1 dataset where applicable; the rest are internal records, summarized here with scope and dates.
Retired & negative results
We Published The Ones That Lost.
A compression idea that raises uncached tokens or thrashes compaction is a finding, not a footnote. These are kept with their configurations and dates so nobody spends a quarter re-running them.
Swarm · measured
Supervisor Cost, Fleet Behaviour.
The swarm layer's own overhead is measured separately from model cost: supervisor RSS, spawn throughput and status latency. The 256-agent soak is deferred by product decision and is not claimed.
SWE-bench-style adapter · measured
Offline Corpus. No Vendor Arms Yet.
The evaluation adapter materializes repositories, applies arm configurations, runs the agent and scores the resulting patch with pytest-style parsing. It is benchmark infrastructure, not part of the shipped binary, and these are resolve counts, not a leaderboard position.
Memory & experience · status
The Honest State Of Memory Measurement.
The memory layer has mechanism evidence, not outcome benchmarks yet. These rows exist so the distinction cannot be missed, and so the missing benchmark is named.
Citing A Record. Or Disputing One.
Cite the measurement, its n, its date and its scope, never a headline number without them. The harness Phase 1 dataset is public at github.com/NaMan6122/Anvaya-Benchmark (MIT code, CC BY 4.0 data). Context, swarm and memory records are internal during beta and summarized here with their dates and sample sizes.
Run Agents That Fit On Your Laptop.
25.6 MB median RSS. 25 agents ran in parallel on a Core 2 Duo with 4 GB RAM. Hundreds on your machine. Zero cloud required on the Ollama path.
Requires Rust/cargo to build from source. Linux and macOS today, Windows not yet supported. Pre-1.0, public beta. Pricing TBD.