Phase 1 — Agent RSS  —  FULL REPORT — UPDATED 2026-10-04

Every Number,Every Failure.

One cumulative list of eleven harnesses. Seven measured on the opencode gateway on 2026-10-04 — 495 unique runs (54 anchor + 441 matrix) across three models and ten task classes, 469 verified, plus copilot (BYOK, 54 runs, 54 verified, two models) — and the four earlier incumbents from 2026-09-16 (238 runs, 228 verified). Median own-process RSS, whole-tree cost, kernel footprint and wall time per harness and per model.

787cumulative runs, both lanes
751/787verified completions
0timeouts (idle-kills rerun, published)
483corpus files — hash 205507698927dd67

All harnesses — cumulative, 787 runs

Own RSS, Tree, Footprint, Wall.

Median of the harness’s own process, with observed range. Tree adds every helper it spawns; footprint is the kernel’s phys_footprint (leader-only). The measured column carries the date because the two lanes ran on different gateways and model sets — 2026-10-04 rows are n=3 per cell (unique cells, a rerun replaces its attempt), 2026-09-16 rows n=1 plus n=3 on the two heaviest classes. Compare footprint to footprint, tree to tree, never across.

HarnessMeasurednVerifiedOwn RSS median (min–max)Tree medFootprint medAux medWall med
anv (--no-mind)2026-10-049090/9023.4 MB (19–30)23.6 MB13.0 MB0.0 MB26.2 s
goose2026-10-048181/8196.5 MB (96–111)96.8 MB61.0 MB0.0 MB26.0 s
pi2026-10-048181/81129.3 MB (116–135)129.5 MB92.0 MB0.0 MB26.2 s
hermes2026-10-048181/81201.4 MB (182–225)204.8 MB174.0 MB2.2 MB38.5 s
aider2026-10-048155/81255.7 MB (217–301)265.9 MB232.0 MB0.1 MB18.3 s
kimi-code2026-10-048181/81393.5 MB (359–444)393.5 MB347.0 MB0.0 MB26.1 s
copilot2026-10-045454/54297.8 MB (288–330)298.1 MB14.0 MB35.8 MB33.3 s
jcode2026-09-164444/4445.0 MB (42.6–49.6)45.2 MB11.0 MB0.0 MB25.1 s
codex2026-09-161717/1797.2 MB (91.1–291)127.6 MB22.0 MB36.5 MB68.0 s
claude2026-09-163432/34432.1 MB (386–500)439.6 MB180.5 MB27.1 MB55.9 s
opencode2026-09-164848/48582.3 MB (523–655)585.9 MB440.5 MB7.2 MB21.3 s

Idle floors (quiet-box anchor, trivial-prompt median own RSS): anv 19–20 — goose 96 — pi 116–128 — hermes 196–200 — aider 250–253 — kimi 380–390 — jcode 43 — codex 93 — claude 390 — opencode 537 MB. copilot 291 MB (BYOK, opencode gateway). long_session is absent for five harnesses (no documented headless session-resume form — documented skip). Node (opencode, claude) and Python (aider) carry runtime floors no context management removes. ±6% measured session-to-session noise; differences below that are not claimed.

Per model — all harnesses

Same Harness, Different Models.

Memory is a runtime property, not a model property — each harness’s median barely moves across its models. MB, median own RSS. The three model columns are each lane’s own model set (the lanes ran on different gateways, so the columns are not a shared axis).

HarnessMeasuredmodel Amodel Bmodel CVerified
anv2026-10-0424232481/81
goose2026-10-0499969772/72
pi2026-10-0412912913072/72
hermes2026-10-0420120120672/72
aider2026-10-0425926025746/72
kimi-code2026-10-0439440139872/72
copilot2026-10-04n/a29629954/54
anv2026-09-1626.823.625.049/49
jcode2026-09-1645.044.845.144/44
codex2026-09-1697.2n/a97.217/17
aider2026-09-16248.5248.9249.538/46
claude2026-09-16438.6n/a428.432/34
opencode2026-09-16588.8573.8582.248/48

2026-10-04 lane: deepseek-v4.1-flash / longcat-2.5-preview-free / mimo-v2.6-flash (opencode gateway) — copilot ran on two of the three (longcat and qwen-3.8-27b via BYOK, shown in the model B and C columns; no deepseek cell). 2026-09-16 lane: deepseek-v4.1-flash / glm-5.3-flash / qwen-3.8-27b (operator gateways). codex ran on qwen-3.8-27b only (it requires the Responses API); claude could not use glm-5.3-flash (opencode’s Anthropic endpoint returns 500). Every skip is printed by the runner and recorded.

Findings

What the field says.

01Model does not move memoryPer-harness medians move a few MB across the three models — the anchor spans 19–20 MB for anv and 382–390 MB for kimi. Memory is a runtime property.
02Runtime ordering holds looselyRust lightest, then Node, then Python — but goose (Rust) is ~5x anv, and pi (Node) is lighter than both Python harnesses. A harness's own engineering moves the number as much as the language.
03anv stays the floor13 MB kernel footprint, no daemon, no helper process, on this gateway and these models. 90/90 verified; the two idle-killed longcat attempts were rerun and passed.
04copilot sits mid-field297.8 MB median on the same gateway — a Node harness with a ~36 MB resident helper, between hermes (200) and claude (432). 54/54 verified across two models and nine tasks; only long_session is skipped (no headless resume). Its 14 MB kernel footprint mirrors anv’s shape: the cost is in the runtime, not a daemon.
05hermes and kimi carry a resident helperOn the heavy tasks each carries a ~190 MB / ~350 MB second process the kernel footprint column makes visible (hermes fp 174 vs own 201; kimi fp 347 vs own 393). Compare footprint to footprint, tree to tree.
06aider + these flash models refuse file taskswide_read and aggregate_scan verified 0/9 across all three models: the model answered conversationally instead of using file tools. Intermittent on the other four tasks. A harness+model-lane property, published with verified:false, not hidden.
07One cumulative listTwo measurement dates appear because four harnesses have not been re-run this round — the 2026-09-16 rows carry their own date, gateway and model set. The aider bridge (253 vs 242 MB idle) sits inside the ±6% noise, so the ordering carries over.

Failures — both lanes

751 Of 787 Verified. Every Miss Listed.

Zero timeouts across both lanes. The aider refusals and the gateway-stall attempts stay in the records with their output, not rewritten.

HarnessCellWhat happenedMeasured
anvlarge_read (longcat, 2 attempts)Idle-killed at 263 s, then rc=1 at 448 s with 253 s of first-token wait. The proxy log shows an upstream read timeout at the same moment — a gateway stall, not anv. The reruns passed and are the published cells.2026-10-04
gooselarge_read (mimo, 1 attempt)Idle-killed at 275 s; the rerun passed in 28.6 s. Gateway stall, same signature.2026-10-04
aiderwide_read / aggregate_scan, all 3 models0/9 verified on both tasks. The model replied “please add the records/ files to the chat” instead of using file tools — reproducible across every trial and model. build_pkg (5/9), noisy_tool (7/9), pressure_read (8/9), edit_tool (8/9) fail intermittently, same pattern. aider is still fastest on wall time (18.3 s) because refusing is fast — the ⚠ marks are the point.2026-10-04
kimi-codeanchor lane, 1 attemptA runner bug pointed kimi at a dead proxy port before it was fixed (Anvaya-Benchmark commit 755a76b). Archived under records/archive/port-bug-attempts-* with the root cause; the cell was re-verified clean.2026-10-04
claudepressure_read (qwen, 2 attempts)Auto-compact guard aborted at 843 s and 546 s; two other attempts passed (628 s, 85 s). Intermittent at a 128K window, not a deterministic ceiling.2026-09-16
aiderwide_read / aggregate_scan (glm, deepseek, qwen)No filesystem search — read the repo-map summary and asked the user to paste files; on aggregate_scan the answer file never appeared. Same capability gap on the earlier model lane.2026-09-16

Stability — n=3 replication (2026-09-16 lane)

What Was Real, What Was One-Off.

The two heaviest classes (pressure_read and edit_tool) were re-run at n=3 across every harness/model pair in the 2026-09-16 lane. Every earlier outlier became a classification.

01codex edit_tool +214% — one-offA single serial-anchor run peaked at 291 MB; three follow-ups landed at 92–146 MB. The cell is not reproducibly memory-hungry.
02claude pressure_read abort — intermittent2 of 4 attempts aborted (843 s, 546 s) with the auto-compact guard; 2 passed (628 s, 85 s). A coin-flip at a 128K window.
03opencode rc=1 — transient, still verifiedOne attempt in four exited non-zero after writing the correct answer; three follow-ups exited cleanly. Recorded as verified with clean_exit=false, not treated as a pattern.

Rerun It. Or Dispute The Method.

Everything above is derived from the published records by the report script. Criteria, corpus generator, runner, and every JSONL record (including failed attempts and superseded runs) are public.

records/549 runs (2026-10-04, incl. copilot) + 238 runs (2026-09-16), append-only JSONL — github.com/NaMan6122/Anvaya-Benchmark
205507698927dd67corpus v3 hash (483 files) — both lanes
a7b1ee2ad176anv build fingerprint (pre-release --no-mind)
MIT — CC BY 4.0code — data licenses

Claim policy: what is supported is memory, speed and verified completion on these models, corpus and machine, not a quality verdict, and not “the world’s lightest” until the mode ships, the comparison is run on default configurations, more harnesses are covered, and a non-macOS platform is tested. Disclosure: anv is our product.

Run Agents That Fit On Your Laptop.

25.6 MB median RSS. 25 agents ran in parallel on a Core 2 Duo with 4 GB RAM. Hundreds on your machine. Zero cloud required on the Ollama path.

Requires Rust/cargo to build from source. Linux and macOS today, Windows not yet supported. Pre-1.0, public beta. Pricing TBD.