Phase 1 — Agent RSS — FULL REPORT — UPDATED 2026-10-04
Every Number,Every Failure.
One cumulative list of eleven harnesses. Seven measured on the opencode gateway on 2026-10-04 — 495 unique runs (54 anchor + 441 matrix) across three models and ten task classes, 469 verified, plus copilot (BYOK, 54 runs, 54 verified, two models) — and the four earlier incumbents from 2026-09-16 (238 runs, 228 verified). Median own-process RSS, whole-tree cost, kernel footprint and wall time per harness and per model.
All harnesses — cumulative, 787 runs
Own RSS, Tree, Footprint, Wall.
Median of the harness’s own process, with observed range. Tree adds every helper it spawns; footprint is the kernel’s phys_footprint (leader-only). The measured column carries the date because the two lanes ran on different gateways and model sets — 2026-10-04 rows are n=3 per cell (unique cells, a rerun replaces its attempt), 2026-09-16 rows n=1 plus n=3 on the two heaviest classes. Compare footprint to footprint, tree to tree, never across.
Idle floors (quiet-box anchor, trivial-prompt median own RSS): anv 19–20 — goose 96 — pi 116–128 — hermes 196–200 — aider 250–253 — kimi 380–390 — jcode 43 — codex 93 — claude 390 — opencode 537 MB. copilot 291 MB (BYOK, opencode gateway). long_session is absent for five harnesses (no documented headless session-resume form — documented skip). Node (opencode, claude) and Python (aider) carry runtime floors no context management removes. ±6% measured session-to-session noise; differences below that are not claimed.
Per model — all harnesses
Same Harness, Different Models.
Memory is a runtime property, not a model property — each harness’s median barely moves across its models. MB, median own RSS. The three model columns are each lane’s own model set (the lanes ran on different gateways, so the columns are not a shared axis).
2026-10-04 lane: deepseek-v4.1-flash / longcat-2.5-preview-free / mimo-v2.6-flash (opencode gateway) — copilot ran on two of the three (longcat and qwen-3.8-27b via BYOK, shown in the model B and C columns; no deepseek cell). 2026-09-16 lane: deepseek-v4.1-flash / glm-5.3-flash / qwen-3.8-27b (operator gateways). codex ran on qwen-3.8-27b only (it requires the Responses API); claude could not use glm-5.3-flash (opencode’s Anthropic endpoint returns 500). Every skip is printed by the runner and recorded.
Findings
What the field says.
Failures — both lanes
751 Of 787 Verified. Every Miss Listed.
Zero timeouts across both lanes. The aider refusals and the gateway-stall attempts stay in the records with their output, not rewritten.
Stability — n=3 replication (2026-09-16 lane)
What Was Real, What Was One-Off.
The two heaviest classes (pressure_read and edit_tool) were re-run at n=3 across every harness/model pair in the 2026-09-16 lane. Every earlier outlier became a classification.
Rerun It. Or Dispute The Method.
Everything above is derived from the published records by the report script. Criteria, corpus generator, runner, and every JSONL record (including failed attempts and superseded runs) are public.
Claim policy: what is supported is memory, speed and verified completion on these models, corpus and machine, not a quality verdict, and not “the world’s lightest” until the mode ships, the comparison is run on default configurations, more harnesses are covered, and a non-macOS platform is tested. Disclosure: anv is our product.
Run Agents That Fit On Your Laptop.
25.6 MB median RSS. 25 agents ran in parallel on a Core 2 Duo with 4 GB RAM. Hundreds on your machine. Zero cloud required on the Ollama path.
Requires Rust/cargo to build from source. Linux and macOS today, Windows not yet supported. Pre-1.0, public beta. Pricing TBD.