Does Rust make the agent smarter?
No. It makes the agent cheaper to keep running and easier to reason about operationally. Model quality comes from the model; harness quality is measured in latency, memory, and whether work actually completes.
Blog · 2026-09-16 · 7 min
12 ms to first frame, 4.2 MB idle, 25.6 MB median under load, what runtime-free buys, what it costs, and why we measured both.
Direct answer
An agent lives beside your editor and your shell for hours. Its baseline cost (memory, startup, idle CPU) is a tax you pay all day whether or not you are asking it anything. Removing the language runtime is the difference between a tool that idles in single-digit megabytes and one that costs half a gigabyte to say READY.
01
Anvaya CLI is a single static Rust binary: 12 ms to first TUI frame, 17 ms for --version, 4.2 MB kernel footprint at idle (12.4 MB ps RSS), 2.8 MB in headless setup, and roughly 11 MB held across a 17.6-hour session. macOS footprint is the honest instrument: RSS includes shared mapped pages and hides compressed ones, so we publish both and the protocol behind them.
In Phase 1 of our benchmark program (six harnesses, three gateway models, ten task classes, 238 runs, external verification) the harness measured 25.6 MB median own-process RSS with a 21 MB idle floor. The Node-based harnesses measured 432 MB and 582 MB; the Python one 249 MB. Their runtimes set floors no context-management improvement removes, which is exactly why the number matters: it is what a user pays before any work happens.
02
Predictable latency and memory without a garbage collector pausing between a tool call and its result. No package tree to audit or pin. No interpreter version matrix to support. A binary you can ship into a minimal container or an air-gapped machine and trust to behave the same way it did on your laptop.
It also buys honesty in the measurement: with no runtime, the harness's memory is the harness's memory. There is no shared V8 isolate to split across processes, no ambiguity about who owns the baseline.
03
Compile times and a smaller extension ecosystem. Rust also raises the bar for contributors, unsafe code, lifetimes, and a type system that makes the wrong abstraction feel expensive. For a process that must stay resident for hours and never surprise the user, that trade is worth it; for a weekend script, it emphatically is not.
And a static binary is not automatically light: the persistent-memory daemon is a separate process that warms to roughly 50 MB with its embedding model mapped. We treat that as its own subsystem with its own budget, and it can be left out entirely.
04
We do not claim 'the world's lightest agentic harness.' That needs a shipped default mode, a default-configuration comparison, more than six harnesses, and at least one non-macOS platform. The ranking we publish is scoped to six harnesses on one machine, with the raw records public so anyone can dispute the method.
Questions
No. It makes the agent cheaper to keep running and easier to reason about operationally. Model quality comes from the model; harness quality is measured in latency, memory, and whether work actually completes.
Startup, usually. Throughput depends on the workload. The win here is a low, predictable baseline for a long-lived process, not raw compute.
The Phase 1 report on our benchmarks page has per-harness and per-model medians, idle floors, whole-tree cost and the full methodology; the raw dataset is public.
25.6 MB median RSS. 25 agents ran in parallel on a Core 2 Duo with 4 GB RAM. Hundreds on your machine. Zero cloud required on the Ollama path.
Requires Rust/cargo to build from source. Linux and macOS today, Windows not yet supported. Pre-1.0, public beta. Pricing TBD.