Do I need a big machine?
No, but you need to size concurrency to your model. The supervisor is lightweight (7.55 MB in the N=32 soak); the bottleneck is inference. Start with one or two agents and add while latency per task stays acceptable.
Blog · 2026-09-18 · 8 min
When a fleet of coding agents beats one, how supervision actually works, and the numbers we have (and have not) measured.
Direct answer
A swarm is not 'more agents.' It is one ledger, one budget, one merge gate, and many isolated workers. The difference between a swarm and four terminals running in parallel is everything that happens when an agent hangs, a task fails, or two edits collide.
01
Four properties separate a supervised swarm from ad-hoc parallel agents. A partitioner turns the job into tasks with typed inputs and outputs. Isolation gives each task its own workspace, shared, git worktree, or full clone. Verification is external: a task is done when a command says so, never when the model says so. And a merge gate decides what lands, on an integration branch, with nothing pushed automatically.
Everything else is supervision detail: liveness heartbeats, retry classification (transient provider errors and crashes retry; approval blocks and budget kills do not), a repair ladder that escalates from same-model retry to a repairer role to widening scope to replanning to a human gate, and a safety envelope that admits waves against disk, file-descriptor and build limits.
02
Swarms pay off when the work partitions cleanly and verification is objective: module-by-module migrations, defect fan-outs, audits across many packages, MVP scaffolding where each slice has its own test command. The measured production run delivered three small SaaS applications this way with zero merge conflicts and green integration suites (35, 43 and 38 tests).
They lose on tightly coupled work. Two agents editing the same symbols is not parallelism; it is a merge conflict with extra tokens. If your task has no objective verifier, a swarm mostly automates the production of plausible-looking diffs, which is worse than slow, careful work by one agent.
03
The supervisor itself is cheap: a 32-agent soak (dummy agents) held supervisor RSS at 7.55 MB max, spawned in 396 ms, and answered status at 4.2 ms median / 14.9 ms p99. Defaults are 24 concurrent agents, 20 spawns per second, 256 hard cap, with a fixed telemetry ring per agent and an idle sampler that stretches its interval rather than polling harder.
What is not measured: 256-agent scale (deferred by product decision), offline throughput for a local fleet, and vendor comparisons. We also do not claim swarms are cheaper per task, partitioning and verification have their own token cost, and the honest way to read that is per merged change, which Phase 6 of our benchmark program will measure.
04
Start smaller than you want to. Validate the manifest offline, run three or four tasks in worktrees, and review the integration branch yourself before merging anything. Watch per-agent RSS and tokens on the board; both are live columns, not logs you reconstruct later. When a task fails, read the repair ladder before overriding it, the cap exists so a confused agent cannot spend your budget learning the same lesson five times.
Questions
No, but you need to size concurrency to your model. The supervisor is lightweight (7.55 MB in the N=32 soak); the bottleneck is inference. Start with one or two agents and add while latency per task stays acceptable.
Never. Approved work lands on an integration branch; pushing is not performed under any policy, including auto-approval.
When verification becomes the bottleneck and per-task latency stops paying for itself. The 256 cap is a supervisor limit, not a hardware promise.
25.6 MB median RSS. 25 agents ran in parallel on a Core 2 Duo with 4 GB RAM. Hundreds on your machine. Zero cloud required on the Ollama path.
Requires Rust/cargo to build from source. Linux and macOS today, Windows not yet supported. Pre-1.0, public beta. Pricing TBD.