Pillar · Local Swarms

A Coding Fleet That Never Leaves The Machine.

Process-per-agent swarm orchestration with local models, local experience and local artifacts. No hosted orchestrator, no account, no code egress, with the GPU bottleneck stated plainly.

Direct answer

Local swarms run multi-agent coding entirely on your own hardware: a local supervisor process, agent processes, local model endpoints and an on-disk experience graph. Anvaya is built for this shape, and built light enough that the fleet fits on the machine you already own: 25 sessions ran in parallel on a Core 2 Duo with 4 GB of RAM, no crashes, no faults, and a 32-agent soak cost 103 MB total agent RSS. The supervisor coordinates 24 agents by default (256 hard cap), each in a shared, worktree or cloned workspace, with all run state under .anvaya/swarm/. Loopback model endpoints are keyless, so the fleet needs no account; sandbox policy can deny network at the tool layer; and the Mind graph it consults lives in SQLite + HNSW under .anvaya/mind/. The honest limit: we have measured the supervisor and the harness footprint (N=32 soak, 7.55 MB max supervisor RSS) but not local-inference throughput across a fleet, so size concurrency to your GPU, not to the 256 cap.

Local By Construction

Six Properties, Not A Promise.

01No hosted control planeThe supervisor is a local process supervising local agent processes. Run state, artifacts, integration reports and halt reports all live under .anvaya/swarm/ in your repository, inspectable, diffable, removable.
02Keyless local inferenceLoopback model endpoints (localhost, 127.0.0.0/8, ::1) are treated as keyless: no API key is built or sent. Point workers at Ollama or any local OpenAI-compatible runtime and the fleet needs no account.
03Memory stays on diskEach worker can consult the same local Mind graph (SQLite + HNSW under .anvaya/mind/). Recall and recording never leave the machine, and federation between projects is opt-in with attenuation and export filters.
04Isolation without the cloudWorktree and clone placement are git-local: each agent gets its own checkout on disk. Sandbox policy profiles (project, computer, strict) constrain reads, writes, exec and network at the tool layer.
05Bounded local cost24 concurrent agents by default, 256 hard cap, 20/s spawn rate, a fixed 512-sample telemetry ring per agent, and an adaptive sampler that stretches to 15 s when everyone is idle. Token and RSS columns are visible on the swarm board.
06Honest bottleneckParallel inference on one box is GPU-bound. Local swarms scale when the model fits your memory and tasks are partitionable; they do not turn a single 7B model into a data center. We have measured the supervisor, not local throughput.

Sizing

Start With One, Then Scale Out.

011 · Prove the modelServe a tool-calling coder locally (Ollama or any OpenAI-compatible runtime). Measure single-agent latency and memory on your real repo.
022 · Prove the memoryRun anv init, let one supervised agent record typed nodes, and query them back. Memory is the layer that makes the second run cheaper.
033 · Pick isolationUse worktree placement for tasks that touch overlapping areas; shared placement when the partition is clean. Start with the repair caps at their defaults.
044 · Add concurrency slowlyWatch per-agent RSS and tokens on the swarm board. Raise concurrency until latency per task stops paying for itself, then stop.

Questions

Asked About Local Fleets.

Q

Can I run a coding agent swarm fully offline?

Yes, by construction: local supervisor, local agent processes, local model endpoint, local memory, local artifacts. The measured caveat: our own N=32 soak used dummy agents, and the production E2E ran on hosted-gateway models, so offline throughput on your hardware is something you size, not something we have benchmarked.

Q

Which local models work for a swarm?

Any tool-calling model your runtime serves over an OpenAI-compatible or Ollama endpoint. Small models per worker can beat one big model for partitioned tasks; sizing depends on how many concurrent requests your GPU can hold at acceptable latency.

Q

How is this different from an ad-hoc parallel agent setup?

A supervisor with a ledger: partition artifacts, per-task verification commands, liveness heartbeats, repair limits, admission control and an integration branch. Parallel terminals have none of that; when one agent hangs you are the dead-man switch.

Q

Do local swarms leak code?

Not by architecture. Nothing in the swarm path requires network: sandbox policy can deny it, and all artifacts are local. If you choose a cloud model endpoint, that is an explicit per-provider decision, not a default.

Q

What hardware should I start with?

Start with one agent plus memory, measure your model's latency and memory ceiling, then add concurrency. The swarm board's per-agent RSS and token columns are the number to watch; the 256-agent cap is a supervisor limit, not a hardware promise.

Run Agents That Fit On Your Laptop.

25.6 MB median RSS. 25 agents ran in parallel on a Core 2 Duo with 4 GB RAM. Hundreds on your machine. Zero cloud required on the Ollama path.

Requires Rust/cargo to build from source. Linux and macOS today, Windows not yet supported. Pre-1.0, public beta. Pricing TBD.