Pillar · Coding Agents

Coding Agents: Pick By Shape, Not By Demo.

IDE, terminal, cloud or fleet, the shape decides what the tool can finish. Here is what each one is for, and the eight questions that separate a demo from something you can leave running.

Direct answer

A coding agent owns a task across files: it reads, edits, runs commands and iterates until an external check says the work is done. The useful categories are shape, not brand, IDE agents keep a human in the visual loop, terminal agents run tasks anywhere a shell exists, cloud agents trade sovereignty for queues, and supervised fleets partition large jobs across isolated workers. Anvaya CLI is a terminal agent: single static Rust binary, 21 native tools, 8 screens, headless CI mode, a local memory graph, and a sandbox enforced independently of approval. Phase 1 of our open benchmark program measured it at 25.6 MB median own-process RSS across 238 runs with 228/238 externally verified completions. Choose your agent on verification, confinement, exit codes and memory, the demo is the easy part.

The Shapes

Four Ways To Deploy An Agent.

Most disappointment with coding agents comes from using the wrong shape, not the wrong model.

01IDE agentsLive in the editor, edit with visual diffs, and keep a human in the loop by default. Best for interactive work you want to watch: feature slices, test writing, refactors with review. The trade is scope, the editor session ends when you close it, and automation stops at the window.
02Terminal agentsOwn a task from the shell: read, edit, run commands, iterate, and exit with a code. Best for multi-file work, headless jobs and anything that must run on a machine without a GUI. Look for a real exit contract and a sandbox you can audit.
03Cloud / async agentsRun in vendor infrastructure, often triggered from an issue or PR. Best for queued work where latency is irrelevant. The trade is sovereignty: your code leaves the machine, and cost scales with vendor inference.
04Supervised fleetsMany agents under a supervisor with task partitioning, isolated workspaces, external verification and a merge gate. Best for partitionable work, migrations, audits, defect fan-outs. The trade is orchestration: without a verifier, a fleet automates the production of plausible diffs.

Evaluation

Eight Questions Before You Commit.

Each one has a failure mode you will otherwise discover on a deadline.

01Objective verificationDoes the tool run something external (tests, a linter, a build) before calling a task done? A model self-report is not a completion signal.
02Confinement you can readIs there a sandbox policy for reads, writes, exec and network, enforced independently of approval? Does auto-approval widen it? (It should not.)
03An exit contractFor CI, headless mode needs pipeable I/O and distinct exit codes for success, error and approval-denied. Without them, a pipeline cannot tell failure from a missing gate.
04FootprintAn agent lives beside your editor for hours. Measure idle memory and startup from a release build, and print both kernel footprint and RSS, they disagree.
05Memory that is inspectableIf the tool claims to remember, ask where the data lives and in what format. Local files you can query and delete beat an opaque cloud profile.
06Provider freedomCan you run local models with no account, and bring your own keys for frontier ones? Model lock-in is workflow lock-in.
07Isolation for parallel workSubagents should get fresh context windows; fleet workers should get isolated workspaces. 'Parallel' on one working tree is a merge conflict with extra steps.
08ObservabilityCan you see per-agent state, token spend and process memory while it runs? Retrofitting observability after a runaway job is called incident response.

At A Glance

What Each Shape Can Finish.

Same model, different ceiling, the shape sets the scope before the model ever runs.

DimensionIDE agentTerminal agentSupervised fleet
ScopeEditor sessionA task to completionA partitionable job
Human loopEvery edit, visualApproval policy per toolReview the integration branch
Runs without youNoYes, headless with exit codesYes, supervised with liveness
IsolationOne working treeSandbox policy per toolShared / worktree / clone per worker
VerificationReviewer's eyesTests and builds as gatesObjective per-task verifier first
Across sessionsTranscriptOptional local memory graphShared memory across workers

Questions

Asked About Coding Agents.

Q

What is a coding agent?

A tool that runs a plan-act-verify loop with real tools (reading files, editing them, running commands) until a task is complete or a limit is hit. It differs from autocomplete (single-file prediction) and chat (advice with no execution) by owning the loop and, importantly, by needing verification.

Q

IDE agent or terminal agent?

Pick by where the work lives. IDE agents win on interactive editing and visual review; terminal agents win on multi-file autonomy, headless CI and remote or air-gapped machines. Many teams run both and bridge them with shared instruction files and a common memory layer.

Q

How do I evaluate one without wasting a week?

Give every candidate the same small, verifiable task on your own repo: one bug with a covering test. Compare whether the work completed, how many tokens it took, what it did when the test failed, and what it left behind. Demos are cheap; completion on your code is the signal.

Q

Do I need a fleet of agents?

Only if your work partitions. Swarms pay off on many independent modules with objective tests; for tightly coupled edits, one agent plus a checkpoint is faster and cheaper. Start with one agent and a memory layer; add workers when the bottleneck is queue depth, not capability.

Q

Can a coding agent run fully offline?

Yes on the local path: a local model runtime such as Ollama needs no account, and a terminal agent with a local memory store never has to send code anywhere. Cloud providers are optional escalations, not prerequisites.

Run Agents That Fit On Your Laptop.

25.6 MB median RSS. 25 agents ran in parallel on a Core 2 Duo with 4 GB RAM. Hundreds on your machine. Zero cloud required on the Ollama path.

Requires Rust/cargo to build from source. Linux and macOS today, Windows not yet supported. Pre-1.0, public beta. Pricing TBD.