Pillar · Automated Development

Agents In The Pipeline, Not Just In The Editor.

Headless runs with a real exit contract, sandbox policy at the tool layer, process monitors for long jobs, and a swarm path for batch work, measured on production E2E.

Direct answer

Automated development means a coding agent executes bounded jobs without a human at the prompt: triage, codemods, failing-test fixes, review passes and batch refactors. Anvaya supports this with a headless contract, anv --no-tui with exit codes 0/1/2 (success/error/approval-denied), a read-only --plan-only pass, and structured output for gating. Long jobs are covered by the monitor tool (up to 8 processes) and --wait-monitors; confinement is enforced by sandbox profiles at the tool layer across all three exec paths; and crash-safe sessions survive pipeline restarts. For fan-out, the swarm supervisor adds per-task verification, repair limits and an integration branch, the measured production E2E delivered three SaaS apps with zero merge conflicts. What is not claimed: unattended production autonomy beyond bounded, verified jobs.

Mechanisms

Eight Levers For Unattended Runs.

Each maps to shipped code; measured records are linked at the end.

01Headless by designanv --no-tui reads a prompt or prompt file, writes to stdout and exits like a Unix citizen: 0 success, 1 error, 2 approval denied. That exit contract is what makes an agent step gateable in CI.
02A read-only planning pass--plan-only blocks write and exec tool calls at dispatch, so a pipeline can ask what would change without risking the tree. Pair it with a required human approval before the mutating run.
03Long jobs that do not hangThe monitor tool (up to 8 armed) tracks long-running processes; run_script executes saved scripts through the same sandbox; --wait-monitors keeps a headless process alive until the last monitored job ends.
04Confinement at the tool layerSandbox profiles (project, computer, strict) gate reads, writes, exec and network for every tool, including all three exec paths (run_command, run_script, monitor). --yolo bypasses approval but does not widen the sandbox.
05Sessions survive crashesA crash-safe session store with retention limits (20 sessions / 512 MiB) lets a pipeline resume rather than restart; anv -s list and --continue address prior runs by code.
06Swarms for batch workFor fan-out jobs, `anv swarm up` runs a supervised fleet with per-task verification, repair limits and an integration branch. The production E2E delivered three SaaS apps with zero conflicts and green integration suites.
07Memory-aware runsHeadless runs can read and write the same local Mind graph, so yesterday's failure is in today's context, the compounding lever, scoped to your measured records.
08Everything observableStructured JSON where it matters (`anv swarm status --json`), append-only run directories under .anvaya/swarm/, per-task logs via tail, and exit codes that distinguish failure from approval denial.

Patterns

Three Jobs To Start With.

01Plan-gate-applyJob 1: run --plan-only and publish the plan. Job 2 (after approval): run the same prompt with the sandbox profile that fits and tests as the verifier. Two steps, one clear gate.
02Nightly triageCron a headless read-only pass over failing tests and open issues; write findings to stdout or a file. Memory-aware: the graph already holds last night's fixes.
03Batch codemod with a swarmPartition by module, place each agent in a worktree, verify with the module's test command, and review the integration branch once. Nothing is pushed.

Questions

Asked About Automation.

Q

Can a coding agent run in CI safely?

With three constraints: read-only planning first (--plan-only), tool-layer sandbox policy, and gated writes where revert is cheap. Anvaya's exit codes separate error (1) from approval-denied (2), so a pipeline can tell the difference between a failed attempt and a blocked one.

Q

How do I automate code review with an agent?

Run a headless agent with a prompt file containing the diff or PR context, --plan-only for a findings-only pass, and a required review artifact on stdout. Keep it read-only: review agents should propose, not push.

Q

What about long-running builds and tests?

The monitor tool watches up to eight processes and reports completion; --wait-monitors keeps the agent process alive until monitored jobs finish. Command output is capped at 30% of the context window with an explicit re-fetch marker, so a verbose log cannot poison the run.

Q

Is this a replacement for CI?

No, it is a step inside CI. Use your existing pipeline for build, test and environment control; add a coding agent for bounded jobs (fix a failing lint, apply a codemod, triage an issue) with its own sandbox profile and exit-code gate.

Q

What has actually been automated end to end?

Measured: three SaaS applications delivered by the swarm production E2E (2026-09-13) with worktree isolation, 35/43/38 integration tests green and zero conflicts. Also measured: 24/24 verified pairs on headless real-world fix tasks, with fix-task uncached tokens down 11–21%. Beyond that we do not claim unattended production autonomy.

Run Agents That Fit On Your Laptop.

25.6 MB median RSS. 25 agents ran in parallel on a Core 2 Duo with 4 GB RAM. Hundreds on your machine. Zero cloud required on the Ollama path.

Requires Rust/cargo to build from source. Linux and macOS today, Windows not yet supported. Pre-1.0, public beta. Pricing TBD.