Blog · 2026-09-17 · 7 min

Self-Healing CI With Coding Agents: A Practical Pattern.

Plan-gate-apply, objective exit codes, monitors for long jobs, and memory that stops nightly retries re-learning the same failure.

Direct answer

The useful version of 'self-healing CI' is not an agent with merge rights. It is a bounded loop: a read-only pass proposes, a human or policy gate decides, a confined run applies, and the test suite is the only thing allowed to call the work done.

01

The pattern

Step one is a read-only planning pass: headless run, no writes, no execution, plan out on stdout or a file. Step two is the gate, a human approval, an issue label, or a path-scoped policy that says this class of change is safe to attempt. Step three is the mutating run inside a sandbox profile that matches the pipeline's tolerance, with the repository's own tests as the verifier.

Exit codes carry the state: 0 success, 1 error, 2 approval denied. A pipeline can tell 'tried and failed' apart from 'blocked by policy,' which matters because those get different responses, one retries or escalates, the other usually means a human forgot a gate.

02

Long jobs and dirty output

CI jobs are long and loud. The monitor tool tracks up to eight concurrent processes and reports completion; a headless run can be told to wait for its monitors instead of exiting early. Command output is capped at a share of the context window with an explicit re-fetch marker, so one verbose test log cannot drown the run that produced it.

Sessions survive pipeline restarts through a crash-safe store with retention limits, and runs can be resumed or listed by code. Retries therefore resume with context instead of starting from a cold repository.

03

Where memory changes the economics

The expensive failure in CI is the repeat: the same flaky test, the same known trap, the same revert that happened last Tuesday. A memory-aware run can carry the project's bug threads and past outcomes into the job, so the agent starts from what the repository already learned. On measured headless fix tasks, the compression lever cut uncached input tokens 11–21% across 24/24 verified pairs, that is the quantity a memory or compression layer actually changes.

The discipline matters more than the mechanism: one job per session, a written spec in, a diff out, and tests as the merge gate. Agents that own the whole pipeline eventually own the outage.

04

What not to automate

Keep agents away from force-pushes, secret-adjacent work, unreviewed production paths, and any merge they can authorize themselves. Auto-approval belongs where revert is one command; it never widens the sandbox, so a pipeline can safely use it for narrow, revertible jobs and still be refused out-of-scope actions.

Questions

Asked About This Post.

Q

Is this a replacement for CI?

No, it is a step inside CI. Your pipeline keeps build, test and environment control; the agent gets bounded jobs with its own sandbox and exit-code gate.

Q

What if the agent's tests pass but the change is wrong?

That is a test-coverage problem, not an agent problem. Verification is only as good as the verifier; keep tasks small enough that a passing suite is real evidence.

Q

How do I start?

Nightly read-only triage first: failing tests and open issues, findings to a file, no mutations. Add the mutating step only after the read-only pass is boringly reliable.

Run Agents That Fit On Your Laptop.

25.6 MB median RSS. 25 agents ran in parallel on a Core 2 Duo with 4 GB RAM. Hundreds on your machine. Zero cloud required on the Ollama path.

Requires Rust/cargo to build from source. Linux and macOS today, Windows not yet supported. Pre-1.0, public beta. Pricing TBD.