Use Case · Automated Review

A Reviewer With No Write Access.

The safe version of automated code review is a read-only agent that produces evidence-backed findings, and a human who still owns the merge.

Direct answer

Automated code review with a coding agent works best as a read-only pass: the agent explores the diff and its dependencies, checks project invariants and completeness, and emits findings with file-and-line evidence, without the ability to write or execute. In Anvaya, that is one headless invocation: anv --no-tui --plan-only with a prompt file containing the diff, sandbox policy limiting reads to the repository, and findings on stdout for the pipeline to post. The checks agents genuinely excel at are mechanical: missing tests, unupdated callers, retired patterns being reintroduced, blast radius beyond the diff. Intent, taste and outside-the-repo context stay human. Gate merges on objective verifiers; treat agent findings as claims to evaluate, not verdicts to obey.

What Agents Are Good At

Checks Humans Skip At 6pm.

01Mechanical completenessAgents are excellent at the boring checks humans skip at 6pm: does every new branch have a test, does every changed function have a caller update, are there leftover debug statements, does the migration have a rollback path.
02Stale-pattern detectionA memory-aware review knows which patterns the project retired. 'This re-introduces the retry shape we removed in March' is a finding no linter will produce.
03Invariant checkingDecisions recorded in the graph (auth boundaries, naming contracts, module dependencies) can be checked against the diff. The agent reviews against the project's own history, not generic style rules.
04Blast-radius summariesStructural retrieval answers who calls what; a review can summarize what the change touches beyond the diff before a human spends attention on it.

What Stays Human

Four Things Review Cannot Automate.

01Intent is not in the diffAn agent can know a change is unusual; it cannot know the product decision behind it. Review comments about intent need a human who was in the room.
02Taste and architectureWhether an abstraction is worth it, whether a module boundary is right, models produce plausible opinions, which is worse than silence when the call is genuinely judgment.
03Outside-the-repo contextCustomer commitments, incident history, compliance constraints. If it is not in the code, decisions or memory, the review cannot weigh it.
04Confident false positivesAn agent will occasionally assert a bug that is not one. Findings must be cheap to dismiss, which is why they should be claims with file:line evidence, not verdicts.

The Pattern

Five Steps, Read-Only.

011 · Freeze the inputGenerate the diff against a known base and write it to a file. Reproducible input beats 'review the repo,' which invites exploration and hallucinated scope.
022 · Run boundedHeadless, plan mode, a token budget and a round cap. The review is a task with a contract, not an open-ended audit.
033 · Require evidenceEvery finding cites file and line, plus the invariant or test it relates to. Findings without evidence are rejected by the prompt contract.
044 · Post, don't blockFindings go to the PR as comments; the pipeline gate stays on tests, lint and types. Humans triage the rest.
055 · Record the missesWhen a reviewer misses a real bug, the bug thread goes into memory so the next review starts from it. Review quality compounds the same way agent knowledge does.

Questions

Asked About Automated Review.

Q

How do I run an agent as a reviewer?

Keep it read-only. Run a headless session with plan mode, give it the diff or PR context as a prompt file, and require findings on stdout as a machine-readable list. Read-only mode blocks write and exec tool calls at dispatch, so review cannot mutate the tree even if the prompt drifts.

Q

What should the findings contain?

A claim, a location, and the evidence that supports it: file and line, which test or invariant is affected, and how confident the finding is. A review that says 'this looks risky' is noise; one that says 'this call path bypasses the auth check added in decision X' is usable.

Q

Should agents block merges?

Only objective checks should block, failing tests, lint, type errors. Agent review findings should inform a human decision, because the cost of a false positive is a wasted review cycle and the cost of a false negative is the bug you already have. Gate the verifier, not the opinion.

Q

Does review need the memory layer?

It makes review meaningfully better: past bug threads, retired patterns and recorded decisions turn generic comments into project-specific ones. Even without memory, the mechanical-completeness checks alone usually pay for the pass.

Run Agents That Fit On Your Laptop.

25.6 MB median RSS. 25 agents ran in parallel on a Core 2 Duo with 4 GB RAM. Hundreds on your machine. Zero cloud required on the Ollama path.

Requires Rust/cargo to build from source. Linux and macOS today, Windows not yet supported. Pre-1.0, public beta. Pricing TBD.