01Objective verificationDoes the tool run something external (tests, a linter, a build) before calling a task done? A model self-report is not a completion signal.
02Confinement you can readIs there a sandbox policy for reads, writes, exec and network, enforced independently of approval? Does auto-approval widen it? (It should not.)
03An exit contractFor CI, headless mode needs pipeable I/O and distinct exit codes for success, error and approval-denied. Without them, a pipeline cannot tell failure from a missing gate.
04FootprintAn agent lives beside your editor for hours. Measure idle memory and startup from a release build, and print both kernel footprint and RSS, they disagree.
05Memory that is inspectableIf the tool claims to remember, ask where the data lives and in what format. Local files you can query and delete beat an opaque cloud profile.
06Provider freedomCan you run local models with no account, and bring your own keys for frontier ones? Model lock-in is workflow lock-in.
07Isolation for parallel workSubagents should get fresh context windows; fleet workers should get isolated workspaces. 'Parallel' on one working tree is a merge conflict with extra steps.
08ObservabilityCan you see per-agent state, token spend and process memory while it runs? Retrofitting observability after a runaway job is called incident response.