Shadow eval vs A/B testing?
Same spirit, zero user impact: historical queries replay against candidates offline. Promote on measured wins.
Glossary · Shadow Evaluation
Canary deploys for knowledge.
Direct answer
Shadow eval scores candidate knowledge against real historical queries in parallel with live serving — measured, never served — promoting only what beats baseline. UNKNOWN is never treated as zero; absence of evidence stays absence, not failure. It is the final gate between extraction ('this might help') and serving ('this earns tokens every turn'), and the mechanism that lets memory systems learn aggressively without ever regressing answers.
In Anvaya
Questions
Same spirit, zero user impact: historical queries replay against candidates offline. Promote on measured wins.
Anvaya's experience benchmark fixtures: fail→fix headless, denial handling, tool-only exploration, producer-off baselines — open trials Q3 2026.
Serving unproven memory taxes every turn's attention. Measure first; attention is the scarcest budget.
One binary. 11+9 Rust crates. 545 tests. Hand-written HNSW index. Three transport modes. Four providers, Ollama, Anthropic, OpenAI, Siemens. Zero API keys required to start. Mind remembers everything after the first session.