Measured · with caveats
Four Baselines, Flaws Attached.
These are the readings that exist. Each one is scoped, dated and labeled with what it does not prove, because the flaw is part of the finding.
MeasurementResultn / scopeStatus
Synthetic ranking baselinenDCG@5 0.648–0.671 · MRR 0.834–0.875100 seeded nodes → 53 queries · fixed seed · queries derived from expected nodes, so the result is circular by constructionCircular by design
Observed ranking baselineInsufficient positives: 185 feedback rows, 2 sessions, 4 query hashes, 0 positive observationsLive store · 2026-08-31Inconclusive
Claim annotation A/BNull result: citation rate 0.0% in both arms · treatment added +32% injection-block tokensn=57 observations · 3 paired prompts · 2026-08-31Negative, published
Injection utility baselineMean 0.189 → 0.186 (≈81% of injected memory unused)918 → 999 samples · live store · 2026-08-31Open finding
Mechanism evidenceSix-stage retrieval, submodular packing under a 4,096-token budget, per-type decay (Decision half-life ≈347 d → Approach ≈23 d), git-anchored drift demotion, Beta-Bernoulli utility updatesShipped paths, default-on unless noted; the learning stack is default-offMechanism, not outcome