Trust claims should be inspectable.
Every published number below traces to a checksummed artifact bundle. Missing evidence stays “unmeasured”—never zero, never implied.
Quality improved without replacing retrieval.
mem0 2.0.12 · infer=False · top-1 · candidate pool 4 · 12 documents · five clean-state provider processes.
Valid Recall@1
Percentage of six cases returning the expected valid evidence
Forbidden-exposure case rate
Percentage of six cases exposing invalid evidence
Source: five digest-bound repeats in publication bundle vaultgovbench-retrieval-v0.1/89b9156. Repeats test deterministic reproducibility over one fixture; they are not independent statistical samples.
Six failure modes retrieval alone cannot govern.
The guard evaluates canonical memory state after provider ranking and before evidence reaches the agent.
Case display reflects the published aggregate 4/6 baseline and 6/6 governed result. Raw per-query records remain the authoritative evidence.
Governance overhead is measured, not hidden.
Latency is reported only as a within-track paired delta. Cross-provider raw speed is not ranked.
Published, diagnostic, and unmeasured stay visibly separate.
| Track | Status | What the evidence supports | Artifact |
|---|---|---|---|
| mem0 + Vault | Published | Governance lift and paired overhead under the frozen six-case fixture | Summary JSON |
| Vault standalone | Published | Built-in keyword track reproducibility, zero forbidden exposure | Summary JSON |
| AgentMemory + Vault | Diagnostic | Local directional result only; clean blinded publication gate not yet met | Diagnostic notes |
| Letta / MemGPT + Vault | Unmeasured | No quality claim | N/A |
This proves a governance contract—not universal memory superiority.
The fixture is public, synthetic, retrieval-only, top-1, and deliberately contains invalid-memory cases. It does not measure end-to-end answer quality and is not an official LoCoMo, LongMemEval, or vendor leaderboard score.