Evidence dashboard

Trust claims should be inspectable.

Every published number below traces to a checksummed artifact bundle. Missing evidence stays “unmeasured”—never zero, never implied.

Valid Recall@1+33.3 pp
4/6 → 6/6
Forbidden exposure−33.3 pp
2/6 → 0/6
Mean overhead+0.1482
milliseconds
Ranking stability30 / 30
queries identical across repeats
Paired result

Quality improved without replacing retrieval.

mem0 2.0.12 · infer=False · top-1 · candidate pool 4 · 12 documents · five clean-state provider processes.

Valid Recall@1

Percentage of six cases returning the expected valid evidence

mem0 only
66.7%
mem0 + Vault
100%

Forbidden-exposure case rate

Percentage of six cases exposing invalid evidence

mem0 only
33.3%
mem0 + Vault
0%

Source: five digest-bound repeats in publication bundle vaultgovbench-retrieval-v0.1/89b9156. Repeats test deterministic reproducibility over one fixture; they are not independent statistical samples.

Case-level evidence

Six failure modes retrieval alone cannot govern.

The guard evaluates canonical memory state after provider ranking and before evidence reaches the agent.

Governance case
mem0 only
mem0 + Vault
Stale temporal fact
Exposed
Blocked / valid selected
Private access boundary
Valid
Valid
Deleted ghost memory
Valid
Valid
TTL expiry
Valid
Valid
Superseded revision
Exposed
Blocked / valid selected
Privacy block
Valid
Valid

Case display reflects the published aggregate 4/6 baseline and 6/6 governed result. Raw per-query records remain the authoritative evidence.

Cost of protection

Governance overhead is measured, not hidden.

Latency is reported only as a within-track paired delta. Cross-provider raw speed is not ranked.

Mean query latency747.7018
ms baseline
Governed mean747.8500
ms with Vault
Paired delta+0.1482
ms mean overhead
P95 paired delta mean: +0.0592 ms. This is the mean of five per-repeat P95 values, not a pooled P95. Provider/API cost was unavailable and is not reported as zero.
Coverage

Published, diagnostic, and unmeasured stay visibly separate.

TrackStatusWhat the evidence supportsArtifact
mem0 + VaultPublishedGovernance lift and paired overhead under the frozen six-case fixtureSummary JSON
Vault standalonePublishedBuilt-in keyword track reproducibility, zero forbidden exposureSummary JSON
AgentMemory + VaultDiagnosticLocal directional result only; clean blinded publication gate not yet metDiagnostic notes
Letta / MemGPT + VaultUnmeasuredNo quality claimN/A
Claim boundary

This proves a governance contract—not universal memory superiority.

The fixture is public, synthetic, retrieval-only, top-1, and deliberately contains invalid-memory cases. It does not measure end-to-end answer quality and is not an official LoCoMo, LongMemEval, or vendor leaderboard score.