Fix the question. Blind the provider. Publish the limits.
The harness is designed to prevent the website from becoming stronger than the evidence.
One controlled question per track.
This publication asks whether a governance layer removes invalid exposure without blocking valid evidence.
6 cases
Public synthetic governance fixture.
12 documents
Valid and deliberately invalid memory states.
Top-1 / K=4
Provider returns four candidates; scorer evaluates top-1.
5 repeats
Sequential, distinct clean-state processes.
Provider retrieval stays provider retrieval.
Blinded provider input
The provider receives redacted documents and queries, not policy labels or the full evaluation fixture. Gold governance labels stay in the policy replay and scorer processes.
Paired comparison
The same ranked candidates feed baseline and governed arms. Vault's guard filters against canonical state; it does not silently swap in a different provider.
Clean-state identity
Each repeat uses a distinct provider store or namespace, verifies index completion, and binds execution evidence to raw runs.
Honest missingness
Unavailable provider cost and unmeasured integrations remain unavailable. Missing data is never rendered as zero.
Fail closed before a number becomes a headline.
| Gate | Purpose | Published bundle |
|---|---|---|
| Checksums and unique artifacts | Detect mutation or accidental reuse | Passed |
| Clean source chain | Bind results to committed benchmark code | Passed |
| Blinded provider input | Keep evaluation labels away from retrieval | Passed |
| Provider integrity and preflight | Verify native retrieval actually ran | Passed |
| Five distinct clean-state repeats | Check deterministic reproducibility | Passed |
| Zero indexing failures | Avoid scoring partial corpora | Passed |
Check all 36 evidence files with one command.
python scripts/verify_publication_bundle.py \ benchmarks/results/vaultgovbench-retrieval-v0.1/89b9156
The verifier fails closed on changed or unlisted files, unsafe paths, invalid indexes, or non-publishable summaries. A pass verifies artifact integrity—not an independent provider rerun.
fixture sha256:
017c0629a561…c3b0b5f
provider input sha256:
ddd3a0799c46…41c6f4fc4
source revision:
89b9156f501b…eb159246b4b1db
Generate the website from the verified evidence.
python scripts/build_benchmark_site_data.py python tests/test_benchmark_site.py
What this evidence does not prove.
Not end-to-end QA
It measures retrieval governance, not final answer generation or agent task success.
Not a universal claim
One published external provider does not prove improvement for every memory engine.
Not production scale
Six synthetic cases and twelve documents do not measure large-corpus throughput or multi-agent contention.
From contract proof to category proof.
Clean AgentMemory and Letta publications; neutral LoCoMo/LongMemEval retrieval tracks; scale curves; concurrent multi-agent operations; failure injection; native extraction cost and quality; third-party reproduction.