Can an independent reviewer replay the result you are about to rely on?
Before you publish an evaluation, answer diligence, or submit assurance evidence, we rebuild and run one declared workflow from fresh repository states — then show exactly what the retained artifacts support and what remains unproved.
One repository. Two fresh replays. One claim-scoped disposition.
Best immediately before a benchmark release, technical diligence decision, procurement response, or regulated assurance gate.
A portable Replay Evidence Packet
Supported, unsupported, and prohibited claims.
Exact test IDs, exits, adverse categories, and output identity.
Artifacts, transcripts, inventory, manifest, and fail-closed verifier.
Isolation, locks, provenance, receipts, workspace protection, semantics.
EVIDENCE-SUPPORTED, EVIDENCE-NOT-SUFFICIENT, or REPLAY-BLOCKED.
Green diagnostics. Still not “verified.”
181/181 tests passed in fresh checkout A.
181/181 tests passed in fresh checkout B.
0 adverse categories. Reports were byte-identical: 437 bytes each.
Yet the terminal result remained NOT VERIFIED: qualifying kernel isolation, locked dependencies, production trust receipts, cache-free source, continuous no-write, and semantic completeness were absent or unproved.
Non-canonical engineering evidence only. No scientific, security, legal, or production certification claim.
Every disposition is limited to the reviewed claim, commit, command, manifest, and observed environment evidence.
Most teams can produce logs. Few can defend the exact claim.
Evidence replay is not observability trace playback. It rebuilds one declared workflow from fresh repository states, reconciles expected versus executed tests, compares canonical outputs, and retains a machine-verifiable chain of artifacts.
One repo · one declared workflow · two fresh replays within the preflight envelope · ≤250 tests · one packet · one findings call.
A negative disposition is valid completed delivery. Excludes remediation, penetration testing, scientific peer review, legal assurance, production certification, and unapproved cloud spend.
Repository. Baseline command. Test manifest. Exact intended claim.
Delivery target begins after access and a runnable baseline are confirmed.