Engineering evidence / 29–30 September 2026
Withdrawal, recovery,
and the work between.
Recorded implementation result: withdrawing one accepted finding withheld its recorded direct and transitive dependents from a fresh reader’s approved context. Unrelated work remained available. Recovery required separate reassessments, while prior admissions stayed in history.
The four-stage result
The fixture contains four invented findings. B depends on A; C depends on B; U is unrelated. Each recorded read ran in a new process against the same persistent SQLite ledger.
| Fresh reader stage | A / Source | B / Recommendation | C / Brief | U / Unrelated |
|---|---|---|---|---|
| Initial handoff | Reusable | Reusable | Reusable | Reusable |
| Source withdrawn | Withdrawn | Blocked | Blocked | Reusable |
| Recommendation reassessed | Withdrawn | Reusable | Blocked | Reusable |
| Brief reassessed | Withdrawn | Reusable | Reusable | Reusable |
Restoring B required a scripted reviewer decision to retire its invalid dependency, relying on independent evidence already bound to B. C remained invalidated until its own reassessment. A stayed withdrawn.
What actually ran
- Unmodified Proofpress at commit ba9895f, using real kernel operations and an isolated SQLite ledger.
- A standalone harness with nine distinct subprocesses and 26 passing checks. This count includes aggregate checks for the upstream demo and selected test run.
- 23 selected upstream tests covering dependency lifecycle, context discovery, the existing conflict demo, and run tracking. This was not a full-suite run; the two counts must not be added together.
- Rejection checks for proposer self-approval and self-reassessment, restoring work with an invalid required dependency, re-admission as a shortcut, stale ledger heads, and conflicting request IDs. Each tested rejection left the ledger head unchanged.
- An identical withdrawal request replayed its prior receipt without adding a ledger event.
This demonstrates behavior already implemented in Richard Tang’s Proofpress code. The additional contribution is the standalone scenario, reproducible checks, and recorded evidence package.
What this does not establish
- All findings, evidence, and review decisions are synthetic. Reviewer identities were self-asserted; authenticated hosted authorization was not tested.
- Dependencies and contradictions are explicit inputs. The kernel does not establish semantic truth or evidence sufficiency.
- No live LLM ran. A fresh process demonstrates durable replay, not that an agent consults approved context or discards cached copies.
- This run does not establish customer savings, general performance improvement, novelty, or superiority over another product.
- AFF export/import, cross-system lifecycle propagation, and workspace-boundary infrastructure were not tested in this demonstration.
The next useful test
Choose one authorized research workflow with real reviewers and a successor agent. Agree on tasks and scoring before the run. Compare correctness, withdrawn-finding reuse, valid-finding retention, and total capture, review, and handoff effort against the team’s current process. Preserve negative and inconclusive results too.
The website replay is a display of these recorded states, not a live kernel execution. The runnable evidence bundle is available for evaluation discussions.
Discuss the evidence