Drop a file of the records a system already writes, in whatever shape it writes them, and this page reports which of the four questions those rows can answer. It runs in the page. Nothing is uploaded and nothing is stored, so a client’s log does not leave the machine it is read on, which is the only footing on which anybody can run this over somebody else’s data.
It is not the assessment. It is the thing worth doing before one, in the minute before a call rather than the day after it. The rubric is what you escalate to when a row here turns out to matter.
· · or drop one onto the box
A gap here is not a verdict on the system. Most agent software was not built to keep who-approved apart from that-something-ran, and until 2026 nobody was asking it to. Free adapters exist today for LangGraph, CrewAI, the OpenAI Agents SDK, Pydantic AI and AutoGen. Each names the approval point the framework already has and takes the approver’s identity from an authenticated session rather than inventing one.
instrument quickstart --framework <name> names the right adapter for a deployer’s own framework, prints the install line and the wiring snippet verbatim from that adapter’s own README, and chains straight into a check on the record it just produced. The adapter is MIT, standalone, and depends on nothing sold here; the library still only measures, and the emitter, which stays free and separate, is what writes.
Once a system emits a Testimony Record, Check reports the conformance level it reaches and every check behind that level, in the same browser, uploading nothing.
It reports the shape of the rows it was given and concludes nothing about the system that wrote them. A system may record an approver somewhere those rows have never been, and this cannot see that and does not claim to. That limit prints with every finding rather than being left for a reader to infer, because a finding that overstates its own scope is worth nothing to the person who has to stand behind it.
What it does establish is narrower and is usually the question anyway: whether the record somebody is handed after an incident answers these, or whether the answer has to come from somebody’s memory.
The same finding comes out of the command line, and the page prints the command that reproduces it. That matters more than the convenience: a finding a client can reproduce, and a supplier can dispute, is worth more than one that arrived from a website.
python3 spec/testimony_convert.py their-logs.jsonl --report
Both are run over the same rows by the test suite and the build fails if they disagree, because two instruments that disagree are worse than one.