Which systems record what

8 systems, read against 20 requirements drawn from the four conformance levels. Every verdict cites a file and a line at a pinned commit, and an absent verdict cites where the assessor looked and did not find it, which is the difference between a measurement and an accusation.

This is not a certification and nobody applied for it. It is a reading of published source, and it is wrong in the ordinary way that readings are wrong. The remedy is below and it does not involve persuading anybody.

SystemTR-1TR-2TR-3TR-4ReachesRead byLast read
AutoGen1/52/41/70/3none yetMichael Brandon Clifford2026-09-04
CrewAI1/50/51/70/3none yetMichael Brandon Clifford2026-09-04
Graphiti4/53/5n/a0/3none yetMichael Brandon Clifford2026-09-04
LangGraph4/50/42/70/3none yetMichael Brandon Clifford2026-09-04
Letta Code3/50/43/71/3none yetMichael Brandon Clifford2026-09-04
mem03/50/5n/a0/3none yetMichael Brandon Clifford2026-09-04
OMEM the author's own5/55/57/73/3TR-4Michael Brandon Clifford2026-09-05
OpenAI Agents SDK2/51/52/70/3none yetMichael Brandon Clifford2026-09-04

A count is met out of applicable. A system that does not act is not marked down for having no gate, and the requirements that do not apply to it are not counted against it.

The question this started from

Does a consequential action produce a durable entry whether or not it ran?

SystemR3.1Where it was read
AutoGenpartpython/packages/autogen-agentchat/src/autogen_agentchat/agents/_code_executor_agent.py:69-73 ApprovalRequest
CrewAIpartlib/crewai/src/crewai/agents/crew_agent_executor.py:984-989
Graphitin/a
LangGraphyeslibs/langgraph/langgraph/types.py:851 interrupt
Letta Codeyessrc/agent/check-approval.ts:26-34
mem0n/a
OMEMyesserver/tests_testimony_export.py
OpenAI Agents SDKpartsrc/agents/run_state.py:1356-1370 _serialize_tool_invocations

If a verdict here is wrong

It is a file. Open a pull request against the subject file naming the requirement and where to look, or write to troy@machinetestimony.com with the same. A correction that lands changes the file, this page, and the date on the row. There is no fee, no membership, and no requirement to use any software of mine.

Arguing with me is not the remedy and does not work. Pointing at code is, and has: this register carries corrections that came from being told I had read something wrong.

Who read these

All of them, Michael Brandon Clifford, which is the weakest thing about this register. A reading nobody has repeated is one person's reading, however carefully it cites its sources.

The instrument is not reserved. The rubric is CC BY 4.0 and the tooling is MIT, commercial use included and expected: if you audit AI systems, or advise on Article 12 or Article 14, or have to answer a procurement question about what a supplier's agent records, you can run these questions yourself and bill for it without asking anybody. How to do that, including how to put the result here with your own name on the row, or keep it and publish it yourself.

If your system is not here

Being absent is not a judgement. It means nobody has done the reading yet. The rubric and the harness are in the repository, so you can run the questions against your own system before anybody else does, and the answer will be the same one I would get.

The conflict of interest

One row is the author's own implementation, of the format the questions derive from. It scores well here the way a dictionary's author spells well, and it carries no evidential weight. It is included so the questions are applied to the system that produced them before they are applied to anybody else's.

That has not been costless. Five findings so far are recorded against this assessment, three of them against its author, including one after publication: on 5 September 2026 the top conformance level was found not to be checking what it claimed, and the author's own passing row was the one affected. It is written up in full rather than quietly repaired.

What a verdict rests on

Some of these questions a reader can settle from a record alone: whether cited evidence exists, whether a refused action is also recorded as executed, whether a digest is the digest of what it covers. Others are attestations, and no reading of source can confirm them: that a risk class really came from a registry, that an approver's name really came from the session it names. The specification marks the difference and the validator reports it, and a conformance claim that does not distinguish them is weaker than it looks.