Machine Testimony Working Paper Series · No. 2 · September 2026
doi 10.5281/zenodo.22738449 · this version 10.5281/zenodo.22738450
Whether deployed agent systems can record who approved an action, and whether any law requires them to
Machine Testimony
Several jurisdictions now give a person the right to human review of an automated decision. This paper asks a narrower question that the deployment of those rights depends on: whether the software that takes the decisions can produce a record showing which human reviewed one.
Ten widely deployed agent and agent-memory systems were read at pinned commits against twenty record-keeping requirements. Eight of the ten take or gate consequential actions. Of those eight, one can identify the person who approved an action. Six cannot: the structure carrying what a human decided has no field for which human decided it. One could not be determined from outside the vendor. The single system that can is the author's own reference implementation, which is disclosed here rather than left to be found, and which carries no evidential weight.
Twenty legal and quasi-legal instruments were then read against the same question. Twelve ask a record to show who intervened in a decision and four of those are law. The two instruments that ask for a record of the review itself, rather than for a right to one, are a private certification scheme and a regulator's published expectation, and neither can require it by law.
The paper reports the method, the per-system verdicts with the commit each was read at, four corrections made by readers after publication, and the limits of what a static reading of source code can establish. Every verdict cites a file and a line so that a wrong one is cheap to demonstrate.
Keywords: AI governance, human oversight, audit trails, agent frameworks, automated decision-making, conformance measurement.
A right to human review of an automated decision is now law in at least four jurisdictions. Article 22(3) of the General Data Protection Regulation has been applicable since 25 May 2018. Quebec's section 12.1 has been in force since 22 September 2023, Articles 22A to 22D of the United Kingdom GDPR since 5 February 2026, and Article 34(1)(4) of South Korea's Framework Act since 22 January 2026. Colorado's Senate Bill 26-189 applies from 1 January 2027 and the European Union's Artificial Intelligence Act applies to high-risk systems from 2 December 2027.
Each of these creates a duty about a review. None of the four that are currently law specifies what a record of that review must contain. The question this paper asks is what happens when somebody tries to satisfy one of them using software they have already deployed: whether the record that software produces can say which person reviewed a particular decision.
This is a narrower question than whether a system is well governed, and deliberately so. It is answerable from published source code, which means it can be measured rather than asserted, and a wrong answer can be corrected by anybody who reads the same code.
Twenty requirements were written before any system was read, grouped into four cumulative conformance levels. They ask what a record contains, not what a system promises: whether entries carry write times, whether refusals are recorded as faithfully as permissions, whether an approval identifies a person, whether the approver's identity comes from an authentication layer rather than from something the proposing model can write, and whether a past state of the record can be verified by a party other than its author.
The requirement this paper turns on is R3.5,
stated in the rubric as: “Does an approval identify a person or a
named role holder?” A verdict of present requires that the
approver be a person or a role rather than a model or a process. The
partial case, also written in advance, is that an approval step exists but
the approver may be an automated principal.
Not every requirement applies to every system. A system that stores material but takes no actions cannot be failed for lacking an approval path. Each subject therefore declares a scope, and requirements are applied only within it. Eight of the ten subjects declare that they act. The results below concern those eight.
This distinction was not in the first version of the rubric. It was added after a reader reported that the validator was refusing a conformance level to a system that had earned it: a record-only system with a genuine hash chain could never reach the third level, and because the levels are cumulative its integrity at the fourth level stayed invisible however good it was. The scope declaration exists because of that report.
Each system was read at a specific commit, recorded with the verdict. A verdict is a claim about one state of one repository on one date and nothing more. Every verdict cites a file and a line at that commit, so a reader who disagrees can open the same line rather than argue about the conclusion.
Ten systems were selected for deployment breadth across agent frameworks and agent-memory systems: AutoGen, CrewAI, Graphiti, Haystack, LangGraph, Letta Code, mem0, OMEM, the OpenAI Agents SDK and Pydantic AI. Eight declare that they act. Graphiti and mem0 declare storage and derivation without an action path, and are outside the results below.
One of the ten, OMEM, is maintained by the author of this paper. That is stated here, in Section 4 where the result appears, and on every published page carrying the number, for the reason given in Section 5.
Of the eight systems that take or gate consequential actions, one records an approval that identifies a person. Six do not. One could not be determined from outside the vendor.
| System | R3.5 | R3.6 | R3.7 | Commit | Read |
|---|---|---|---|---|---|
| AutoGen | absent | absent | absent | 027ecf0a379b | 2026-09-04 |
| CrewAI | absent | absent | absent | 92eb5f91830c | 2026-09-04 |
| Haystack | absent | absent | absent | 82da3adc2fac | 2026-09-06 |
| LangGraph | absent | absent | absent | 81bf17b23123 | 2026-09-04 |
| Letta Code | undetermined | undetermined | undetermined | e0a0e1e62278 | 2026-09-04 |
| OMEM | present | present | present | 9a77c2066553 | 2026-09-04 |
| OpenAI Agents SDK | absent | absent | absent | 89c02c828ee8 | 2026-09-04 |
| Pydantic AI | absent | absent | absent | c0e4d824eaa0 | 2026-09-06 |
R3.6 asks whether the approver's identity comes from the authentication layer rather than from something the proposing model can write. R3.7 asks whether the acting agent is prevented from approving its own action. The three move together in every subject, which is the expected shape: a system with no field for an approver has nowhere to record where that identity came from, and no principal to compare against the proposer.
The failure is structural rather than incidental. In the six systems marked absent, the object that carries a human's decision has no member for the human. An integration that wishes to record one must place it somewhere the framework does not define, which means a later reader must know where to look and must trust that it was not written by the model.
Twenty legal and quasi-legal instruments were read against the question of whether the text asks a record to show who intervened in a decision. Twelve of the twenty do. Four of those twelve are law, as at 13 September 2026: Article 22(3) of the GDPR, Quebec's section 12.1, Articles 22A to 22D of the UK GDPR, and Article 34(1)(4) of South Korea's Framework Act.
Two further instruments are in effect and are not law. One is AIUC-1, a private certification scheme whose mandatory control E015.2 requires structured logs capturing approver identity, timestamp and decision outcome, with a separate control E015.4 requiring those logs to be tamper-evident. The other is the Information Commissioner's guidance on automated decision-making, which states that a controller should keep a record of how a human reviewed a decision. These are the only two instruments in the set that ask for a record of the review itself rather than for a right to one, and neither can require it by law.
Four substantive corrections were made to this work by readers after publication, each of which changed a published claim. They are listed because a measurement programme that reports only its findings and not its errors is asking for a trust it has not earned.
The subject files, the rubric, and the per-verdict evidence are public. The census manifest records a digest over the scored material, and the check recomputes it.
git clone https://github.com/troybrandonc-bit/machine-testimony
python3 census/manifest.py --check
The headline number and its method are published at machinetestimony.org/named-approver/, the full register at /register/, and the dated edition of the census at /census/2026-09/.