Machine Testimony Working Paper Series  ·  No. 2  ·  September 2026
doi 10.5281/zenodo.22738449  ·  this version 10.5281/zenodo.22738450

One of Eight

Whether deployed agent systems can record who approved an action, and whether any law requires them to

Troy Clifford1

Machine Testimony


Abstract

Several jurisdictions now give a person the right to human review of an automated decision. This paper asks a narrower question that the deployment of those rights depends on: whether the software that takes the decisions can produce a record showing which human reviewed one.

Ten widely deployed agent and agent-memory systems were read at pinned commits against twenty record-keeping requirements. Eight of the ten take or gate consequential actions. Of those eight, one can identify the person who approved an action. Six cannot: the structure carrying what a human decided has no field for which human decided it. One could not be determined from outside the vendor. The single system that can is the author's own reference implementation, which is disclosed here rather than left to be found, and which carries no evidential weight.

Twenty legal and quasi-legal instruments were then read against the same question. Twelve ask a record to show who intervened in a decision and four of those are law. The two instruments that ask for a record of the review itself, rather than for a right to one, are a private certification scheme and a regulator's published expectation, and neither can require it by law.

The paper reports the method, the per-system verdicts with the commit each was read at, four corrections made by readers after publication, and the limits of what a static reading of source code can establish. Every verdict cites a file and a line so that a wrong one is cheap to demonstrate.

Keywords: AI governance, human oversight, audit trails, agent frameworks, automated decision-making, conformance measurement.


1. Introduction

A right to human review of an automated decision is now law in at least four jurisdictions. Article 22(3) of the General Data Protection Regulation has been applicable since 25 May 2018. Quebec's section 12.1 has been in force since 22 September 2023, Articles 22A to 22D of the United Kingdom GDPR since 5 February 2026, and Article 34(1)(4) of South Korea's Framework Act since 22 January 2026. Colorado's Senate Bill 26-189 applies from 1 January 2027 and the European Union's Artificial Intelligence Act applies to high-risk systems from 2 December 2027.

Each of these creates a duty about a review. None of the four that are currently law specifies what a record of that review must contain. The question this paper asks is what happens when somebody tries to satisfy one of them using software they have already deployed: whether the record that software produces can say which person reviewed a particular decision.

This is a narrower question than whether a system is well governed, and deliberately so. It is answerable from published source code, which means it can be measured rather than asserted, and a wrong answer can be corrected by anybody who reads the same code.

2. Method

2.1 The rubric

Twenty requirements were written before any system was read, grouped into four cumulative conformance levels. They ask what a record contains, not what a system promises: whether entries carry write times, whether refusals are recorded as faithfully as permissions, whether an approval identifies a person, whether the approver's identity comes from an authentication layer rather than from something the proposing model can write, and whether a past state of the record can be verified by a party other than its author.

The requirement this paper turns on is R3.5, stated in the rubric as: “Does an approval identify a person or a named role holder?” A verdict of present requires that the approver be a person or a role rather than a model or a process. The partial case, also written in advance, is that an approval step exists but the approver may be an automated principal.

2.2 Scope declarations

Not every requirement applies to every system. A system that stores material but takes no actions cannot be failed for lacking an approval path. Each subject therefore declares a scope, and requirements are applied only within it. Eight of the ten subjects declare that they act. The results below concern those eight.

This distinction was not in the first version of the rubric. It was added after a reader reported that the validator was refusing a conformance level to a system that had earned it: a record-only system with a genuine hash chain could never reach the third level, and because the levels are cumulative its integrity at the fourth level stayed invisible however good it was. The scope declaration exists because of that report.

2.3 Pinned reading

Each system was read at a specific commit, recorded with the verdict. A verdict is a claim about one state of one repository on one date and nothing more. Every verdict cites a file and a line at that commit, so a reader who disagrees can open the same line rather than argue about the conclusion.

3. Subjects

Ten systems were selected for deployment breadth across agent frameworks and agent-memory systems: AutoGen, CrewAI, Graphiti, Haystack, LangGraph, Letta Code, mem0, OMEM, the OpenAI Agents SDK and Pydantic AI. Eight declare that they act. Graphiti and mem0 declare storage and derivation without an action path, and are outside the results below.

One of the ten, OMEM, is maintained by the author of this paper. That is stated here, in Section 4 where the result appears, and on every published page carrying the number, for the reason given in Section 5.

4. Results

Of the eight systems that take or gate consequential actions, one records an approval that identifies a person. Six do not. One could not be determined from outside the vendor.

Table 1. R3.5, whether an approval identifies a person, across the eight acting subjects. R3.6 is whether that identity comes from an authentication layer; R3.7 is whether the agent is prevented from approving its own action.
SystemR3.5R3.6R3.7CommitRead
AutoGenabsentabsentabsent027ecf0a379b2026-09-04
CrewAIabsentabsentabsent92eb5f91830c2026-09-04
Haystackabsentabsentabsent82da3adc2fac2026-09-06
LangGraphabsentabsentabsent81bf17b231232026-09-04
Letta Codeundeterminedundeterminedundeterminede0a0e1e622782026-09-04
OMEMpresentpresentpresent9a77c20665532026-09-04
OpenAI Agents SDKabsentabsentabsent89c02c828ee82026-09-04
Pydantic AIabsentabsentabsentc0e4d824eaa02026-09-06

R3.6 asks whether the approver's identity comes from the authentication layer rather than from something the proposing model can write. R3.7 asks whether the acting agent is prevented from approving its own action. The three move together in every subject, which is the expected shape: a system with no field for an approver has nowhere to record where that identity came from, and no principal to compare against the proposer.

The failure is structural rather than incidental. In the six systems marked absent, the object that carries a human's decision has no member for the human. An integration that wishes to record one must place it somewhere the framework does not define, which means a later reader must know where to look and must trust that it was not written by the model.

4.1 What a right to review meets in practice

Twenty legal and quasi-legal instruments were read against the question of whether the text asks a record to show who intervened in a decision. Twelve of the twenty do. Four of those twelve are law, as at 13 September 2026: Article 22(3) of the GDPR, Quebec's section 12.1, Articles 22A to 22D of the UK GDPR, and Article 34(1)(4) of South Korea's Framework Act.

Two further instruments are in effect and are not law. One is AIUC-1, a private certification scheme whose mandatory control E015.2 requires structured logs capturing approver identity, timestamp and decision outcome, with a separate control E015.4 requiring those logs to be tamper-evident. The other is the Information Commissioner's guidance on automated decision-making, which states that a controller should keep a record of how a human reviewed a decision. These are the only two instruments in the set that ask for a record of the review itself rather than for a right to one, and neither can require it by law.

5. Limitations

  1. The only passing subject is the author's own. OMEM is maintained by the author of this paper. A measurement whose single positive result belongs to the party conducting it is worth what its disclosure is worth, and the appropriate reading of the headline number is one of eight, and that one is mine. It is reported as a row rather than excluded, because removing it would hide the fact that the property is achievable.
  2. A static reading establishes what a structure can carry, not what a deployment does. A framework with no approver field cannot record one; a framework with a field may still be deployed in a way that never populates it. The absent verdicts are therefore stronger than the present one.
  3. Each verdict is against one commit. A system that added the field after the date recorded beside it is not described by this paper. The correction mechanism is a pull request against the subject file rather than a dispute about the conclusion.
  4. One subject is undetermined, and that is a finding rather than a gap. Letta Code could not be determined from outside the vendor. It is reported as undetermined rather than absent, because reporting an unknown as a failure would inflate the headline number in the direction that favours the author's argument.
  5. Ten systems are not a census of the field. They are ten widely deployed systems, selected before reading and named in full. The healthcare, insurance and financial platforms where most consequential automated decisions are actually taken are mostly not readable from outside at all, and this paper says nothing about them.
  6. This paper does not establish that a record naming an approver would make a review meaningful. Whether a reviewer had authority, training, or the independence to disagree are facts about a person and an organisation. No record format produces them and none should claim to.

6. Corrections

Four substantive corrections were made to this work by readers after publication, each of which changed a published claim. They are listed because a measurement programme that reports only its findings and not its errors is asking for a trust it has not earned.

  1. A reader reported that the reference validator refused the third conformance level to any record containing no decisions, a requirement that appeared nowhere in the specification text. The scope declaration described in Section 2.2 exists because of that report.
  2. A reader pointed out that an expired evidence deadline and a failed action are different facts and that the format could express neither, because the member recording execution was a required boolean. An optional outcome member was added.
  3. A reader observed that every timestamp in a record is written by the party whose conduct is in question, and that the section listing what a reader cannot settle from a record did not say so. The specification now names it, and where a record carries an RFC 3161 token the validator checks that no entry claims a write time later than the authority's. Back-dating remains open and the specification says so.
  4. A reader independently reproduced the conformance corpus at a pinned commit, matching the published totals, then contested one case and was correct. The same reader subsequently corrected this programme's reading of Article 12(3)(d) of the EU AI Act, and then verified the fix and found a further error in it.

7. Reproduction

The subject files, the rubric, and the per-verdict evidence are public. The census manifest records a digest over the scored material, and the check recomputes it.

git clone https://github.com/troybrandonc-bit/machine-testimony
python3 census/manifest.py --check

The headline number and its method are published at machinetestimony.org/named-approver/, the full register at /register/, and the dated edition of the census at /census/2026-09/.

Notes

  1. Machine Testimony is a research programme, not a registered institute or a certification body. It publishes readings of public texts and measurements of public software. It does not certify, audit, or offer an opinion about any deployment. Correspondence: troy@machinetestimony.com.
  2. The author maintains OMEM, one of the ten subjects, and a record format and Internet-Draft in the same area. Both interests are disclosed here, in Section 4, in Section 5, and on every published page that carries the headline number.