Assessment, 2026-09-04 · 0.22.0 (89c02c8)
| Read at | 89c02c8 |
|---|---|
| Repository | github.com/openai/openai-agents-python |
| Licence | MIT |
| Assessed as | stores, derives, acts |
| Reaches | no level yet |
| Read by | Troy Brandon Clifford |
| Level | Meets | What the level asks |
|---|---|---|
| TR-1 Recorded | 2/5 | The record exists and is append-only. |
| TR-2 Explained | 1/5 | Every belief resolves to its evidence, and disagreements survive. |
| TR-3 Gated | 2/7 | Actions carry a verdict, and approvals carry a name. |
| TR-4 Verifiable | 0/3 | The record can be shown not to have changed. |
A count is requirements fully met out of those that apply. This system is assessed as stores, derives, acts, and requirements outside that are not counted against it.
This is the first system assessed here that gates actions, so it is the first with anything to say at TR-3, and the shape of the answer is consistent: the SDK knows a great deal about WHAT was proposed and almost nothing about WHO allowed it. Tools carry needs_approval set in developer code, a run keeps a canonical ledger of tool invocations recording whether each executed, and a rejection can carry a reason. But approve() takes no approver, and the only thing called an identity in the approval path identifies the tool call rather than a person. The other half of the picture is the session layer, where the default store is append-only but the optional compaction session clears the whole history and writes back a model-written summary in its place. Both observations describe an SDK doing its job: it is a framework for building agents, not a record of what they did, and it does not claim otherwise.
Twenty requirements, each stated as a capability rather than a format, so a system that holds the information in its own shape counts as having it. An absent verdict cites where the assessor looked and did not find it, which is the difference between a measurement and an accusation.
The record exists and is append-only.
R1.1 · present. When the system stores a fact, does it durably record when that happened?
R1.2 · partial. When a stored fact changes, is the previous version still readable?
In the default session there is nothing to preserve, because nothing changes. Where a session does change, in compaction, the prior state is destroyed rather than superseded.
R1.3 · partial. Does the ordinary write path ever destroy what was previously recorded?
The default path appends and is clean. What makes this partial is that a supported and automatic path destroys: with the compaction session, reaching the item threshold clears the session and replaces it with a summary, so the conversation an auditor would want to read is gone by design rather than by accident.
R1.4 · present. Are entries distinguishable by kind, or is everything one undifferentiated blob of text?
R1.5 · absent. When data must be destroyed for a legal reason, is the destruction itself recorded?
Every belief resolves to its evidence, and disagreements survive.
R2.1 · partial. Can the source a stored fact came from be recovered from the store, by following a link rather than by guessing?
For verbatim items the question barely applies, which is a fair pass. For the one kind of item the SDK generates itself, the sources are neither cited nor retained.
R2.2 · present. Can a fact the system inferred be told apart from one it was told?
Model-written content in the history is labelled as model-written. That is the whole of what this requirement asks, and the SDK does it.
R2.3 · absent. When two stored facts about the same proposition disagree, do both survive?
This is a consequence of the SDK storing conversation rather than belief. Both statements survive as text, but nothing in the system can tell that they are about the same thing, so nothing is retained as a disagreement.
R2.4 · absent. Is the disagreement itself queryable, or must a reader diff rows to notice it?
R2.5 · partial. When a conflict is resolved, does the record say who resolved it and by what method?
Never resolving would be a pass. Compaction is the exception: it is the one place where an earlier statement can quietly stop being part of the record, decided by a model, with nothing written down about the decision.
Actions carry a verdict, and approvals carry a name.
R3.1 · partial. Does a consequential action produce a durable entry whether or not it ran?
The information exists and is well modelled, which is most of the work. What it is not is durable: the ledger lives in a run object, and whether it survives the process is a decision the SDK leaves to whoever embeds it.
R3.2 · present. Does an action's risk class come from somewhere the proposing model cannot write to?
This is a boolean gate rather than a graded risk class, but the substance the requirement asks for is met: whether an action needs approval is decided outside anything the model writes, and the model cannot clear its own gate.
R3.3 · partial. Are refusals recorded as faithfully as permissions?
Refusals are modelled as faithfully as permissions, and the same caveat applies as at R3.1: this lives in the run state, not in a record that outlives it.
R3.4 · present. Does a refusal record why it was refused?
R3.5 · absent. Does an approval identify a person or a named role holder?
The SDK pauses the run and hands the decision to the embedding application. Who that application asked, and whether it asked anybody, is not part of what it records.
R3.6 · absent. Does the approver's identity come from the authentication layer rather than from something the model can write?
R3.7 · absent. Is the acting agent prevented from approving its own action?
This is the requirement most likely to be assumed satisfied by anyone who sees a human-in-the-loop feature and stops reading. The pause is real; the attribution is not.
The record can be shown not to have changed.
R4.1 · absent. Does the system publish a scheme under which the record's past state can be verified?
R4.2 · absent. Can an independent party run that verification without the vendor's cooperation, and without the operator's?
R4.3 · absent. Would alteration of a past entry be detectable after the fact?
Read of src/agents/tool.py, src/agents/run_state.py, src/agents/_tool_invocation.py, src/agents/memory/sqlite_session.py and src/agents/memory/openai_responses_compaction_session.py at 89c02c8, cloned from the public repository. Line numbers are that commit's.
Then it is wrong in the ordinary way readings are wrong, and every verdict cites a file and a line at a pinned commit precisely so that being wrong is cheap to demonstrate. The remedy is a pull request against the subject file, and it does not involve persuading anybody. Nobody applied for this and it is not a certification.
The standing register carries every system side by side, and the rubric is the twenty requirements in full, free to apply to anything, including to this assessment.