What OpenAI Agents SDK records, and what it does not

Read at89c02c8
Repositorygithub.com/openai/openai-agents-python
LicenceMIT
Assessed asstores, derives, acts
Reachesno level yet
Read byTroy Brandon Clifford
LevelMeetsWhat the level asks
TR-1 Recorded2/5The record exists and is append-only.
TR-2 Explained1/5Every belief resolves to its evidence, and disagreements survive.
TR-3 Gated2/7Actions carry a verdict, and approvals carry a name.
TR-4 Verifiable0/3The record can be shown not to have changed.

A count is requirements fully met out of those that apply. This system is assessed as stores, derives, acts, and requirements outside that are not counted against it.

This is the first system assessed here that gates actions, so it is the first with anything to say at TR-3, and the shape of the answer is consistent: the SDK knows a great deal about WHAT was proposed and almost nothing about WHO allowed it. Tools carry needs_approval set in developer code, a run keeps a canonical ledger of tool invocations recording whether each executed, and a rejection can carry a reason. But approve() takes no approver, and the only thing called an identity in the approval path identifies the tool call rather than a person. The other half of the picture is the session layer, where the default store is append-only but the optional compaction session clears the whole history and writes back a model-written summary in its place. Both observations describe an SDK doing its job: it is a framework for building agents, not a record of what they did, and it does not claim otherwise.

Every verdict, and what it rests on

Twenty requirements, each stated as a capability rather than a format, so a system that holds the information in its own shape counts as having it. An absent verdict cites where the assessor looked and did not find it, which is the difference between a measurement and an accusation.

TR-1 Recorded

The record exists and is append-only.

R1.1 · present. When the system stores a fact, does it durably record when that happened?

R1.2 · partial. When a stored fact changes, is the previous version still readable?

In the default session there is nothing to preserve, because nothing changes. Where a session does change, in compaction, the prior state is destroyed rather than superseded.

R1.3 · partial. Does the ordinary write path ever destroy what was previously recorded?

The default path appends and is clean. What makes this partial is that a supported and automatic path destroys: with the compaction session, reaching the item threshold clears the session and replaces it with a summary, so the conversation an auditor would want to read is gone by design rather than by accident.

R1.4 · present. Are entries distinguishable by kind, or is everything one undifferentiated blob of text?

R1.5 · absent. When data must be destroyed for a legal reason, is the destruction itself recorded?

  • source src/agents/memory/sqlite_session.py:435-446 clear_session
    deletes the session's messages and the session row itself, with ON DELETE CASCADE on the foreign key
  • searched grep -rniE 'tombstone|deleted_at|erasure|audit' src/agents/memory/ --include=*.py
    no match; nothing survives a cleared session to record that it existed

TR-2 Explained

Every belief resolves to its evidence, and disagreements survive.

R2.1 · partial. Can the source a stored fact came from be recovered from the store, by following a link rather than by guessing?

For verbatim items the question barely applies, which is a fair pass. For the one kind of item the SDK generates itself, the sources are neither cited nor retained.

R2.2 · present. Can a fact the system inferred be told apart from one it was told?

Model-written content in the history is labelled as model-written. That is the whole of what this requirement asks, and the SDK does it.

R2.3 · absent. When two stored facts about the same proposition disagree, do both survive?

This is a consequence of the SDK storing conversation rather than belief. Both statements survive as text, but nothing in the system can tell that they are about the same thing, so nothing is retained as a disagreement.

  • searched grep -rniE 'class Conflict|contradiction|disagree' src/agents/ --include=*.py
    the single match is an unrelated code comment at run_state.py:2514; there is no notion of a proposition, so there is nothing that could disagree
  • source src/agents/memory/sqlite_session.py:245-253
    the session is an ordered list of conversation items; two statements that contradict each other are simply two items

R2.4 · absent. Is the disagreement itself queryable, or must a reader diff rows to notice it?

  • searched grep -rniE 'class Conflict|contradiction|disagree' src/agents/ --include=*.py
    no conflict type, record or query surface exists anywhere in the package

R2.5 · partial. When a conflict is resolved, does the record say who resolved it and by what method?

Never resolving would be a pass. Compaction is the exception: it is the one place where an earlier statement can quietly stop being part of the record, decided by a model, with nothing written down about the decision.

TR-3 Gated

Actions carry a verdict, and approvals carry a name.

R3.1 · partial. Does a consequential action produce a durable entry whether or not it ran?

The information exists and is well modelled, which is most of the work. What it is not is durable: the ledger lives in a run object, and whether it survives the process is a decision the SDK leaves to whoever embeds it.

  • source src/agents/run_state.py:1356-1370 _serialize_tool_invocations
    a canonical per-run ledger records each call id with its type, approval scope, fingerprint, and whether it executed and completed. An attempted call that never ran is present in it
  • source src/agents/run_state.py:1330-1354 _serialize_approvals
    approvals and rejections serialise alongside, including rejection messages
  • searched grep -rniE 'tombstone|deleted_at|erasure|audit' src/agents/memory/ --include=*.py
    no persistence layer writes the ledger anywhere; RunState serialises to JSON on request and it is the application's job to keep it

R3.2 · present. Does an action's risk class come from somewhere the proposing model cannot write to?

This is a boolean gate rather than a graded risk class, but the substance the requirement asks for is met: whether an action needs approval is decided outside anything the model writes, and the model cannot clear its own gate.

  • source src/agents/tool.py:486-494
    needs_approval is declared on the tool, as a bool or a callable, and the docstring states the call must be approved via RunState.approve() or rejected via RunState.reject() before continuing
  • source src/agents/tool.py:707-743
    the function_tool decorator takes needs_approval and passes it into the constructed tool, so it is fixed in developer code at definition time

R3.3 · partial. Are refusals recorded as faithfully as permissions?

Refusals are modelled as faithfully as permissions, and the same caveat applies as at R3.1: this lives in the run state, not in a record that outlives it.

R3.4 · present. Does a refusal record why it was refused?

R3.5 · absent. Does an approval identify a person or a named role holder?

The SDK pauses the run and hands the decision to the embedding application. Who that application asked, and whether it asked anybody, is not part of what it records.

  • source src/agents/run_state.py:1285 approve
    the signature is approve(self, approval_item, always_approve=False). There is no parameter for who is approving
  • source src/agents/run_context.py:1043 approve_tool
    the underlying context call takes the same arguments and records approved/rejected as booleans or call-id lists, with no actor
  • searched grep -rniE 'approver|approved_by|principal' src/agents/run_state.py src/agents/run_context.py src/agents/tool.py
    the matches are approval_identity at run_state.py:1080 and :1190, which come from _tool_invocation.py:135 tool_invocation_identity and identify WHICH TOOL CALL an approval belongs to, by server label, tool name and request id. Nothing in the approval path identifies a person

R3.6 · absent. Does the approver's identity come from the authentication layer rather than from something the model can write?

  • searched grep -rniE 'approver|approved_by|principal' src/agents/run_state.py src/agents/run_context.py src/agents/tool.py
    no human identity is captured at any point, so there is no identity for an authentication layer to be the source of
  • source src/agents/run_state.py:1285-1298
    approve resolves the approval item and calls through to the context; nothing about the caller is examined or stored

R3.7 · absent. Is the acting agent prevented from approving its own action?

This is the requirement most likely to be assumed satisfied by anyone who sees a human-in-the-loop feature and stops reading. The pause is real; the attribution is not.

  • source src/agents/run_state.py:1285 approve
    always_approve=True marks a tool approved for the remainder of the run, so a program can clear the gate for every subsequent call in one statement
  • searched grep -rniE 'approver|approved_by|principal' src/agents/run_state.py src/agents/run_context.py src/agents/tool.py
    with no principals in the model, there is nothing to compare, so a proposer and an approver cannot be told apart even in principle

TR-4 Verifiable

The record can be shown not to have changed.

R4.1 · absent. Does the system publish a scheme under which the record's past state can be verified?

  • searched grep -rniE 'hash_chain|merkle|tamper|checksum|signature' src/agents/ --include=*.py
    matches are Python introspection (inspect.signature) and a ToolCallSignature tuple used to index tool calls; nothing relates to record integrity

R4.2 · absent. Can an independent party run that verification without the vendor's cooperation, and without the operator's?

  • searched grep -rniE 'hash_chain|merkle|tamper|checksum|signature' src/agents/ --include=*.py
    no verification surface exists, so there is nothing for an independent party to run

R4.3 · absent. Would alteration of a past entry be detectable after the fact?

  • source src/agents/memory/sqlite_session.py:245-253
    session rows carry an autoincrement id and a timestamp, and nothing binding one row to the next; an UPDATE against the table would leave no evidence
  • searched grep -rniE 'hash_chain|merkle|tamper|checksum' src/agents/ --include=*.py
    no match

How this was read

Read of src/agents/tool.py, src/agents/run_state.py, src/agents/_tool_invocation.py, src/agents/memory/sqlite_session.py and src/agents/memory/openai_responses_compaction_session.py at 89c02c8, cloned from the public repository. Line numbers are that commit's.

Found an error? Challenge a finding

Then it is wrong in the ordinary way readings are wrong, and every verdict cites a file and a line at a pinned commit precisely so that being wrong is cheap to demonstrate. The remedy is a pull request against the subject file, and it does not involve persuading anybody. Nobody applied for this and it is not a certification.

The standing register carries every system side by side, and the rubric is the twenty requirements in full, free to apply to anything, including to this assessment.