Assessment, 2026-09-06 · 2.40.0+ (c0e4d82)
| Read at | c0e4d82 |
|---|---|
| Repository | github.com/pydantic/pydantic-ai |
| Licence | MIT |
| Assessed as | stores, acts |
| Reaches | no level yet |
| Read by | Troy Brandon Clifford |
| Level | Meets | What the level asks |
|---|---|---|
| TR-1 Recorded | 3/5 | The record exists and is append-only. |
| TR-2 Explained | 0/4 | Every belief resolves to its evidence, and disagreements survive. |
| TR-3 Gated | 3/7 | Actions carry a verdict, and approvals carry a name. |
| TR-4 Verifiable | 0/3 | The record can be shown not to have changed. |
A count is requirements fully met out of those that apply. This system is assessed as stores, acts, and requirements outside that are not counted against it.
Pydantic AI is assessed as storing and acting but not deriving. It keeps a typed message history across turns and it gates tool calls that have effects, but the framework makes no claims about the world beyond what the model produced, so R2.2 does not apply and is not scored. The same reading was applied to LangGraph.
This subject scores better on TR-1, and on the recording half of TR-3, than anything else assessed so far. Entries are typed with a discriminator the parser depends on, every part carries its own timestamp, and how a tool call ended is a closed literal in which denied sits beside success, so a refusal is a record of the same standing as a permission rather than an exception or a log line. Compaction inserts a marker and retains what it summarises. In-place mutation of history is detected and warned about, which no other subject does.
It stops in the same place as the rest. An approval is typed as a mapping to bool or DeferredToolApprovalResult, so a bare True approves, and a boolean has nowhere to put a person. Everything in TR-3 that concerns who authorised an action is therefore absent, and TR-4 is absent as it is everywhere.
One thing here does not fit the pattern and is worth saying plainly. ToolApproved carries override_args, which means the design already understands that what a person approved and what the model proposed can differ, a distinction most systems with an approval gate have not drawn. The gap is not that this framework has thought about approval carelessly. It is that it has thought about the action carefully and about the approver not at all.
Twenty requirements, each stated as a capability rather than a format, so a system that holds the information in its own shape counts as having it. An absent verdict cites where the assessor looked and did not find it, which is the difference between a measurement and an accusation.
The record exists and is append-only.
R1.1 · present. When the system stores a fact, does it durably record when that happened?
The strongest R1.1 in the census so far. The time is a field on the entry itself, defaulted at construction, and it round-trips through the framework's own serialiser. Nothing here depends on an application remembering to stamp anything.
R1.2 · partial. When a stored fact changes, is the previous version still readable?
Two documented mechanisms that change history, behaving in opposite ways. Compaction is done the way this requirement asks: a typed marker, the prior parts retained, and a function for reading either view. History processors replace outright and the docs put keeping the original on the caller. Partial because which answer a deployment gets depends on which of the two it uses.
R1.3 · present. Does the ordinary write path ever destroy what was previously recorded?
The ordinary path only appends. Worth recording that the framework actively detects in-place mutation and warns, which no other subject here does, though the stated motive is serialisation caching rather than integrity: the warning exists because instrumentation serialises each message once and a field mutated in place would not be picked up. Detection as a side effect of a performance decision is still detection, but it is not a guarantee, and nothing makes the objects immutable.
R1.4 · present. Are entries distinguishable by kind, or is everything one undifferentiated blob of text?
An instruction, a tool call and a tool result are separate types with a discriminator the parser depends on, which is what this requirement asks for. The outcome literal is a genuine refinement over the rest of the census: most systems record that a call returned, not how it ended.
R1.5 · absent. When data must be destroyed for a legal reason, is the destruction itself recorded?
The documented route to removing sensitive data leaves no mark that anything was removed. A reader of the resulting history cannot tell a conversation that was always short from one that was filtered. This is the ordinary shape of the finding across the census rather than anything unusual to this framework.
Every belief resolves to its evidence, and disagreements survive.
R2.1 · partial. Can the source a stored fact came from be recovered from the store, by following a link rather than by guessing?
Provenance is good exactly where the framework has structure to hang it on, and absent where it does not. A tool result names its call. A sentence the model wrote names nothing, so what a claim in the output rests on is recoverable only by reading the surrounding history and inferring.
R2.3 · partial. When two stored facts about the same proposition disagree, do both survive?
Both survive, but incidentally. The same verdict and the same reasoning as LangGraph: an append-only log preserves contradictions without knowing that is what it is doing, and a summarisation step configured by the deployment can end that.
R2.4 · absent. Is the disagreement itself queryable, or must a reader diff rows to notice it?
A reader has to notice a contradiction by reading the transcript. That is not a criticism of a message history, which is what this is; it is the fact the requirement is asking for.
R2.5 · absent. When a conflict is resolved, does the record say who resolved it and by what method?
Scored absent rather than present under the clause that credits a system which never resolves conflicts, because that clause is for a system that leaves a contradiction standing as a visible unresolved state. Here nothing is standing, because nothing was ever represented as a contradiction.
R2.2 is not assessed here, because this system is not in that business and marking it down for that would be dishonest.
Actions carry a verdict, and approvals carry a name.
R3.1 · present. Does a consequential action produce a durable entry whether or not it ran?
The clearest yes in the census on this requirement. Proposed, refused and executed calls are all entries, and which of the three happened is readable from a typed field rather than from whether a row is missing.
R3.2 · present. Does an action's risk class come from somewhere the proposing model cannot write to?
Present, with one honest qualification. The default and the primary mechanism put risk entirely outside model output. The optional predicate is handed tool_args, which the model wrote, so a deployment can choose to classify on model-produced values. The framework neither does that nor encourages it, and a requirement about where risk comes from should score what the framework establishes, not what a caller can opt into. Binary rather than graded, as everywhere else in this census.
R3.3 · present. Are refusals recorded as faithfully as permissions?
The best answer to this requirement in the census. Elsewhere a refusal is a state change, an exception or a line in application logging; here it is the same record type as a permission, distinguished by a typed field, and it goes back to the model as well as into the history.
R3.4 · partial. Does a refusal record why it was refused?
Ahead of most of the census, which records no reason at all, and short of the requirement, which asks for a machine-readable one. A deployment can tell that a call was refused by testing a field, and can tell why only by parsing English, which may be a default string nobody chose.
R3.5 · absent. Does an approval identify a person or a named role holder?
The same finding as almost every subject that has an approval gate, and the reason the adapter for this framework exists. ToolApproved is a richer object than most, carrying override_args so that what was approved can differ from what was proposed, which shows the design already takes the approval seriously as an object. It simply has no field for the person. The consequence is visible in approve_all: a deployment that approves a batch produces records identical to one where somebody read every call, and no reader of the history can tell the two apart.
R3.6 · absent. Does the approver's identity come from the authentication layer rather than from something the model can write?
Not reachable while R3.5 is absent: an identity that is never recorded cannot be sourced from anywhere. Recorded separately because the two failures come apart in systems that do capture an approver and then take the name from the request body.
R3.7 · absent. Is the acting agent prevented from approving its own action?
No separation is enforced or claimed. With no principals on either side there is nothing to compare, so an agent loop that supplies its own DeferredToolResults approves its own actions and the record cannot show that this is what happened.
The record can be shown not to have changed.
R4.1 · absent. Does the system publish a scheme under which the record's past state can be verified?
There is a documented way to persist a record and no documented way to verify one. That is the census-wide pattern rather than a fault peculiar to this framework: TR-4 is the level nothing assessed has reached.
R4.2 · absent. Can an independent party run that verification without the vendor's cooperation, and without the operator's?
Nothing here needs the vendor's cooperation, because there is no verification to withhold. Scored absent on the same basis as every other subject: absence of a scheme, not a restriction on access to one.
R4.3 · absent. Would alteration of a past entry be detectable after the fact?
The closest thing to alteration detection in the census, and still absent under the requirement. The warning fires inside a live run, is driven by a serialisation cache rather than by an integrity scheme, and says nothing about a record at rest, which is where the requirement is aimed. A stored history can be edited freely and nothing would show it.
Read of pydantic_ai_slim/pydantic_ai/messages.py, _tool_execution.py, _deferred.py, tools.py, toolsets/approval_required.py and docs/message-history.md at c0e4d824eaa0401d4481d401e5b3894ab32ab59d. The source archive for that commit was downloaded from the public repository rather than an installed release, so every line number refers to the pinned commit and every search below was run over all 305 modules of the package rather than over the handful of files that were read closely. That commit is main, dated 5 September 2026, one day after the v2.40.0 release, which is why the version is written as 2.40.0+ rather than as a release. Documentation was read only where it states a behaviour the source does not settle on its own, which here is what happens to history when a processor runs.
Then it is wrong in the ordinary way readings are wrong, and every verdict cites a file and a line at a pinned commit precisely so that being wrong is cheap to demonstrate. The remedy is a pull request against the subject file, and it does not involve persuading anybody. Nobody applied for this and it is not a certification.
The standing register carries every system side by side, and the rubric is the twenty requirements in full, free to apply to anything, including to this assessment.