Does the approver see the arguments that execute?

An approval is worth what it was an approval of. If the arguments a person was shown are not the arguments that run, then they approved something that did not happen, and the record of their approval describes an action nobody agreed to.

That is the failure mode approval exists to prevent, so it is worth knowing which frameworks prevent it. Six were read in published source at a pinned commit, against one question.

Between the moment a human approver is shown a tool call and the moment that tool runs, can the arguments change?

What the reading found

FrameworkVerdictWhy
OpenAI Agents SDKbound The approval item and the invocation read the same tool call. Nothing between them rewrites it.
Pydantic AIpartial Validation runs after approval and can transform the arguments, but the approver can state exactly what should run.
Haystackunbound Values named by inputs_from_state are injected after the confirmation hook has run.
LangGraphunbound Resuming re-executes the whole node, so a value derived from live state can differ the second time. Documented.
AutoGenno boundary No approval pause in core. Its maintainers have it open as two issues.
CrewAIno boundary No approval pause and no before-tool hook.

A framework with no approval pause is not failing this. It is not an approval gate, and marking it down for lacking one would be the same error as marking a vector store down for not authorising actions. Both are recorded because the absence is the relevant fact, and in both cases their own maintainers have it open as a request.

The three shapes this takes

Bound. The OpenAI Agents SDK builds its approval item from the model's tool call and invokes the tool with the arguments off that same object. There is no hook in between that could rewrite them. It also carries a setting for whether input guardrails run before approval, which is evidence that the ordering was considered rather than left to fall out.

A mechanism, not a default. Pydantic AI hands the approver the model's call and then validates it after approval, and validation can transform. What saves it from the third category is that an approver can supply the arguments that should run, which binds the two when it is used. Their AI lead has this open, and describes it exactly: the human approves one thing and the system executes another.

Unbound. In Haystack, a tool that reads from state executes with arguments the confirming person never saw: the confirmation runs as a before-tool hook over what the model produced, and the state values are injected afterwards. In LangGraph the mechanism is different and the result is the same. Resuming an interrupt re-runs the node from the start, so anything computed before the pause is computed again, and a value that reads live state can legitimately differ.

The LangGraph case deserves its distinction. It is unbound, and it is documented as unbound, in the framework's own docstring. A reader who checks is told. That is a materially different thing from a silent divergence, even though it leaves the same problem on the caller's desk.

Nobody has been comparing notes

Three of these are open issues right now, filed by three different people who do not cite each other. One is a maintainer describing it as a bug. One is a user who found it through inputs_from_state. The third is a design thread that arrived at the requirement from the other direction, asking for a canonical digest of the validated arguments so that a changed digest invalidates the decision.

That is the same defect being discovered separately in three places, which is usually what a missing shared idea looks like. The idea here is small: an approval is a claim about a specific action, so the record of it has to fix which action, not merely that permission was given.

What this does not show

One question, read once, in one direction. It is not a conformance assessment and it says nothing about whether any of these is well built. A framework that comes out badly here may be excellent at what it was made for, and none of them claims to be an authorisation layer.

Only paths reachable from each framework's documented approval feature were read. A deployment that builds its own gate is outside this entirely, and may be bound or unbound regardless of what its framework does.

And a reading is wrong in the ordinary way readings are wrong. Every verdict cites a file and a line at a commit precisely so that being wrong is cheap to demonstrate.

The second reading, 9 September 2026

The question above is one leg of a wider one, and five parties arrived at the wider version within a week of each other without citing one another. So the same six frameworks were read again against five requirements rather than one. B2 below is the original question, carried over unchanged.

B1Is the material shown to the reviewer persisted, or only displayed?
B2Between being shown a call and it running, can the arguments change?
B3Can a reviewer who MODIFIED an action be told apart from one who approved it as proposed?
B4Is an approval bound to the specific call it permitted, or to the tool?
B5Does the record distinguish an action that was permitted from one observed to run?
FrameworkB1B2B3B4B5
LangGraphpresentabsentpresentpartialpartial
OpenAI Agents SDKpartialpresentpresentpresentpartial
Pydantic AIpartialpartialabsentpresentpresent
Haystackpartialabsentpartialpartialabsent
CrewAIabsentabsentabsentabsentpartial
AutoGenno pauseno pauseno pauseno pauseno pause

Two rows of the first reading above are now wrong, and they are corrected here rather than quietly left. Re-reading found both. Nobody reported them.

CrewAI was recorded as having no approval boundary at all. As of 1.15.20 it has before_tool_call hooks that can block execution, so the questions now arise. A hook is handed tool_input as a mutable dict, mutations persist even when it blocks, and it returns one boolean. Reproduced: shown 0 and 0, mutated in place to 41 and 1, executed 41 and 1, with nothing recording that a mutation happened.

LangGraph is close to the reverse of the impression the first reading leaves. The checkpoint keeps the whole interrupt as a pending write, so what was shown is retained verbatim, and the typed resume value sits on the same checkpoint written by the same task. It is also the only one of the six that refuses an ambiguous approval: with several interrupts pending it will not accept a bare resume value and names the requirement to specify which. B4 is partial only because the interrupt id hashes the graph position and not the arguments.

Pydantic AI is the weakest on B3 and the strongest on B5, which is the argument for reading requirement by requirement instead of forming an impression. Its message history states arguments that did not execute, and it is the only subject whose record can say it does not know whether an effect occurred.

Check the second reading

Every verdict, its locator, and which were reproduced rather than reasoned about are in census/binding/readings-2.json, against the rubric in binding/rubric.py. A suite holds the two to each other.

Check it

The commit, the two locators and the reasoning for each subject are in census/binding/readings.json. The counts on this page are recomputed from that file by the test suite rather than typed.

The related question of whether a system can say who approved, rather than what they approved, is the conformance census. Of the eight systems there that take or gate actions, one can name the person.

Citing this

The readings are CC BY 4.0, which asks for attribution, so the reference is here rather than left to be composed. This block is generated from the page it sits on, so a date that moves here moves in the citation too.

Clifford, T. (2026). Human in the loop approval: does the approver see the arguments that execute. Machine Testimony. https://machinetestimony.org/approval-binding/
@misc{clifford2026approvalbinding,
  author = {Clifford, Troy},
  title  = {Human in the loop approval: does the approver see the arguments that execute},
  year   = {2026},
  note   = {Machine Testimony, read 6 September 2026},
  url    = {https://machinetestimony.org/approval-binding/},
}

This page carries no DOI. It cites its dated URL, and saying so is the point: a citation naming a deposit that does not exist is worse than one naming a page that does.