Assessment, 2026-09-06 · 3.1.1+ (82da3ad)
| Read at | 82da3ad |
|---|---|
| Repository | github.com/deepset-ai/haystack |
| Licence | Apache-2.0 |
| Assessed as | stores, derives, acts |
| Reaches | no level yet |
| Read by | Troy Brandon Clifford |
| Level | Meets | What the level asks |
|---|---|---|
| TR-1 Recorded | 0/5 | The record exists and is append-only. |
| TR-2 Explained | 0/5 | Every belief resolves to its evidence, and disagreements survive. |
| TR-3 Gated | 3/7 | Actions carry a verdict, and approvals carry a name. |
| TR-4 Verifiable | 0/3 | The record can be shown not to have changed. |
A count is requirements fully met out of those that apply. This system is assessed as stores, derives, acts, and requirements outside that are not counted against it.
Haystack is the first subject assessed as storing, deriving and acting, so all twenty requirements apply. It stores documents in a document store and messages in a chat history, it infers metadata about documents with a language model through LLMMetadataExtractor, and it gates tool calls through a human-in-the-loop hook.
It has the best developed approval object in the census and the weakest storage guarantees. ToolExecutionDecision carries the verdict, the tool call id, a feedback field meant for the reason and the final parameters, it serialises, and _classify_decision returns a closed literal of confirm, modify and reject. Nothing else assessed distinguishes an approver who changed the arguments from one who accepted them. A rejected call leaves the same two message types in the chat history that a permitted call leaves.
The storage side is the opposite. Neither a Document nor a ChatMessage carries a timestamp. Writing at an existing id under DuplicatePolicy.OVERWRITE deletes the previous document before writing the new one, delete_documents removes a document with no trace, and the approval path rebuilds the chat history rather than appending to it, which means an approver who modifies a tool call leaves a record in which their arguments are the only ones present.
This is the first subject to score nothing at TR-1, and that number needs reading carefully. A Haystack document store is a search index, and an index is built to answer what is true now rather than to say what was true in March. Overwriting in place and carrying no write time are reasonable properties for what it is. The census records them because a deployer who reaches for that store as the place their agent remembers things has acquired a retention obligation the store was never built to carry, and nothing in the library will tell them so. The same reasoning is why frameworks with no approval pause are not marked down for lacking one.
Two findings are worth separating from the scores. The first is that LLMMetadataExtractor merges what a model inferred into the same meta dict that holds what the document was told, by plain dict assignment, so an inferred value can overwrite a told one and nothing marks either. That is the requirement the derives claim exists to test. The second is that a Document's id is a sha256 over its own content, which is the closest thing to tamper evidence anywhere in this census, and the framework's own enrichment path routinely produces documents whose id no longer matches their content. A mismatch therefore proves nothing, and the distance to a real answer is a decision about what the id is for rather than a new mechanism.
Twenty requirements, each stated as a capability rather than a format, so a system that holds the information in its own shape counts as having it. An absent verdict cites where the assessor looked and did not find it, which is the difference between a measurement and an accusation.
The record exists and is append-only.
R1.1 · absent. When the system stores a fact, does it durably record when that happened?
The clearest absent on R1.1 in the census. A time can always be put in the meta dict by an application that thinks to, but nothing in the framework does it, so two documents written a year apart are indistinguishable once stored.
R1.2 · absent. When a stored fact changes, is the previous version still readable?
Both storage surfaces answer this the same way. The document store overwrites by deleting first, and the approval path rebuilds the chat history rather than appending to it. In neither case does the framework retain what was there before.
R1.3 · absent. Does the ordinary write path ever destroy what was previously recorded?
This is the requirement Haystack is furthest from, and the approval path is the sharpest instance. An approver who changes a tool call leaves a record in which their arguments are the only ones present. A reader six months later sees a proposal that matches what ran, because the proposal was rewritten to match.
R1.4 · partial. Are entries distinguishable by kind, or is everything one undifferentiated blob of text?
Typed where the conversation is, untyped where the facts are stored. The split matters because this subject claims to store, and the store is the half with one type and a free-form dict.
R1.5 · absent. When data must be destroyed for a legal reason, is the destruction itself recorded?
The ordinary shape of this finding across the census. Worth stating plainly because a deployer answering an erasure request has nothing to show a regulator except the absence itself.
Every belief resolves to its evidence, and disagreements survive.
R2.1 · partial. Can the source a stored fact came from be recovered from the store, by following a link rather than by guessing?
Two different answers again. A tool result names its call properly. A stored document names its source by a convention the type does not enforce, and the default converter setting reduces that name to a basename.
R2.2 · absent. Can a fact the system inferred be told apart from one it was told?
An inferred key also silently overwrites a told key of the same name, because the merge is plain dict assignment. This is the requirement the derives claim exists to test, and the answer is that once extraction has run, the two kinds of fact are the same kind of fact.
R2.3 · partial. When two stored facts about the same proposition disagree, do both survive?
Both survive incidentally, in the same way an append-only log preserves a contradiction: because nothing looked. The framework has no notion of two documents being about the same proposition, so it cannot be said to have kept them deliberately.
R2.4 · absent. Is the disagreement itself queryable, or must a reader diff rows to notice it?
A reader has to find contradictions by reading. That is expected of a document store and is recorded because the requirement asks.
R2.5 · absent. When a conflict is resolved, does the record say who resolved it and by what method?
Scored absent rather than under the clause that credits a system for never resolving conflicts, because that clause is for a system which leaves a contradiction standing as a visible unresolved state. Nothing here is standing, because nothing was represented as a contradiction in the first place.
Actions carry a verdict, and approvals carry a name.
R3.1 · present. Does a consequential action produce a durable entry whether or not it ran?
Among the strongest answers here. An action that was stopped leaves the same kinds of message an action that ran leaves, in the history rather than in a log. Haystack is also the only subject that distinguishes an approver who changed the arguments from one who accepted them, which is precisely the distinction the approval binding reading was asking about.
R3.2 · present. Does an action's risk class come from somewhere the proposing model cannot write to?
Present with the same qualification recorded for Pydantic AI. The supplied policies ignore the model entirely and route on the tool name. A deployment can write a policy that reads the arguments, and the framework neither does that nor recommends it.
R3.3 · present. Are refusals recorded as faithfully as permissions?
A refusal has the same standing as a permission because it is made of the same message types. What it does not have is a distinguishing mark of its own, which is R3.4's problem rather than this one.
R3.4 · partial. Does a refusal record why it was refused?
Ahead of most of the census for having a field meant for the reason, and short of the requirement because the field is free text and the machine-readable part, error=True, conflates refusal with failure.
R3.5 · absent. Does an approval identify a person or a named role holder?
The same stopping point as every other subject with a gate. The decision object is the most detailed in this census and still has nowhere to put the person, so an approval that a console user typed and one a NeverAskPolicy produced are the same record.
R3.6 · absent. Does the approver's identity come from the authentication layer rather than from something the model can write?
Not reachable while R3.5 is absent: an identity that is never captured cannot be sourced from an authentication layer. Recorded separately because the two come apart in systems that do capture an approver and then trust the request body for the name.
R3.7 · absent. Is the acting agent prevented from approving its own action?
With no principals on either side there is nothing to compare. The passthrough case is recorded because it is a decision the framework makes on its own behalf, and it is a defensible design choice: it lets the tool resolution layer report an unknown tool uniformly instead of failing inside the hook.
The record can be shown not to have changed.
R4.1 · absent. Does the system publish a scheme under which the record's past state can be verified?
Closer to the shape of TR-4 than most subjects and still absent, because a content hash used as a primary key is not a published scheme for verifying a past state. It is the same observation recorded against LangGraph: the structure is most of the way there and nothing is built on it.
R4.2 · absent. Can an independent party run that verification without the vendor's cooperation, and without the operator's?
Nothing here needs deepset's cooperation, because there is nothing to withhold. Scored absent on the same basis as every other subject: absence of a scheme rather than a restriction on access to one.
R4.3 · absent. Would alteration of a past entry be detectable after the fact?
The most interesting negative in this assessment. A digest over the content exists, and the ordinary enrichment path produces documents that fail it, which means a mismatch cannot be read as evidence of tampering even by somebody who thought to check. The distance to a real answer here is not a new mechanism but a decision about what the id is for.
Read of haystack/dataclasses/document.py and chat_message.py, haystack/document_stores/in_memory/document_store.py, haystack/hooks/human_in_the_loop/ (dataclasses.py, hooks.py, policies.py, strategies.py) and haystack/components/extractors/llm_metadata_extractor.py at 82da3adc2fac4675b80ff5573b790ec07113697b, the commit the approval binding reading already pins, so the two agree on which source they describe. The source archive for that commit was downloaded rather than an installed release, so every line number refers to the pinned commit and every search recorded below was run over all 286 modules of the package. That commit is main, dated 4 September 2026, one day after the v3.1.1 release, which is why the version is written as 3.1.1+ rather than as a release.
Then it is wrong in the ordinary way readings are wrong, and every verdict cites a file and a line at a pinned commit precisely so that being wrong is cheap to demonstrate. The remedy is a pull request against the subject file, and it does not involve persuading anybody. Nobody applied for this and it is not a certification.
The standing register carries every system side by side, and the rubric is the twenty requirements in full, free to apply to anything, including to this assessment.