What Haystack records, and what it does not

Read at82da3ad
Repositorygithub.com/deepset-ai/haystack
LicenceApache-2.0
Assessed asstores, derives, acts
Reachesno level yet
Read byTroy Brandon Clifford
LevelMeetsWhat the level asks
TR-1 Recorded0/5The record exists and is append-only.
TR-2 Explained0/5Every belief resolves to its evidence, and disagreements survive.
TR-3 Gated3/7Actions carry a verdict, and approvals carry a name.
TR-4 Verifiable0/3The record can be shown not to have changed.

A count is requirements fully met out of those that apply. This system is assessed as stores, derives, acts, and requirements outside that are not counted against it.

Haystack is the first subject assessed as storing, deriving and acting, so all twenty requirements apply. It stores documents in a document store and messages in a chat history, it infers metadata about documents with a language model through LLMMetadataExtractor, and it gates tool calls through a human-in-the-loop hook.

It has the best developed approval object in the census and the weakest storage guarantees. ToolExecutionDecision carries the verdict, the tool call id, a feedback field meant for the reason and the final parameters, it serialises, and _classify_decision returns a closed literal of confirm, modify and reject. Nothing else assessed distinguishes an approver who changed the arguments from one who accepted them. A rejected call leaves the same two message types in the chat history that a permitted call leaves.

The storage side is the opposite. Neither a Document nor a ChatMessage carries a timestamp. Writing at an existing id under DuplicatePolicy.OVERWRITE deletes the previous document before writing the new one, delete_documents removes a document with no trace, and the approval path rebuilds the chat history rather than appending to it, which means an approver who modifies a tool call leaves a record in which their arguments are the only ones present.

This is the first subject to score nothing at TR-1, and that number needs reading carefully. A Haystack document store is a search index, and an index is built to answer what is true now rather than to say what was true in March. Overwriting in place and carrying no write time are reasonable properties for what it is. The census records them because a deployer who reaches for that store as the place their agent remembers things has acquired a retention obligation the store was never built to carry, and nothing in the library will tell them so. The same reasoning is why frameworks with no approval pause are not marked down for lacking one.

Two findings are worth separating from the scores. The first is that LLMMetadataExtractor merges what a model inferred into the same meta dict that holds what the document was told, by plain dict assignment, so an inferred value can overwrite a told one and nothing marks either. That is the requirement the derives claim exists to test. The second is that a Document's id is a sha256 over its own content, which is the closest thing to tamper evidence anywhere in this census, and the framework's own enrichment path routinely produces documents whose id no longer matches their content. A mismatch therefore proves nothing, and the distance to a real answer is a decision about what the id is for rather than a new mechanism.

Every verdict, and what it rests on

Twenty requirements, each stated as a capability rather than a format, so a system that holds the information in its own shape counts as having it. An absent verdict cites where the assessor looked and did not find it, which is the difference between a measurement and an accusation.

TR-1 Recorded

The record exists and is append-only.

R1.1 · absent. When the system stores a fact, does it durably record when that happened?

The clearest absent on R1.1 in the census. A time can always be put in the meta dict by an application that thinks to, but nothing in the framework does it, so two documents written a year apart are indistinguishable once stored.

  • source haystack/dataclasses/document.py:48-54
    a Document is id, content, blob, meta, score, embedding and sparse_embedding. There is no field for when it was written
  • source haystack/dataclasses/chat_message.py:291-294
    a ChatMessage is _role, _content, _name and _meta. No timestamp either, so neither the document store nor the conversation records when anything happened
  • searched grep -rniE 'timestamp|created_at|written_at' haystack/dataclasses/*.py
    six hits, all in breakpoints.py, where PipelineSnapshot.timestamp is datetime | None = None and is set at breakpoint.py:261 when a pipeline is paused for debugging. That is a debugging artefact of a run, not a write time stored with a fact

R1.2 · absent. When a stored fact changes, is the previous version still readable?

Both storage surfaces answer this the same way. The document store overwrites by deleting first, and the approval path rebuilds the chat history rather than appending to it. In neither case does the framework retain what was there before.

  • source haystack/document_stores/in_memory/document_store.py:491-498
    under DuplicatePolicy.OVERWRITE the store calls delete_documents on the existing id and then assigns the new document to the same key. The previous version is gone before the new one is written
  • source haystack/hooks/human_in_the_loop/strategies.py:663
    on the conversation side _update_chat_history returns chat_history[:insertion_point + 1] plus the new messages, so the assistant message carrying what the model originally proposed is dropped rather than kept beside its replacement
  • searched grep -rniE 'def .*(version|history|revision|previous|supersed)' haystack/document_stores/
    zero hits. No document store exposes a previous version, a revision or any way to ask what a document said before it was overwritten

R1.3 · absent. Does the ordinary write path ever destroy what was previously recorded?

This is the requirement Haystack is furthest from, and the approval path is the sharpest instance. An approver who changes a tool call leaves a record in which their arguments are the only ones present. A reader six months later sees a proposal that matches what ran, because the proposal was rewritten to match.

  • source haystack/document_stores/in_memory/document_store.py:491-492
    the routine write path destroys: if document.id in self.storage, delete_documents([document.id]) runs before the write. This is the ordinary call, not a maintenance operation
  • source haystack/hooks/human_in_the_loop/strategies.py:619
    a modified tool call is written as replace(tc, arguments=final_args) into a rebuilt history, so the arguments the model proposed are replaced in place by the ones the approver allowed. A user message explaining the change is added, in prose, but the original arguments are not kept
  • source haystack/document_stores/in_memory/document_store.py:507-516
    delete_documents is a routine public method that does del self.storage[doc_id], with no soft delete and no marker left behind
  • searched grep -rniE 'append_only|soft.delete|immutable|is_deleted|deleted_at' haystack
    zero hits across all 286 modules. The DocumentStore protocol at document_stores/types/protocol.py:109, 127 has exactly two mutating methods, write_documents and delete_documents, and neither has a non-destructive mode

R1.4 · partial. Are entries distinguishable by kind, or is everything one undifferentiated blob of text?

Typed where the conversation is, untyped where the facts are stored. The split matters because this subject claims to store, and the store is the half with one type and a free-form dict.

  • source haystack/dataclasses/chat_message.py:204
    ChatMessageContentT is a union of TextContent, ToolCall, ToolCallResult, ImageContent, ReasoningContent and FileContent, so on the conversation side an assertion, an action and a result are genuinely separate types
  • source haystack/dataclasses/chat_message.py:19-34
    ChatRole is a closed enum of user, system, assistant and tool
  • source haystack/dataclasses/document.py:48-54
    the storage side has one type. Everything a document knows beyond its content lives in meta, an untyped dict[str, Any] in which a value the system was told and a value it inferred sit together

R1.5 · absent. When data must be destroyed for a legal reason, is the destruction itself recorded?

The ordinary shape of this finding across the census. Worth stating plainly because a deployer answering an erasure request has nothing to show a regulator except the absence itself.

  • source haystack/document_stores/in_memory/document_store.py:513-516
    delete_documents iterates the ids and does del self.storage[doc_id]. Nothing is written to say a deletion happened, and an id that was never present is skipped silently
  • searched grep -rniE 'audit|deletion_log|erasure|tombstone' haystack
    zero hits across all 286 modules. There is no audit surface of any kind, so an erasure carried out for a legal reason leaves the store identical to one where the document never existed

TR-2 Explained

Every belief resolves to its evidence, and disagreements survive.

R2.1 · partial. Can the source a stored fact came from be recovered from the store, by following a link rather than by guessing?

Two different answers again. A tool result names its call properly. A stored document names its source by a convention the type does not enforce, and the default converter setting reduces that name to a basename.

R2.2 · absent. Can a fact the system inferred be told apart from one it was told?

An inferred key also silently overwrites a told key of the same name, because the merge is plain dict assignment. This is the requirement the derives claim exists to test, and the answer is that once extraction has run, the two kinds of fact are the same kind of fact.

  • source haystack/components/extractors/llm_metadata_extractor.py:379
    new_meta[key] = parsed_metadata[key] merges what a language model inferred about a document straight into the same meta dict that holds what the document was given, with no prefix, namespace or marker
  • source haystack/components/extractors/llm_metadata_extractor.py:380-382
    the only markers it does write, metadata_extraction_error and metadata_extraction_response, are set on failure and removed on success, so a successful extraction leaves nothing at all to say it happened
  • searched grep -rniE 'derived|inferred|provenance|is_generated' haystack/dataclasses/ haystack/components/extractors/
    zero hits. No field anywhere marks a value as inferred rather than told

R2.3 · partial. When two stored facts about the same proposition disagree, do both survive?

Both survive incidentally, in the same way an append-only log preserves a contradiction: because nothing looked. The framework has no notion of two documents being about the same proposition, so it cannot be said to have kept them deliberately.

  • source haystack/dataclasses/document.py:103-116
    the id is a sha256 over content, meta and embeddings, so two documents that disagree have different ids and both can be written and both survive
  • source haystack/document_stores/in_memory/document_store.py:491-498
    unless one is written under DuplicatePolicy.OVERWRITE at the same id, in which case the earlier one is deleted first
  • searched grep -rniE '\bconflict|contradict|disagree' haystack
    ten hits across two modules and none is about stored facts. They are prompt variables clashing with state schema names in agent.py and input and output name collisions in super_component.py

R2.4 · absent. Is the disagreement itself queryable, or must a reader diff rows to notice it?

A reader has to find contradictions by reading. That is expected of a document store and is recorded because the requirement asks.

  • searched grep -rniE '\bconflict|contradict|disagree' haystack
    ten hits, all of them pipeline wiring: variable name clashes in agent.py:580-585 and type or output name collisions in super_component.py:248, 344. Nothing models a disagreement between stored facts
  • source haystack/dataclasses/document.py:48-54
    there is no conflict type and no field on a document that could point at another document it contradicts, so a disagreement is not something the store can be asked about

R2.5 · absent. When a conflict is resolved, does the record say who resolved it and by what method?

Scored absent rather than under the clause that credits a system for never resolving conflicts, because that clause is for a system which leaves a contradiction standing as a visible unresolved state. Nothing here is standing, because nothing was represented as a contradiction in the first place.

  • searched grep -rniE '\bconflict|contradict|disagree' haystack
    the same ten wiring hits. No resolution is recorded because no conflict is represented, so there is no actor and no method to record
  • source haystack/document_stores/in_memory/document_store.py:491-498
    where a disagreement is settled it is settled silently and by recency: writing at the same id under OVERWRITE deletes the earlier document and leaves the later one looking uncontested

TR-3 Gated

Actions carry a verdict, and approvals carry a name.

R3.1 · present. Does a consequential action produce a durable entry whether or not it ran?

Among the strongest answers here. An action that was stopped leaves the same kinds of message an action that ran leaves, in the history rather than in a log. Haystack is also the only subject that distinguishes an approver who changed the arguments from one who accepted them, which is precisely the distinction the approval binding reading was asking about.

R3.2 · present. Does an action's risk class come from somewhere the proposing model cannot write to?

Present with the same qualification recorded for Pydantic AI. The supplied policies ignore the model entirely and route on the tool name. A deployment can write a policy that reads the arguments, and the framework neither does that nor recommends it.

R3.3 · present. Are refusals recorded as faithfully as permissions?

A refusal has the same standing as a permission because it is made of the same message types. What it does not have is a distinguishing mark of its own, which is R3.4's problem rather than this one.

  • source haystack/hooks/human_in_the_loop/strategies.py:601-607
    the rejection is recorded as an assistant message plus a tool message, the same two message types a permitted call produces, and it is placed in the chat history by _update_chat_history rather than emitted to a logger
  • source haystack/hooks/human_in_the_loop/strategies.py:298-304
    the decision is additionally written to a tracing span as haystack.agent.hook.human_in_the_loop.strategy.decision with the whole decision as a content tag, which is observability rather than record, but it means a refusal is visible in two places

R3.4 · partial. Does a refusal record why it was refused?

Ahead of most of the census for having a field meant for the reason, and short of the requirement because the field is free text and the machine-readable part, error=True, conflates refusal with failure.

  • source haystack/hooks/human_in_the_loop/dataclasses.py:53
    ToolExecutionDecision.feedback is a dedicated field for the reason, documented as carrying it when execution is rejected. Most of the census has no such field at all
  • source haystack/hooks/human_in_the_loop/strategies.py:601
    tool_result_text = ted.feedback or REJECTION_FEEDBACK_TEMPLATE.format(tool_name=tc.tool_name), so a refusal with no reason still produces prose, and the reason reaches the record as free text in a tool result
  • searched grep -rniE 'reason|rejection_code|denial' haystack/hooks/human_in_the_loop/dataclasses.py
    the only structured marker on the resulting message is error=True, which a tool that raised also sets, so the record cannot distinguish a call a person refused from one that crashed

R3.5 · absent. Does an approval identify a person or a named role holder?

The same stopping point as every other subject with a gate. The decision object is the most detailed in this census and still has nowhere to put the person, so an approval that a console user typed and one a NeverAskPolicy produced are the same record.

  • source haystack/hooks/human_in_the_loop/dataclasses.py:9-27
    ConfirmationUIResult is action, feedback and new_tool_params. The object that carries what a person decided has no field for which person decided it
  • source haystack/hooks/human_in_the_loop/dataclasses.py:30-54
    ToolExecutionDecision, which is what survives the interaction, has tool_name, execute, tool_call_id, feedback and final_tool_params, and no approver either
  • searched grep -rniE 'approver|principal|user_id|approved_by|authenticat' haystack/hooks/
    two hits, and both are false positives on the substring user_id: last_user_idx and insertion_point in strategies.py:658 and 661. No identity of any kind appears on the approval path

R3.6 · absent. Does the approver's identity come from the authentication layer rather than from something the model can write?

Not reachable while R3.5 is absent: an identity that is never captured cannot be sourced from an authentication layer. Recorded separately because the two come apart in systems that do capture an approver and then trust the request body for the name.

  • source haystack/hooks/human_in_the_loop/user_interfaces.py:22,124
    the supplied interfaces are console based, RichConsoleUI at 22 and SimpleConsoleUI at 124, and read a decision from whoever is at the terminal. There is no session, token or authenticated caller in the picture
  • searched grep -rniE 'approver|principal|user_id|approved_by|authenticat' haystack/hooks/
    two hits, both the last_user_idx false positive. Nothing on the approval path reads an identity from anywhere, authenticated or not

R3.7 · absent. Is the acting agent prevented from approving its own action?

With no principals on either side there is nothing to compare. The passthrough case is recorded because it is a decision the framework makes on its own behalf, and it is a defensible design choice: it lets the tool resolution layer report an unknown tool uniformly instead of failing inside the hook.

  • source haystack/hooks/human_in_the_loop/policies.py:26-33
    NeverAskPolicy is a supplied policy whose should_ask returns false, so a configuration in which nothing is ever put to a person is a first-class option rather than a misuse
  • source haystack/hooks/human_in_the_loop/strategies.py:232-245
    _passthrough_tool_call builds a decision with execute=True and the model's own arguments for calls that do not resolve to a known tool, so some calls are approved by the framework itself
  • searched grep -rniE 'separation|segregat|self.approv|four.eyes' haystack
    nothing expresses a separation between the party proposing an action and the party allowing it, which follows from neither party being represented

TR-4 Verifiable

The record can be shown not to have changed.

R4.1 · absent. Does the system publish a scheme under which the record's past state can be verified?

Closer to the shape of TR-4 than most subjects and still absent, because a content hash used as a primary key is not a published scheme for verifying a past state. It is the same observation recorded against LangGraph: the structure is most of the way there and nothing is built on it.

  • source haystack/dataclasses/document.py:103-116
    the primitive exists. _create_id builds a sha256 over content, blob, mime type, sorted meta and embeddings, and the algorithm is in published source, so anybody can recompute it
  • source haystack/dataclasses/document.py:68
    but it is identity, not integrity: self.id = self.id or self._create_id() means a caller-supplied id is kept as given, and the hash is only computed when one is absent
  • searched grep -rniE 'hmac|merkle|tamper|signature|verify_integrity' haystack
    85 hits and not one is an integrity mechanism: 84 are inspect.signature and the single tamper match is the substring inside structlog's TimeStamper at logging.py:386. hmac, merkle and verify_integrity return nothing

R4.2 · absent. Can an independent party run that verification without the vendor's cooperation, and without the operator's?

Nothing here needs deepset's cooperation, because there is nothing to withhold. Scored absent on the same basis as every other subject: absence of a scheme rather than a restriction on access to one.

  • searched grep -rn '_create_id' haystack
    two hits: the definition and the single call at document.py:68. The hash is never recomputed and never compared, so there is no verification for an independent party to run
  • source haystack/document_stores/in_memory/document_store.py:462
    the store's public surface is write, delete, filter and count. It offers no verify operation, so the question of whether one needs the vendor does not arise

R4.3 · absent. Would alteration of a past entry be detectable after the fact?

The most interesting negative in this assessment. A digest over the content exists, and the ordinary enrichment path produces documents that fail it, which means a mismatch cannot be read as evidence of tampering even by somebody who thought to check. The distance to a real answer here is not a new mechanism but a decision about what the id is for.

  • source haystack/dataclasses/document.py:103-116
    in principle an altered document would no longer hash to its id, which is the closest any subject in this census comes to alteration being detectable at rest
  • source haystack/components/extractors/llm_metadata_extractor.py:379-383
    in practice the framework breaks that relationship itself. replace(document, meta=new_meta) keeps the original id while changing the meta the id was computed over, so a document whose id does not match its content is the normal result of enrichment
  • searched grep -rn '_create_id' haystack
    and nothing recomputes it, so no code path would notice either the legitimate mismatch or a malicious one

How this was read

Read of haystack/dataclasses/document.py and chat_message.py, haystack/document_stores/in_memory/document_store.py, haystack/hooks/human_in_the_loop/ (dataclasses.py, hooks.py, policies.py, strategies.py) and haystack/components/extractors/llm_metadata_extractor.py at 82da3adc2fac4675b80ff5573b790ec07113697b, the commit the approval binding reading already pins, so the two agree on which source they describe. The source archive for that commit was downloaded rather than an installed release, so every line number refers to the pinned commit and every search recorded below was run over all 286 modules of the package. That commit is main, dated 4 September 2026, one day after the v3.1.1 release, which is why the version is written as 3.1.1+ rather than as a release.

Found an error? Challenge a finding

Then it is wrong in the ordinary way readings are wrong, and every verdict cites a file and a line at a pinned commit precisely so that being wrong is cheap to demonstrate. The remedy is a pull request against the subject file, and it does not involve persuading anybody. Nobody applied for this and it is not a certification.

The standing register carries every system side by side, and the rubric is the twenty requirements in full, free to apply to anything, including to this assessment.