Assessment, 2026-09-04 · main at 92eb5f9
| Read at | 92eb5f9 |
|---|---|
| Repository | github.com/crewAIInc/crewAI |
| Licence | MIT |
| Assessed as | stores, derives, acts |
| Reaches | no level yet |
| Read by | Troy Brandon Clifford |
| Level | Meets | What the level asks |
|---|---|---|
| TR-1 Recorded | 1/5 | The record exists and is append-only. |
| TR-2 Explained | 0/5 | Every belief resolves to its evidence, and disagreements survive. |
| TR-3 Gated | 1/7 | Actions carry a verdict, and approvals carry a name. |
| TR-4 Verifiable | 0/3 | The record can be shown not to have changed. |
A count is requirements fully met out of those that apply. This system is assessed as stores, derives, acts, and requirements outside that are not counted against it.
CrewAI has a real gate and a real consolidation step, and both are worth understanding before deploying it where a record matters. The gate is the before_tool_call hook: developer code, filtered by tool name, that can block a call, which satisfies the requirement that risk not be decided by the model. What it does with a block is the problem. HookAborted carries a reason, run_before_tool_call_hooks discards it, and the executor substitutes the fixed string 'Tool execution blocked by hook', which goes back to the model as an ordinary tool result. The refusal survives as narration in a transcript rather than as a decision anybody can query. On the memory side, an LLM emits a keep/update/delete plan over existing records and updates overwrite content in place while preserving the record id and created_at, with no history kept anywhere. Both behaviours are reasonable for a framework optimising for agents that work; neither leaves a record.
Twenty requirements, each stated as a capability rather than a format, so a system that holds the information in its own shape counts as having it. An absent verdict cites where the assessor looked and did not find it, which is the difference between a measurement and an accusation.
The record exists and is append-only.
R1.1 · present. When the system stores a fact, does it durably record when that happened?
R1.2 · absent. When a stored fact changes, is the previous version still readable?
Once a record is updated, what it previously said is gone. Preserving created_at while replacing content is worse than either alternative for a reader: the record looks as old as the original claim while carrying the new one.
R1.3 · absent. Does the ordinary write path ever destroy what was previously recorded?
R1.4 · partial. Are entries distinguishable by kind, or is everything one undifferentiated blob of text?
There is more than one kind of record in the system, but within memory a claim and whatever produced it are the same shape, distinguished only by whatever a caller happened to put in the metadata dict.
R1.5 · absent. When data must be destroyed for a legal reason, is the destruction itself recorded?
Every belief resolves to its evidence, and disagreements survive.
R2.1 · partial. Can the source a stored fact came from be recovered from the store, by following a link rather than by guessing?
A user id or a session id narrows the origin to a person or a conversation. It does not identify the message, document or event the claim came from, so recovering that means searching the session by hand if it still exists.
R2.2 · partial. Can a fact the system inferred be told apart from one it was told?
The metadata is labelled and the content is not, which is the wrong way round for this purpose. After consolidation a record can hold model-written text while presenting as the original observation.
R2.3 · partial. When two stored facts about the same proposition disagree, do both survive?
Two disagreeing memories both persist until the consolidation step runs over them, at which point the plan may delete one or overwrite it. This is the textbook case for a partial: survival is real but conditional on a step that runs automatically.
R2.4 · absent. Is the disagreement itself queryable, or must a reader diff rows to notice it?
R2.5 · partial. When a conflict is resolved, does the record say who resolved it and by what method?
The reason is computed and then not stored. That is a narrower gap than not having one, and closing it would mean carrying insert_reason onto the updated record rather than discarding it after the branch.
Actions carry a verdict, and approvals carry a name.
R3.1 · partial. Does a consequential action produce a durable entry whether or not it ran?
A stopped attempt does leave a mark, which is why this is not absent. What it leaves is a sentence in the model's context, structurally identical to a tool that ran and returned that text. Nothing an auditor could query separates the two.
R3.2 · present. Does an action's risk class come from somewhere the proposing model cannot write to?
Coarser than a graded risk class, but the substance holds: which tools are gated is fixed in code the model cannot write.
R3.3 · partial. Are refusals recorded as faithfully as permissions?
A refusal is recorded in the same channel as a permission, which is the letter of this requirement, but with no marker that it was a refusal rather than an output.
R3.4 · absent. Does a refusal record why it was refused?
This is the most fixable finding in this file. The reason is constructed and then thrown away two frames later. Returning it instead of a boolean would give every refusal a machine-readable cause at no design cost.
R3.5 · absent. Does an approval identify a person or a named role holder?
The approval is a line entered at a terminal. Whoever is sitting at that terminal is the approver, and the system records neither the fact nor the identity.
R3.6 · absent. Does the approver's identity come from the authentication layer rather than from something the model can write?
R3.7 · absent. Is the acting agent prevented from approving its own action?
The record can be shown not to have changed.
R4.1 · absent. Does the system publish a scheme under which the record's past state can be verified?
R4.2 · absent. Can an independent party run that verification without the vendor's cooperation, and without the operator's?
R4.3 · absent. Would alteration of a past entry be detectable after the fact?
Read of lib/crewai/src/crewai/memory/types.py, memory/encoding_flow.py, memory/analyze.py, memory/storage/lancedb_storage.py, hooks/decorators.py, hooks/tool_hooks.py, hooks/llm_hooks.py and agents/crew_agent_executor.py at 92eb5f9, cloned from the public repository. Line numbers are that commit's. The project versions from SCM and the clone carries no tag, so the commit is the identifier.
Then it is wrong in the ordinary way readings are wrong, and every verdict cites a file and a line at a pinned commit precisely so that being wrong is cheap to demonstrate. The remedy is a pull request against the subject file, and it does not involve persuading anybody. Nobody applied for this and it is not a certification.
The standing register carries every system side by side, and the rubric is the twenty requirements in full, free to apply to anything, including to this assessment.