Assessment, 2026-09-04 · 2.0.20 (9a7924b)
| Read at | 9a7924b |
|---|---|
| Repository | github.com/mem0ai/mem0 |
| Licence | Apache-2.0 |
| Assessed as | stores, derives |
| Reaches | no level yet |
| Read by | Troy Brandon Clifford |
| Level | Meets | What the level asks |
|---|---|---|
| TR-1 Recorded | 3/5 | The record exists and is append-only. |
| TR-2 Explained | 0/5 | Every belief resolves to its evidence, and disagreements survive. |
| TR-3 Gated | n/a | Actions carry a verdict, and approvals carry a name. |
| TR-4 Verifiable | 0/3 | The record can be shown not to have changed. |
A count is requirements fully met out of those that apply. This system is assessed as stores, derives, and requirements outside that are not counted against it.
mem0 is a memory layer built for retrieval quality and low integration cost, and its update phase consolidates on purpose: when new information contradicts an existing memory, the prompt at configs/prompts.py:264 instructs the model to delete the old one. Most of the gaps here follow directly from that goal rather than being oversights against it, and mem0 has never claimed to be an audit record. What the assessment says is narrow: a deployment that later needs to show what the agent believed and why will not find those facts here, because the design traded them away for something else it wanted more. Two findings are worth the maintainers' attention independently of any of that, and both are noted on the requirements they sit under: R1.5, where a deleted memory's text is retained in the local history database, and R2.2, where nothing distinguishes an LLM-extracted fact from a message stored verbatim.
Twenty requirements, each stated as a capability rather than a format, so a system that holds the information in its own shape counts as having it. An absent verdict cites where the assessor looked and did not find it, which is the difference between a measurement and an accusation.
The record exists and is append-only.
R1.1 · present. When the system stores a fact, does it durably record when that happened?
R1.2 · present. When a stored fact changes, is the previous version still readable?
R1.3 · partial. Does the ordinary write path ever destroy what was previously recorded?
The routine write path overwrites. A correction is an edit to the existing record rather than a new entry, and the previous state lives in a separate history database instead of in the store being queried. This is the requirement that keeps mem0 below TR-1, and it is a deliberate design choice rather than a defect: consolidating in place is what makes the retrieval surface small.
R1.4 · present. Are entries distinguishable by kind, or is everything one undifferentiated blob of text?
The kinds are separate. Whether they are linked is a different question, assessed at R2.1.
R1.5 · partial. When data must be destroyed for a legal reason, is the destruction itself recorded?
The deletion is recorded, which is half of what this asks. The other half is not met, and in a direction worth flagging to the maintainers: _delete_memory writes the deleted memory's own text into history.old_memory, so a deletion carried out to satisfy a privacy request leaves that content sitting in the local history database. Anyone relying on delete_all for erasure should know the text survives there. This is the one verdict in the census demonstrated by running the assessed system rather than by reading it, because it is a claim that something is written rather than that something is missing, and it concerns data somebody asked to have deleted.
Every belief resolves to its evidence, and disagreements survive.
R2.1 · partial. Can the source a stored fact came from be recovered from the store, by following a link rather than by guessing?
The raw source is kept but the fact does not point at it. Recovering which message produced a given memory means matching on actor, role and approximate time, which is guessing rather than following a link.
R2.2 · absent. Can a fact the system inferred be told apart from one it was told?
Both kinds of memory sit in one store looking identical. A consumer reading a mem0 memory cannot tell whether the model wrote it or a person did, which matters most for exactly the memories a model got wrong.
R2.3 · partial. When two stored facts about the same proposition disagree, do both survive?
Only one side survives as a belief. The other survives as a text column in a history row, which is an audit trail rather than a retained memory: a later query about that subject returns the winner and gives no sign there was ever a loser.
R2.4 · absent. Is the disagreement itself queryable, or must a reader diff rows to notice it?
A disagreement leaves an UPDATE or DELETE event indistinguishable from any other. Noticing that two sources ever disagreed means reading the history and inferring it.
R2.5 · partial. When a conflict is resolved, does the record say who resolved it and by what method?
What changed is recorded. Who decided and how is not: the decision is taken by an LLM against the prompt at configs/prompts.py, and neither the model, the prompt version nor the reason is written down anywhere.
Actions carry a verdict, and approvals carry a name.
R3.1, R3.2, R3.3, R3.4, R3.5, R3.6, R3.7 are not assessed here, because this system is not in that business and marking it down for that would be dishonest.
The record can be shown not to have changed.
R4.1 · absent. Does the system publish a scheme under which the record's past state can be verified?
No scheme is published, which is consistent with the product not claiming one. The md5 in the payload is a content hash a holder recomputes from the current text, so it detects nothing about alteration.
R4.2 · absent. Can an independent party run that verification without the vendor's cooperation, and without the operator's?
R4.3 · absent. Would alteration of a past entry be detectable after the fact?
Read of mem0/memory/main.py, mem0/memory/storage.py, mem0/configs/base.py and mem0/configs/prompts.py at 9a7924b, cloned from the public repository. Line numbers are that commit's.
Then it is wrong in the ordinary way readings are wrong, and every verdict cites a file and a line at a pinned commit precisely so that being wrong is cheap to demonstrate. The remedy is a pull request against the subject file, and it does not involve persuading anybody. Nobody applied for this and it is not a certification.
The standing register carries every system side by side, and the rubric is the twenty requirements in full, free to apply to anything, including to this assessment.