What mem0 records, and what it does not

Read at9a7924b
Repositorygithub.com/mem0ai/mem0
LicenceApache-2.0
Assessed asstores, derives
Reachesno level yet
Read byTroy Brandon Clifford
LevelMeetsWhat the level asks
TR-1 Recorded3/5The record exists and is append-only.
TR-2 Explained0/5Every belief resolves to its evidence, and disagreements survive.
TR-3 Gatedn/aActions carry a verdict, and approvals carry a name.
TR-4 Verifiable0/3The record can be shown not to have changed.

A count is requirements fully met out of those that apply. This system is assessed as stores, derives, and requirements outside that are not counted against it.

mem0 is a memory layer built for retrieval quality and low integration cost, and its update phase consolidates on purpose: when new information contradicts an existing memory, the prompt at configs/prompts.py:264 instructs the model to delete the old one. Most of the gaps here follow directly from that goal rather than being oversights against it, and mem0 has never claimed to be an audit record. What the assessment says is narrow: a deployment that later needs to show what the agent believed and why will not find those facts here, because the design traded them away for something else it wanted more. Two findings are worth the maintainers' attention independently of any of that, and both are noted on the requirements they sit under: R1.5, where a deleted memory's text is retained in the local history database, and R2.2, where nothing distinguishes an LLM-extracted fact from a message stored verbatim.

Every verdict, and what it rests on

Twenty requirements, each stated as a capability rather than a format, so a system that holds the information in its own shape counts as having it. An absent verdict cites where the assessor looked and did not find it, which is the difference between a measurement and an accusation.

TR-1 Recorded

The record exists and is append-only.

R1.1 · present. When the system stores a fact, does it durably record when that happened?

R1.2 · present. When a stored fact changes, is the previous version still readable?

  • source mem0/memory/storage.py:108-119
    the history table holds memory_id, old_memory, new_memory, event, created_at, updated_at, is_deleted, actor_id and role
  • api Memory.history(memory_id), mem0/memory/main.py:1946
    public method returning the change history for one memory
  • source mem0/configs/base.py:42-45
    history_db_path defaults to ~/.mem0/history.db, so the history persists rather than living only in the process

R1.3 · partial. Does the ordinary write path ever destroy what was previously recorded?

The routine write path overwrites. A correction is an edit to the existing record rather than a new entry, and the previous state lives in a separate history database instead of in the store being queried. This is the requirement that keeps mem0 below TR-1, and it is a deliberate design choice rather than a defect: consolidating in place is what makes the retrieval surface small.

R1.4 · present. Are entries distinguishable by kind, or is everything one undifferentiated blob of text?

The kinds are separate. Whether they are linked is a different question, assessed at R2.1.

  • source mem0/memory/storage.py:17-18
    history and messages are separate tables created alongside the vector store, so a memory, the message it came from and the record of a change are three distinct record kinds

R1.5 · partial. When data must be destroyed for a legal reason, is the destruction itself recorded?

The deletion is recorded, which is half of what this asks. The other half is not met, and in a direction worth flagging to the maintainers: _delete_memory writes the deleted memory's own text into history.old_memory, so a deletion carried out to satisfy a privacy request leaves that content sitting in the local history database. Anyone relying on delete_all for erasure should know the text survives there. This is the one verdict in the census demonstrated by running the assessed system rather than by reading it, because it is a claim that something is written rather than that something is missing, and it concerns data somebody asked to have deleted.

  • source mem0/memory/main.py:1890 delete_all
    the per-user deletion path, which calls _delete_memory for each memory so every deletion is written to history
  • source mem0/memory/main.py:2100 _delete_memory
    add_history is called with prev_value as old_memory and is_deleted=1, so the destruction is recorded
  • source mem0/memory/storage.py:326 reset
    DROP TABLE IF EXISTS history; a full reset destroys the record of what was destroyed
  • run benchmarks/census/verify/mem0_delete_retains_text.py --repo ./mem0
    executed against this commit: add_history for a DELETE event writes the memory's own text into history.old_memory and it is readable from the database afterwards. Runs mem0's SQLiteManager directly, with no network, no LLM and no API key. Output is reproducible by checking out 9a7924be and running the script

TR-2 Explained

Every belief resolves to its evidence, and disagreements survive.

R2.1 · partial. Can the source a stored fact came from be recovered from the store, by following a link rather than by guessing?

The raw source is kept but the fact does not point at it. Recovering which message produced a given memory means matching on actor, role and approximate time, which is guessing rather than following a link.

  • source mem0/memory/storage.py:19 _create_messages_table
    raw messages are retained, and get_last_messages reads them back
  • source mem0/memory/main.py:1961 _create_memory
    the payload written for a memory is data, hash, created_at, updated_at and text_lemmatized, plus role and actor_id; there is no message or source identifier among them
  • searched grep -n 'message_id\|source_id' mem0/memory/main.py
    no match; the only 'source' hits are a local variable in merge_filters

R2.2 · absent. Can a fact the system inferred be told apart from one it was told?

Both kinds of memory sit in one store looking identical. A consumer reading a mem0 memory cannot tell whether the model wrote it or a person did, which matters most for exactly the memories a model got wrong.

  • source mem0/memory/main.py:879-914
    with infer=False the raw message content is stored verbatim as a memory; with infer=True the LLM-extracted fact is stored. Both go through _create_memory and produce the same payload shape
  • searched grep -rn 'is_inferred\|inferred\|"infer"' mem0/memory/main.py mem0/memory/storage.py
    no match; the infer argument controls the code path but is never recorded on the resulting memory

R2.3 · partial. When two stored facts about the same proposition disagree, do both survive?

Only one side survives as a belief. The other survives as a text column in a history row, which is an audit trail rather than a retained memory: a later query about that subject returns the winner and gives no sign there was ever a loser.

  • docs mem0/configs/prompts.py:264
    "Delete: If the retrieved facts contain information that contradicts the information present in the memory, then you have to delete it"
  • source mem0/memory/main.py:2100 _delete_memory
    the contradicted memory is removed from the vector store; its text is written to history.old_memory with is_deleted=1

R2.4 · absent. Is the disagreement itself queryable, or must a reader diff rows to notice it?

A disagreement leaves an UPDATE or DELETE event indistinguishable from any other. Noticing that two sources ever disagreed means reading the history and inferring it.

  • searched grep -rniE 'class Conflict|conflict_id|def conflicts|contradiction' mem0/ --include=*.py
    the only match is descriptive text inside an LLM prompt at configs/prompts.py:699; there is no conflict record, id or listing anywhere in the code
  • source mem0/memory/storage.py:108-119
    the history schema has no column that would mark a change as having been caused by a disagreement

R2.5 · partial. When a conflict is resolved, does the record say who resolved it and by what method?

What changed is recorded. Who decided and how is not: the decision is taken by an LLM against the prompt at configs/prompts.py, and neither the model, the prompt version nor the reason is written down anywhere.

TR-3 Gated

Actions carry a verdict, and approvals carry a name.

R3.1, R3.2, R3.3, R3.4, R3.5, R3.6, R3.7 are not assessed here, because this system is not in that business and marking it down for that would be dishonest.

TR-4 Verifiable

The record can be shown not to have changed.

R4.1 · absent. Does the system publish a scheme under which the record's past state can be verified?

No scheme is published, which is consistent with the product not claiming one. The md5 in the payload is a content hash a holder recomputes from the current text, so it detects nothing about alteration.

  • searched grep -rniE 'hash_chain|merkle|signature|tamper|checksum|anchor' mem0/ --include=*.py
    no match relating to record integrity. The hashes present are an md5 of memory text used for deduplication (main.py:1961), a sha1 used to bucket telemetry notices (memory/notices.py:94) and a sha256 used to anonymise a telemetry id (memory/setup.py:97)
  • docs mem0 README.md
    no integrity, verification or tamper-evidence claim is made

R4.2 · absent. Can an independent party run that verification without the vendor's cooperation, and without the operator's?

  • searched grep -rniE 'verify|verification' mem0/memory/ --include=*.py
    no verification surface exists to be run by anybody, so the question of who can run it does not arise

R4.3 · absent. Would alteration of a past entry be detectable after the fact?

  • searched grep -rniE 'hash_chain|merkle|signature|tamper|checksum|anchor' mem0/ --include=*.py
    no match; nothing links one record to the next, so editing a row in the history database or the vector store leaves no evidence
  • source mem0/memory/storage.py:108-119
    history rows carry no digest of prior rows and no ordering guarantee beyond their timestamps

How this was read

Read of mem0/memory/main.py, mem0/memory/storage.py, mem0/configs/base.py and mem0/configs/prompts.py at 9a7924b, cloned from the public repository. Line numbers are that commit's.

Found an error? Challenge a finding

Then it is wrong in the ordinary way readings are wrong, and every verdict cites a file and a line at a pinned commit precisely so that being wrong is cheap to demonstrate. The remedy is a pull request against the subject file, and it does not involve persuading anybody. Nobody applied for this and it is not a certification.

The standing register carries every system side by side, and the rubric is the twenty requirements in full, free to apply to anything, including to this assessment.