Assessment, 2026-09-04 · 0.31.12 (e0a0e1e)
| Read at | e0a0e1e |
|---|---|
| Repository | github.com/letta-ai/letta-code |
| Licence | Apache-2.0 |
| Assessed as | stores, acts |
| Reaches | no level yet |
| Read by | Troy Brandon Clifford |
| Level | Meets | What the level asks |
|---|---|---|
| TR-1 Recorded | 3/5 | The record exists and is append-only. |
| TR-2 Explained | 0/4 | Every belief resolves to its evidence, and disagreements survive. |
| TR-3 Gated | 3/7 | Actions carry a verdict, and approvals carry a name. |
| TR-4 Verifiable | 1/3 | The record can be shown not to have changed. |
A count is requirements fully met out of those that apply. This system is assessed as stores, acts, and requirements outside that are not counted against it.
The memory here is a git repository, and that one decision answers most of TR-1 by itself: every memory write is a commit, so nothing is overwritten, every prior version is readable, and the history has an author. It is the only assessed system whose store is content-addressed by construction, which puts it closer to TR-4 than anything else in this census. Two things stop it going further and both are deliberate. Commit signing is not merely absent but explicitly disabled, in memory-git-signing.ts, because the harness-managed committer identities have no key, so a rewritten history is indistinguishable from an original one to anyone who did not record a SHA beforehand. And git's durability cuts the other way on erasure: deleting a memory is a commit, which records the deletion and keeps the deleted content in history, so honouring a real erasure request means rewriting the history that was the point of using git. The approval flow could not be settled from this repository and is marked accordingly rather than assumed.
Twenty requirements, each stated as a capability rather than a format, so a system that holds the information in its own shape counts as having it. An absent verdict cites where the assessor looked and did not find it, which is the difference between a measurement and an accusation.
The record exists and is append-only.
R1.1 · present. When the system stores a fact, does it durably record when that happened?
R1.2 · present. When a stored fact changes, is the previous version still readable?
R1.3 · present. Does the ordinary write path ever destroy what was previously recorded?
The strongest answer to this requirement in the census, and the one obtained most cheaply: using git means append-only was never a thing that had to be built.
R1.4 · partial. Are entries distinguishable by kind, or is everything one undifferentiated blob of text?
Files are distinguishable from one another. A claim and its source are not distinguishable within a file, because both are prose.
R1.5 · partial. When data must be destroyed for a legal reason, is the destruction itself recorded?
The trace is excellent and the other half fails in the way git always fails it: the deleted content stays in history, and on any mirror it was pushed to. Satisfying a real erasure request means rewriting the history, which destroys the record that made this requirement passable. This is the same bind mem0 is in, reached from the opposite direction.
Every belief resolves to its evidence, and disagreements survive.
R2.1 · absent. Can the source a stored fact came from be recovered from the store, by following a link rather than by guessing?
The commit history says when a memory appeared and which agent identity wrote it. That is closer to provenance than most systems here manage, and it is still not a link to the source: reading a memory gives no way to reach the conversation that produced it.
R2.3 · partial. When two stored facts about the same proposition disagree, do both survive?
Both survive as history because git keeps everything. Only one survives as memory: reading the current file gives the latest text, and finding what it replaced means reading the log.
R2.4 · absent. Is the disagreement itself queryable, or must a reader diff rows to notice it?
R2.5 · partial. When a conflict is resolved, does the record say who resolved it and by what method?
There is an actor on every change, which is more than most systems here record. The actor is the agent, and since nothing recognises a disagreement in the first place, there is no method recorded for how one was settled.
R2.2 is not assessed here, because this system is not in that business and marking it down for that would be dishonest.
Actions carry a verdict, and approvals carry a name.
R3.1 · present. Does a consequential action produce a durable entry whether or not it ran?
A proposed action that has not run is a durable record, not a run-scoped one, and the harness depends on that being true in order to resume.
R3.2 · present. Does an action's risk class come from somewhere the proposing model cannot write to?
Binary rather than graded, like every other framework assessed here, and set outside anything a plan can write, which is the substance of the requirement.
R3.3 · present. Are refusals recorded as faithfully as permissions?
The only system in this census where a refusal has its own record type rather than being narrated into a tool result.
R3.4 · partial. Does a refusal record why it was refused?
A reason has a defined place and survives into the record, which is better than CrewAI discarding it or LangGraph never modelling it. It is optional and it arrives as prose inside a tool return rather than as a field, so it cannot be relied on or read mechanically.
R3.5 · undetermined. Does an approval identify a person or a named role holder?
This is the one system in the census where the answer might be yes, and it cannot be settled from the source available. An acting-user identity exists, is server-validated, and is explicitly designed to name the human rather than the credential, which is precisely what this requirement asks for. Whether it lands on an approval response is not visible here. Recording that honestly is worth more than guessing in either direction, and the Letta maintainers can close it in one sentence.
R3.6 · undetermined. Does the approver's identity come from the authentication layer rather than from something the model can write?
The mechanism that would satisfy this requirement demonstrably exists and is validated in the right place. What is unverifiable from here is whether it is applied at the approval boundary.
R3.7 · undetermined. Is the acting agent prevented from approving its own action?
The record can be shown not to have changed.
R4.1 · partial. Does the system publish a scheme under which the record's past state can be verified?
The substrate provides a real integrity structure and the implementation neither publishes it as one nor takes the step that would give it force. The reason for disabling signing is sound: the harness-managed committer identities have no key, so a global commit.gpgsign=true would break memory initialisation outright. The consequence is still that nothing attests to who made a commit beyond a string in the author field.
R4.2 · present. Can an independent party run that verification without the vendor's cooperation, and without the operator's?
Whatever verification the format supports can be run entirely by the record holder against their own clone. That is the requirement, and using git means it is satisfied by every tool the reader already has.
R4.3 · partial. Would alteration of a past entry be detectable after the fact?
Closer to tamper-evidence than anything else assessed here, and short of it by one step. The push mirror is already most of an external anchor; recording the pushed head somewhere the harness cannot force-push would finish the job without changing the write path at all.
Read of src/agent/memory-git.ts, memory-git-signing.ts, memory-git-hooks.ts, acting-user.ts, check-approval.ts, approval-result-normalization.ts and src/tools/manager.ts at e0a0e1e, cloned from the public repository. Line numbers are that commit's. SCOPE: this repository is the harness. Agent messages, including approval requests and responses, are held by the Letta server, which is not in it. Requirements that turn on the shape of that record are marked undetermined rather than guessed, and say so.
Then it is wrong in the ordinary way readings are wrong, and every verdict cites a file and a line at a pinned commit precisely so that being wrong is cheap to demonstrate. The remedy is a pull request against the subject file, and it does not involve persuading anybody. Nobody applied for this and it is not a certification.
The standing register carries every system side by side, and the rubric is the twenty requirements in full, free to apply to anything, including to this assessment.