Assessment, 2026-09-04 · autogen-core 0.7.5 (027ecf0)
| Read at | 027ecf0 |
|---|---|
| Repository | github.com/microsoft/autogen |
| Licence | MIT (code, LICENSE-CODE); CC BY 4.0 (documentation, LICENSE) |
| Assessed as | stores, acts |
| Reaches | no level yet |
| Read by | Troy Brandon Clifford |
| Level | Meets | What the level asks |
|---|---|---|
| TR-1 Recorded | 1/5 | The record exists and is append-only. |
| TR-2 Explained | 2/4 | Every belief resolves to its evidence, and disagreements survive. |
| TR-3 Gated | 1/7 | Actions carry a verdict, and approvals carry a name. |
| TR-4 Verifiable | 0/3 | The record can be shown not to have changed. |
A count is requirements fully met out of those that apply. This system is assessed as stores, acts, and requirements outside that are not counted against it.
AutoGen is assessed as storing and acting but not deriving: its memory holds what a caller puts in it, and any inference is done by the user's agents rather than by the framework. It has the best-modelled approval request in this census. ApprovalRequest carries the code and the full LLM context, and ApprovalResponse requires a reason rather than accepting a bare boolean, which is more than the other frameworks here ask for. Two things stop that translating into a record. The reason reaches the transcript interpolated into an output string on a result carrying exit_code=1, which is the same shape an execution error takes, so a refusal and a crash are indistinguishable to anything reading afterwards. And ApprovalResponse has no approver, so the fact that a person decided is not recorded even though the fact that a decision happened is. On the memory side, the built-in ListMemory is an append-only in-process list whose contents carry no timestamp at all.
Twenty requirements, each stated as a capability rather than a format, so a system that holds the information in its own shape counts as having it. An absent verdict cites where the assessor looked and did not find it, which is the difference between a measurement and an accusation.
The record exists and is append-only.
R1.1 · absent. When the system stores a fact, does it durably record when that happened?
A caller can put a timestamp in metadata, but then the caller is recording it, not the system, and nothing obliges or checks it. Reading a MemoryContent back gives no answer to when it was stored.
R1.2 · partial. When a stored fact changes, is the previous version still readable?
Nothing is overwritten, because nothing can be. What keeps this from a pass is that readable later is bounded by the life of the process: the built-in implementation is a list, and preserving it across a restart is the embedding application's problem.
R1.3 · present. Does the ordinary write path ever destroy what was previously recorded?
The ordinary write path is append-only by omission rather than by design: there is no update operation to misuse. That is still the property the requirement asks for.
R1.4 · partial. Are entries distinguishable by kind, or is everything one undifferentiated blob of text?
The type discriminates the format of the payload, not the kind of record. A claim, the document it came from and a note about an action are all one MemoryContent differing only in what a caller happened to put in metadata.
R1.5 · absent. When data must be destroyed for a legal reason, is the destruction itself recorded?
There is also no way to remove one subject's items short of clearing everything, which makes an erasure request an all-or-nothing operation on the whole memory.
Every belief resolves to its evidence, and disagreements survive.
R2.1 · absent. Can the source a stored fact came from be recovered from the store, by following a link rather than by guessing?
R2.3 · present. When two stored facts about the same proposition disagree, do both survive?
Both survive, which is what this asks. They survive because nothing in the framework is capable of noticing they disagree, which is what the next requirement is for.
R2.4 · absent. Is the disagreement itself queryable, or must a reader diff rows to notice it?
R2.5 · present. When a conflict is resolved, does the record say who resolved it and by what method?
A pass under the specification's own wording, which records resolution only if it happens. This is the conservative outcome, reached because the framework declines to do the thing rather than because it does it carefully.
R2.2 is not assessed here, because this system is not in that business and marking it down for that would be dishonest.
Actions carry a verdict, and approvals carry a name.
R3.1 · partial. Does a consequential action produce a durable entry whether or not it ran?
The attempt is recorded and the record is well populated. Whether it outlives the run is not the framework's decision.
R3.2 · present. Does an action's risk class come from somewhere the proposing model cannot write to?
Scoped to code execution rather than to tool calls in general, which is narrower than the requirement imagines but is genuinely outside model control within that scope.
R3.3 · partial. Are refusals recorded as faithfully as permissions?
The refusal is recorded in the same channel and with the same standing as a permission, which satisfies the letter. Structurally it is indistinguishable from a crash: anyone counting failed executions afterwards will count refusals among them.
R3.4 · partial. Does a refusal record why it was refused?
Requiring the reason at the point of decision is better than every other framework assessed here, two of which discard it or never model it. What is lost is the structure: a reason parsed back out of a formatted sentence is not a machine-readable field.
R3.5 · absent. Does an approval identify a person or a named role holder?
The documented example function reads a console prompt and returns ApprovalResponse(approved=True, reason='Approved by user'). The word user appears in the reason string, which is a sentence rather than an identity.
R3.6 · absent. Does the approver's identity come from the authentication layer rather than from something the model can write?
R3.7 · absent. Is the acting agent prevented from approving its own action?
The record can be shown not to have changed.
R4.1 · absent. Does the system publish a scheme under which the record's past state can be verified?
R4.2 · absent. Can an independent party run that verification without the vendor's cooperation, and without the operator's?
R4.3 · absent. Would alteration of a past entry be detectable after the fact?
Read of python/packages/autogen-core/src/autogen_core/memory/_base_memory.py and _list_memory.py, and python/packages/autogen-agentchat/src/autogen_agentchat/agents/_code_executor_agent.py at 027ecf0, cloned from the public repository. Line numbers are that commit's.
Then it is wrong in the ordinary way readings are wrong, and every verdict cites a file and a line at a pinned commit precisely so that being wrong is cheap to demonstrate. The remedy is a pull request against the subject file, and it does not involve persuading anybody. Nobody applied for this and it is not a certification.
The standing register carries every system side by side, and the rubric is the twenty requirements in full, free to apply to anything, including to this assessment.