What your stack can show on 1 January, and what it cannot

From 1 January 2027, a deployer of automated decision-making technology in Colorado must give a consumer an opportunity for meaningful human review of an adverse decision, and must keep records of compliance for three years. That is section 6-1-1703 of Senate Bill 26-189. A right to human review is older and wider than Colorado: the United Kingdom has required a controller to enable one since 5 February 2026 and requires nothing kept. What is new here is the retention.

Every advisory you will read between now and then tells you what the law requires. This one tells you something else, and as far as we can establish nobody else is measuring it: whether the software you already bought can produce the record.

The measurement

In August and September 2026, ten widely deployed agent and agent-memory systems were read at pinned commits against twenty record-keeping requirements, with every verdict published next to the file and line it rests on. Eight of the ten take or gate consequential actions.

1 of 8can identify the person who approved an action
6 of 8cannot identify anyone: the structure carrying what a human decided has no field for which human decided it
1 of 8could not be determined from outside the vendor

The one that can is the reference implementation maintained by the author of this page, which is disclosed here rather than left to be found. The finding is not that these are bad systems. None of them was built to answer this, and until 2026 nobody asked them to.

If you run one of these, this is what it does today

Six agent frameworks were read again on the narrower question of whether a record can relate what a reviewer was shown, what they approved, and what actually executed.

FrameworkWhat was shownApproved or modifiedBound to the call
LangGraph retained in the checkpoint typed, so an edit is distinguishable partial: the id names the graph position, not the arguments
OpenAI Agents SDK recoverable, not the rendering no modify path exists, so the two cannot be confused bound to the call id
Pydantic AI recoverable, not the rendering absent: the history states arguments that did not execute bound to the call id
Haystack the decision keeps the final parameters, not the proposed ones the vocabulary exists and the persisted object keeps one bool optional call id
CrewAI nothing on the approval path persists it a hook mutates the arguments in place and returns one bool no approval artifact to bind
AutoGen no approval pause in core, so the question does not arise as above as above

Every verdict has a locator and a version, and several were reproduced by running the code rather than reasoned about. They are at the approval binding reading, free to reuse and free to dispute.

What changes it, and it is smaller than you think

Adapters exist for five of these frameworks. They are MIT licensed, published on PyPI, and they do one thing: make the approval carry the identity your authentication layer already knows, and write a record that anybody can check.

pip install testimony-langgraph

Measured rather than claimed. A record produced through any of the five adapters, read against the eleven Colorado obligations this project has published a reading of:

6obligations the record could evidence
1it could not: section 6-1-1701(15)(c), that the reviewer does not default to the system output. No single record shows this. A population of them does.
4are not facts about a record at all, and are set aside rather than reported as failures

Those four are worth being clear about, because an advisory that counts them against you is selling something. Designation, training and authority live in your HR and identity systems, and the enumerated records in 6-1-1703 are about the software's version identifiers and changelogs. No record format produces any of them and none should claim to.

What this does not do

It does not make you compliant, and nobody selling you software can. Whether an obligation is met is a judgement somebody qualified makes on evidence. This gathers the evidence, and it is honest about the parts it cannot reach.

It is not legal advice. It is a reading of a public act and a measurement of public software, both published with their dates so you can check them or pay somebody to disagree with them.

Doing it yourself, for nothing

Everything above is free and stays free. The two-file route takes an afternoon. You can check a record in your browser, which uploads nothing, and read your own logs the same way. The rubric is Creative Commons and you may run it against your own suppliers and never tell us.

Or have it read for you

A Colorado readiness reading: your records, or the stack you run, assessed against the published reading of SB 26-189, item by item. You get an evidence pack containing the reading, the record it rests on, a digest over it, and the commands to recompute every number in it using a validator published elsewhere. Your auditor can reproduce every figure without trusting us or you.

Colorado is rarely the only text that reaches a deployer running agents. The same reading now checks your records against the other nineteen instruments in this library's private copy in a single pass, grouped by whether each currently binds, so you see in one document which texts you already answer, which you partly answer, and which ask nothing of a record at all. If only one applies to you, the reading can be scoped to that one instrument alone instead: an operator who only needs South Korea's Framework Act, say, gets an answer sized to that Act, not to twenty.

It contains no telemetry: the reading records which obligations were evidenced and by which field name, never the field's contents, so the pack can leave your infrastructure when the data it was derived from cannot.

Colorado's rules are not final. When the Attorney General's revised draft lands, the reading behind this page is re-read against it, and what changed between the two versions is produced from the two committed texts rather than summarised from memory: which obligations were added, which were removed, and whether the rule now binds when it did not before. That is the difference between a subscription and a download, made checkable rather than asserted.

The first three are free, and that is a trade rather than a discount: these instruments have never run against a real estate, your data will find what a fixture cannot, and a reference name is worth more than a first year of fees. It is said plainly here so nobody has to guess.

Ask: troy@machinetestimony.com. Tell us which frameworks you run and whether you already keep records. If the answer is that you do not need this, we will say so.

Sources, and where to go next

The act, as signed, is at leg.colorado.gov/bills/sb26-189, and the rulemaking is at coag.gov/ai. The reading of the act is at /colorado/, the dated census this page's measurement comes from is at /census/2026-09/, the live per-system verdicts at /register/, and the conformance level the records reach at /tr-3/.

To find out what your own stack's records establish rather than read about somebody else's, Review reads the logs you already have, and Check verifies a Testimony Record if you already produce one. Both run in the browser and upload nothing.