The evidence an AI policy asks for, and who can produce it

Policies covering autonomous agents are being written now, and the drafting has converged on a condition precedent: the insured must preserve the records that show what the agent did and who allowed it. The model clause quoted throughout this note is from Quanyan Zhu, Insurance of Agentic AI (arXiv:2606.05449), whose Table 2 makes it a condition precedent to indemnity that the insured "shall preserve materially relevant prompts, tool traces, access logs, model/version records, approval records, and rollback history, where reasonably available", on the drafting caution that "without logs, causation and allocation become nearly impossible". Its control schedule separately requires "documented human confirmation or equivalent approved control" for consequential decisions.

Four words carry more weight than their length suggests. Preservation is conditioned on the record being where reasonably available, and whether that caveat protects an insured or voids the clause for most of the market is an empirical question rather than a drafting one. If the software has nowhere to put the identity of an approver, the record was not available on the day the policy incepted, at any price. Read one way the caveat excuses exactly the deployers the condition precedent was written to bind. Read the other, the deployer is uninsured on a point they had no way to discover. The measurement below is what decides which.

That clause assumes the records exist. This note is about whether they do. Ten widely deployed agent systems were read against a published rubric in August and September 2026, and the finding an underwriter should care about is narrow and unambiguous.

Of the eight assessed systems that take or gate actions, one can name the person who approved an action. Six cannot, because nothing on the approval path identifies a person: the type that carries what a human decided has no field for which human decided it. In the eighth it could not be determined from outside the vendor.

This is not negligence by the deployers. It is that the software has nowhere to put the information. A deployer running one of those systems cannot satisfy the condition precedent at any price, and will not find out until a claim.

The requirement, restated as something checkable

A policy requirement is a sentence. An assessment needs a question with an artifact for an answer. These are the questions behind the clause. The three approval and three tool-trace questions apply only to the eight systems that take or gate actions; the last two apply to all ten, because anything that stores a fact can lose it.

What the policy asks forThe checkable question OfCanPartlyCannot
Approval records Does an approval identify a person or a named role holder? 810 6, and 1 undetermined
Approval records Does the approver's identity come from the authentication layer rather than from something the model wrote? 810 6, and 1 undetermined
Approval records Is the acting agent prevented from approving its own action? 810 6, and 1 undetermined
Tool traces Does a consequential action produce a durable entry whether or not it ran? 8530
Tool traces Are refusals recorded as faithfully as permissions? 8440
Tool traces Does a refusal record why it was refused? 8242
Rollback history When a stored fact changes, is the previous version still readable? 10532
Causation Can the source a stored fact came from be recovered by following a link rather than by inference? 10262

The requirement the clause does not contain, and should

Preserving a record is not the same as being able to read it. The clause obliges the insured to keep approval records; it says nothing about who can interpret them afterwards.

Of the ten systems assessed, eight produce records that cannot be verified without the vendor's cooperation.

At the moment that matters, the vendor may be insolvent, uncooperative, or the adverse party. A condition precedent satisfied only while a third company chooses to help is a condition precedent that fails exactly when it is invoked. This is the difference between a log and evidence, and it is the question a loss adjuster will ask second and an opposing expert will ask first.

It is also cheap to test, which is the point of putting it on a proposal form rather than discovering it during a claim.

What a decline cannot tell you

Added 9 September 2026.

Underwriters are already declining this risk, and the stated reason repays reading closely. Armilla, describing its head of AI policy Philip Dawson, published in January 2026 that applications for AI insurance “have been declined not because the models were inherently unsafe, but because governance and monitoring processes were too thin to support risk transfer”, and that insurable AI “requires evidence that systems are tested, monitored, and governed over time, not just certified at a single point”. Both are right.

But a decline on those grounds is a judgement about the applicant, and for six of the systems assessed here it cannot be one.

A deployer running one of those six can have a real approval gate, a written policy, a trained reviewer and a person who genuinely clicks approve, and still be unable to say afterwards who that person was. Their governance is not thin. Their software has nowhere to put the fact. Both cases produce the same submission: a policy document, and no record naming anybody.

Weak process and absent capability are indistinguishable from a submission, and only one of them is the applicant’s to fix.

That cuts both ways and both ways are expensive. A deployer with good governance is declined, or loaded, for an omission in software they bought, which is a mispricing they cannot correct by improving anything they control. And a deployer with thick documentation and weak practice is accepted, holding a policy whose evidence clause cannot be satisfied at any price on the day it incepts.

The fact that separates them is not in the submission. It is in the product the deployer bought, it is the same fact for every customer running that product, and it is established once by reading the software rather than repeatedly by interviewing applicants. Which is what the questions below are for.

Six questions for a proposal form

Each is answerable from an artifact rather than an assurance. A deployer who can answer them can satisfy the clause; one who cannot, cannot, whatever their policies say. Nothing here requires any particular software, and no answer is improved by naming a vendor.

  1. Show one approval from last month. Who approved it, and where did that name come from: a session, a token, an operator console?
  2. Show one action that was refused. If refusals are not on the record, the record cannot distinguish a system that declined from one that was never asked.
  3. Who set the risk class that decided the action needed approval, and can the model that proposed the action write to it?
  4. Can the agent approve its own action? Answer from the code path, not the policy document.
  5. Hand the record to somebody with none of your software. Can they confirm it has not changed since it was written?
  6. What was believed at the time, and on what evidence? A decision that looks wrong afterwards may have been right on what was known, and only the record can carry that difference.

Question five is the one that most often changes an answer, because almost every system passes questions one to four on paper and fails five in practice.

The method, and what it does not show

Twenty requirements, published, each stated as a capability rather than a format. A system that holds the information in its own shape counts as having it; translation into any particular format is a separate and smaller problem. Systems are assessed only against capabilities they claim: a vector store is not failing at approval gates, it is not an approval gate, and requirements outside a subject's declared scope resolve to not applicable.

What this is not. It is not a security review, a benchmark, or a ranking. It says nothing about whether a system is well built or fit for a purpose, and a system that scores badly here may be excellent at what it was made for. It measures one thing: whether a record exists that somebody outside the vendor could rely on afterwards.

Where it can be wrong. Every verdict was reached by reading published source and documentation at a named commit. A capability held in a component that cannot be read from outside is recorded as undetermined rather than guessed, which happened once. Versions move; each assessment carries the version and the date it was read, and the register is re-read on a published schedule so a stale finding fails the build rather than sitting there.

The assessments, the rubric, the evidence behind every verdict and the re-reading date are at the register. They are free, they name their sources, and they can be disputed with a commit reference.

If you are pricing this risk

The useful thing is not a rating. It is that the six questions above can go on a proposal form tomorrow, cost nothing to ask, and separate a deployer whose evidence will survive a claim from one whose will not. The rubric behind them is public so that a broker, a deployer or a competing assessor can apply it without asking anyone's permission, which is the only way a test of this kind stays worth anything.

If a question is wrong, or a verdict is out of date, that is worth more to me than agreement: the correction is public and dated like everything else. How to assess a system yourself is written down, and nothing in it needs my software.