Can a provider's logs prove what your AI agent did? No. They only show what happened on that provider's own side of the workflow. They cannot show whether the agent's actions across your estate could be rebuilt in full, stopped in time, and undone. Before you approve wider autonomy for a business-critical AI workflow, work through the seven boardroom questions below and score the result.
Why provider logs are not enough
Every AI provider records its own side of an interaction: prompts sent, tokens used, tool calls it initiated, sometimes an output. That is real telemetry, and it is worth collecting. But it stops at the provider's boundary. It does not tell you what happened next: which system the agent's output touched, who approved the next step, whether a write action could be undone, or how long it took a human to notice and intervene if something went wrong.
A CISO signing off on wider AI autonomy is being asked to trust a chain of evidence across their own estate: identity, the target systems, the change itself, and the human oversight around it. Provider logs are one link in that chain, not the chain.
SenseOn's Decision Trace records every action SenseOn's own platform and agents take, not the internal workings of a third-party AI provider's runtime that SenseOn does not ingest, and not every AI agent operating elsewhere in your estate.
The seven boardroom questions
Pick one business-critical workflow an AI agent already touches, or one you are about to authorise. Answer each question yes or no, based on what you can produce today, not on what your architecture is designed to do eventually. Score one point per yes.
| # | Question | What a pass requires |
|---|---|---|
| 1 | Reconstruction. Can you reconstruct every action the agent took on this workflow, in order, with no gaps? | A complete, ordered action list, not a sample or a summary |
| 2 | Attribution. Can you show which decision was made by the AI and which by a human, at every step? | Each action tagged to its decision-maker, not inferred after the fact |
| 3 | Timing. Can you show when each action happened, precisely enough to build a timeline from trigger to resolution? | Timestamped actions, not an approximate window |
| 4 | Containment. Could you stop this agent's actions inside your target containment time, without shutting down the whole workflow around it? | A tested control that isolates the agent, with a measured time, not a theoretical kill switch |
| 5 | Reversal. Can you undo every write action the agent made, and prove the undo worked? | A rollback that has actually been exercised and verified, not assumed |
| 6 | Escalation. Can you show why the agent escalated to a human, or why it did not, against a stated policy threshold? | A documented threshold and a record of the decision against it |
| 7 | Immutability. Is the record of all of the above append-only and tamper-evident, or is it an application log that can be edited, rotated out, or lost? | An evidence store nobody, including an administrator, can quietly rewrite |
Scoring:
- 7/7: board-ready. You can approve wider autonomy for this workflow on the evidence you already hold.
- 5-6/7: workable, but close the specific gaps before you scale the workflow's autonomy further.
- 4 or fewer: do not approve wider autonomy for this workflow yet. You cannot currently prove what the agent did, only what you designed it to do.
A workflow can be well-designed and still fail these questions if nobody has exercised the rollback or measured the containment time.

The five board measures behind the questions
The seven questions are the audit; the board needs them expressed as ongoing measures, not a one-off pass or fail. Track these per workflow, on a rolling basis:
- Evidence completeness - the proportion of an agent's actions that appear in the reconstructable record, not a sample.
- Time to reconstruct - how long it takes a person to build a full, ordered timeline of what an agent did on a given case.
- Time to contain or revoke - how long from decision to actually stopping or revoking the agent's access.
- Reversal success rate - the proportion of write actions that were successfully undone when a rollback was tested or required.
- Human escalation rate - how often the workflow escalates to a human, and whether that rate matches the policy threshold you set, not one that drifted.
A workflow that scores 7/7 once and is never re-measured is not evidence of control. These five measures turn a single pass into a signal a board can track over time.
A worked example: what one evidence chain produced
Here is what an append-only evidence chain produces at scale in a Fortune 500 environment, where the record is complete rather than sampled:
| Stage | Volume |
|---|---|
| Events analysed | 33.4B per month |
| Alerts raised from those events | 36,000 |
| Cases requiring human judgement in 30 days | 173 |
| Human escalations per day | ~6 |
Behind that funnel, every AI and analyst action, including AI agent governance decisions, writes to SenseOn's Decision Trace: append-only, immutable, and built for chain-of-custody under NIS2, DORA, the EU AI Act and ISO 27001. That is what questions 1, 3 and 7 in the table above ask you to prove for your own workflow: not a sampling rate, but a record that exists and cannot be quietly edited.
That funnel is SenseOn's own workflow, measured on its own actions. Yours will look different in scale, and that is the point: the questions ask whether you can produce this kind of ordered, complete, reversible record, not whether you match someone else's numbers.
Where to start
Pick the one AI workflow closest to production autonomy today. Run the seven questions against it this week, before the next authorisation decision, not after. A score below 5 is not a reason to stop using the workflow; it is a reason to fix the specific gap the failed question names before you widen what the agent is allowed to do.
This article is part of The CISO's AI Accountability Playbook. For how to bring the five board measures into a board pack, see CISO board reporting.
Related reading: