Your model is probabilistic. Your attestation is not.
A CFO signs under §302 that controls were effective. That signature does not become conditional because a model was involved. This scenario shows where a deterministic layer sits so the attestation stays defensible.
What you are already obliged to produce
None of these are new because of AI. What is new is that a probabilistic component now sits inside the path they govern.
| Authority | What it obliges you to produce |
|---|---|
| SOX §302 | Officer certification that disclosure controls were effective in the period. An effectiveness claim needs evidence that survives re-examination, not a model's recollection. |
| MiFID II Art. 16 | Records of orders and client communications, retained and retrievable in a form an authority can inspect. |
| DORA Art. 11 | ICT incident response and reporting: what happened, under which controls, in what order, reconstructable after the fact. |
| SR 11-7 | Model risk management — validation, ongoing monitoring, and a clear account of what the model was permitted to decide. |
Why the usual answer does not close
A guardrail that returns a different answer on a different day cannot support a certification that must hold for a period.
A probabilistic guardrail re-scored months later may not reproduce the decision it made at the time. An examiner asking why an action was permitted receives an approximation.
Security review blocks expansion because tool permissions are broad and the blast radius of a wrong action is settlement-sized, not ticket-sized.
Evidence lives in logs written by the same system whose behaviour is in question, with no binding between the verdict and the policy version that produced it.
One action at the boundary
The engine does not review the model's reasoning. It evaluates the action the model is about to take, against frozen thresholds, before the action reaches the world.
Candidate action
wire.transfer.execute · $2.4M · digest previously consumed
Schema and proof anchor validated. Replay uniqueness did not: the WORM ledger already held this payload digest. Ordered validation means a failure here blocks by construction rather than by a rule someone remembered to write — and the disposition is bound to the exact policy version in force, so the block is replayable years later.
The same measurements, whatever the sector
These are figures from the internal technical evidence report, not projections modelled for this scenario. They describe one pipeline, so they do not change when the mandate does.
0 / 18
Attack episodes passed
Stage 0 SHADOW evaluation of action traces authored by real generative planners. None of nine held-out benign episodes was blocked.
0 / 10
Prohibited cases passed, dual reference
Controlled ablation on the same held-out set. A single-reference gate passed four of ten under identical calibration.
0.47 ms
Mean decision latency
Across 5,000 measured decisions; 1.15 ms at p99. The gate itself completes in 258 nanoseconds.
1,000 / 1,000
Identical digests on replay
Same input, pinned environment. Ten thousand ledger records re-verified in 0.195 seconds.
What this scenario is not
Read this before you quote it
- An illustrative scenario, not a customer engagement. No client is named because none is being described.
- No deployment in this sector is claimed, and no regulator has reviewed or endorsed this material.
- The figures are prototype measurements on one commodity workstation, not production or distributed results.
- Nothing here has been independently reproduced by a third party.
- The authorities cited describe the obligation you carry — not a determination that we satisfy them.
- Operational accuracy on a validated sector corpus remains outstanding; a pilot requires your own reference data.
Check the numbers before you trust the scenario
Every figure above is drawn from the technical evidence page, where the same measurements appear with their tail distribution and their stated limits.