The safest clinical model still needs a boundary it cannot cross.
Clinical judgement is not what governance disputes. What it governs is the moment an action leaves the reasoning layer and touches a record, a device, or a patient. This scenario shows that moment.
What you are already obliged to produce
These obligations predate the model. They do not soften because the actor is automated.
| Authority | What it obliges you to produce |
|---|---|
| HIPAA §164.312(b) | Audit controls: hardware, software and procedural mechanisms that record and examine activity in systems holding protected health information. |
| HIPAA §164.312(a) | Access control — technical policies restricting PHI to those with a right to it, enforced rather than documented. |
| 21 CFR Part 11 | Electronic records and signatures: attributable, legible, contemporaneous, original and accurate, with an audit trail that cannot be silently altered. |
| IRB protocol | Research use constrained to the approved protocol, with any deviation recorded and reviewable. |
Why the usual answer does not close
An audit control that a model can talk its way past is a description of a control, not a control.
Prompt-level restrictions live inside the same probabilistic system they are meant to restrain, and are re-interpreted every call rather than enforced once.
PHI custody becomes the blocker: any layer that retains payload to evaluate it enlarges the surface the privacy office has to defend.
A deviation found in a quarterly review is a deviation that already reached the patient record. The control has to sit before the effect, not after it.
One action at the boundary
Evaluation happens on a compact candidate state. Payload is read, scored, and not retained — the ledger holds proofs and dispositions, not the record.
Candidate action
patient.records.export · 1,240 rows · endpoint outside deployment boundary
The tool name was not on any denylist — this is the exact class our own gate failed on before the Phi-Invariant remediation, published in full on the security page. Under default-deny, the question is no longer whether the tool is forbidden but whether this effect, to this target, under this authority, is explicitly allowed. It was not.
The same measurements, whatever the sector
These are figures from the internal technical evidence report, not projections modelled for this scenario. They describe one pipeline, so they do not change when the mandate does.
0 / 18
Attack episodes passed
Stage 0 SHADOW evaluation of action traces authored by real generative planners. None of nine held-out benign episodes was blocked.
0 / 10
Prohibited cases passed, dual reference
Controlled ablation on the same held-out set. A single-reference gate passed four of ten under identical calibration.
0.47 ms
Mean decision latency
Across 5,000 measured decisions; 1.15 ms at p99. The gate itself completes in 258 nanoseconds.
1,000 / 1,000
Identical digests on replay
Same input, pinned environment. Ten thousand ledger records re-verified in 0.195 seconds.
What this scenario is not
Read this before you quote it
- An illustrative scenario, not a customer engagement. No client is named because none is being described.
- No deployment in this sector is claimed, and no regulator has reviewed or endorsed this material.
- The figures are prototype measurements on one commodity workstation, not production or distributed results.
- Nothing here has been independently reproduced by a third party.
- The authorities cited describe the obligation you carry — not a determination that we satisfy them.
- Operational accuracy on a validated sector corpus remains outstanding; a pilot requires your own reference data.
Check the numbers before you trust the scenario
Every figure above is drawn from the technical evidence page, where the same measurements appear with their tail distribution and their stated limits.