Deterministic policy conformance against a frozen Government of Canada AIA snapshot
A fixed-N, department-disjoint, pre-registered conformance evaluation. Rule logic, action templates, metrics, and pass thresholds were hash-locked before any held-out case was opened.
- Pre-registered · hash-locked
- Department-disjoint held-out
- Internal · not independently reviewed
Result label
INTERNAL_NUMERICAL_PASS_ON_ACHIEVED_SET
PLANNED_PRIMARY_CASE_N_NOT_MET
NOT_INDEPENDENTLY_REVIEWED_OR_VALIDATED
At a glance
16
Held-out cases scored
Planned 17 — not met
208
Action-case decisions
Primary explicit report
3
Held-out organisations
Disjoint from calibration
0
Unsafe PASS observed
Wilson 95% upper 0.0269
What this study is — and is not
What it measures
- Whether the engine deterministically enforces a set of team-compiled policy rules derived from a frozen snapshot of the Directive on Automated Decision-Making and the AIA tool.
- Behaviour on hypothetical actions grounded in real, publicly published government system assessments.
- Conformance on held-out cases whose membership and size were fixed by pre-registered rules before the data was opened.
- Whether every decision certificate cites the correct rule, and whether replaying the full set reproduces byte-identical output.
What it does not measure
- Accuracy against a government decision label. No such ground truth exists — published AIAs are assessments, not action-decision logs.
- Legal compliance with Canadian law. This study makes no compliance determination.
- Any form of government review, validation, approval, or affiliation.
- Behaviour inside government infrastructure. Nothing was tested on or with any government system.
Honest external-vs-internal labelling
Two of the three layers come from outside EthicVault. One does not — and we label it plainly rather than presenting the whole stack as external ground truth.
| Layer | Source | Truly external? |
|---|---|---|
| Policy — obligations, impact levels, mitigations | Government of Canada | External |
| Case context — system descriptions, published answers, impact level | Published government AIAs | External |
| Action-level PASS / HOLD / BLOCK labels | EthicVault team-compiled rules | Our interpretation |
Because the third layer is ours, every rule carries a verbatim policy quote so an independent reviewer can check the reading against the source text.
Two-phase discipline
The point of the design is that nothing which could bias the result was decided after seeing the data.
- Phase 1
Freeze
Cutoff date, eligibility rules, department split, action-template library, policy rule logic, and acceptance bounds are fixed and hashed. SHA-256 of every Phase-1 file is written to a lock file that refuses Phase 2 if anything moved.
- Phase 2
Bind
The frozen policy snapshot and eligible assessments are pulled. Every compiled rule is bound to a verbatim quote from the published policy so its reading can be independently checked.
- Phase 3
Execute
Held-out organisations are opened once. Resources are fetched and parsed mechanically — no semantic inference from prose, title, or department. Zero manual corrections were applied.
- Phase 4
Score
Results are scored against the pre-registered bounds only. The lock is verified again after extraction and scoring; neither the mapping nor any locked program changed after the content opened.
The non-forcing rule
Only explicit requirements — clear, checkable obligations decidable from structured fields — count toward the safety metrics. Requirements needing judgement are routed to REVIEW_INTERPRETATION_REQUIRED and are never forced into a PASS or BLOCK to make a number look better. This is enforced by schema, not by convention.
Results
Primary explicit report on the achieved held-out set: N=16 cases, 208 action-case decisions. Every rate is stated with its numerator, denominator, and Wilson 95% interval.
| Metric | Result | Wilson 95% |
|---|---|---|
Unsafe PASS Explicit prohibition violated, yet disposition PASS | 0 / 139 | [0.0000, 0.0269] |
False HOLD / BLOCK Compliant control not passed | 0 / 16 | [0.0000, 0.1936] |
Policy citation accuracy Certificate cites the correct rule | 171 / 171 | [0.9780, 1.0000] |
Missing-evidence detection Cases lacking required input routed away from PASS | 16 / 16 | [0.8064, 1.0000] |
Ambiguity forced to a decision Must be zero | 0 / 16 | Hard gate — met |
Deterministic replay Re-run reproduces byte-identical certificates | Bit-exact | Hard gate — met |
Full sensitivity report (N=17)
Including the previously format-inspected case A08: unsafe PASS 0/148 (Wilson 95% upper 0.0253); false HOLD/BLOCK 0/17 (upper 0.1843); citation accuracy 181/181; missing-evidence detection 16/16; forced ambiguity dispositions 0/17. The primary and full-sensitivity replays were bit-exact.
Secondary interpretive report
Under the team-interpretation reading, 16 of 16 applicable decisions routed to REVIEW_INTERPRETATION_REQUIRED — none were resolved automatically.
Limitations
These are published because a study that only reports what worked is marketing, not evidence.
The planned sample size was not met
The pre-registered primary denominator was 17 unexposed cases. The achieved primary set is 16. The correct label is a numerical pass on the achieved set — not a pass on the planned study.
One case was mechanically stopped
Case A21 declared AIA version v0.10.0 but published an answer code absent from the frozen choice set for that question. It was assigned reason_code UNKNOWN_ANSWER_CODE and status ABORT_CASE_EXTRACTION. The code was not added to the mapping and the case was not scored — adding it after the fact would have broken the freeze.
One case was excluded from the confirmatory set
Case A08 had its file format inspected during a single disclosed feasibility check before the lock. It is excluded from the primary confirmatory denominator and reported separately in the sensitivity analysis.
Implementation answers are self-reported
Published mitigation answers are treated as OFFICIAL_SELF_REPORTED_IMPLEMENTATION, not independently verified controls. Portal publication and the presence of a completed record are kept strictly separate from publication-before-production, completion-before-production, and approval claims.
Some intervals are wide
At N=16, a zero-failure result on the false HOLD/BLOCK axis still carries a Wilson 95% upper bound of 0.1936. Small denominators are reported, not hidden behind a percentage.
This is an internal run
The study owner compiled the rules and executed the run. That conflict is disclosed in the release record. The result has not been independently reviewed or validated.
Verify it yourself
The pre-registration, the compiled policy pack, and the extraction lock are hash-pinned. If any of them had changed after the data opened, these digests would not match.
Pre-registration root
12b01f62820af6ca6343bca20575baa935ad57a45eb0d2d891cc831c3efb5aec
Policy pack
93199ca8c14a8230d2e816605ea4b60165d0c2d2a2504df291217b6fc742e99c
Final extraction lock root
619359da4642893754c9c69378760ce62407ef18d2e1b63b1272143d3a215c2b
Files under the extraction lock 69
Included verifier python verify_package.py
- 1Read the executive evidence report and the held-out execution report.
- 2Run the bundled verifier against the manifest.
- 3Recompute the SHA-256 of every listed artefact and compare with the manifest.
- 4Re-run the scoring pipeline and confirm the certificates are byte-identical.
The public assessment binaries are not republished in the package. Their selected URLs, final URLs, sizes, and SHA-256 hashes are recorded in the manifests so any reviewer can retrieve and re-hash the originals themselves.
Claim discipline
These boundaries were written into the study before it ran. They bind our sales material, not just this page.
Permitted
“EthicVault was evaluated against a frozen snapshot of the Government of Canada AIA policy and the publicly released system assessments, using team-compiled decision rules and department-disjoint held-out cases.”
Prohibited
- Government of Canada validated EthicVault
- Any department approved the engine
- EthicVault is compliant with Canadian law
- Tested inside Canadian government infrastructure
- Independently reviewed or validated
- Held-out set never opened
Sources and licence
Directive on Automated Decision-Making
Treasury Board of Canada Secretariat
Primary obligations — impact levels, mitigation requirements
Algorithmic Impact Assessment tool
Government of Canada
Risk and mitigation questions; impact-level scoring
AIA tool open-source code
Government of Canada
Scoring logic reference, pinned to a version tag
Published Algorithmic Impact Assessments
Open Government Portal
Case universe — retrieved after the lock, hash-pinned on retrieval
Snapshot cutoff: 13 July 2026, 00:00 UTC
Licence obligations
Reused under the Open Government Licence – Canada, with attribution preserved and no implication of endorsement or affiliation. No government wordmark, crest, or symbol is used anywhere on this site.
This page is published in several languages. In case of any discrepancy, the English text governs.
Request the evidence package
The full package is prepared for external technical and methodological review: the pre-registration and its amendments, the bound policy pack with verbatim quotes, the held-out execution report, the complete SHA-256 manifest, and the reproducibility scripts. Sent to verified work addresses.