Defending AI-assisted decisions to a regulator
The question that decides these matters is rarely "did you use AI?" It is "how did you reach this decision, and can you show me?"
That is a records question. And it is answered — or lost — long before anyone asks it.
What regulators actually evaluate
Regulators, tribunals, and auditors do not assess your model architecture. They assess whether the decision can be reconstructed: what information it rested on, how that information was verified, what the automated system contributed versus the human, and whether the affected party was told.
The reason is straightforward. Administrative law has always required decision-makers to give reasons and to base decisions on evidence properly before them. Automation does not create a new standard; it creates a new way to fail the existing one — by producing an output nobody can trace back to the evidence that generated it.
What Canada requires today
For federal institutions and the vendors serving them, the Treasury Board's Directive on Automated Decision-Making is the operative instrument. It requires institutions to assess and mitigate the impacts of automated decision systems used in administrative decision-making.
The core mechanism is the Algorithmic Impact Assessment: a mandatory questionnaire of 65 risk questions and 41 mitigation questions that classifies a system into one of four impact levels, with requirements scaling by level — peer review, transparency to affected parties, human oversight, and ongoing monitoring.
Two points matter for anyone selling into government:
- Vendors are in scope by contract. System vendors and integrators are contractually obligated to provide the documentation and monitoring data the Directive specifies.
- Procurement is the trigger. Every procurement or material modification involving an automated decision system must be integrated with Directive protocols throughout the project lifecycle — not assessed once at the end.
Systems deployed or procured before June 24, 2025 were required to reach full compliance with the Directive's requirements. If your product sits inside a federal workflow and its documentation was written for a sales conversation rather than an assessment, that gap is now visible to your buyer.
The five parts of a defensible file
A file that holds up has these, captured at the time:
- The inputs. What was actually submitted or retrieved, in the form the system received it.
- The evidence, with stable identifiers. Not "the policy manual" — the specific passage, with a durable reference someone else can pull up and read.
- The system and version. Which model, which retrieval configuration, which date. Behaviour changes between versions; a file that does not name the version cannot explain its own output.
- The human review step. Who reviewed it, what they reviewed, and the scope of their authority to override. "A human was in the loop" is not a finding; a named reviewer with a defined role is.
- Contemporaneous reasons. Why this decision, recorded then — not reconstructed after a challenge arrives.
That last one carries disproportionate weight. Reasons written in response to a complaint are, correctly, treated as advocacy. Reasons written at the time are treated as evidence.
The failure mode nobody plans for
The common failure is not a dramatic hallucination reaching a decision. It is this: the system produced a reasonable output, an experienced person read it and agreed, the decision was made, and nothing recorded what the person checked. Six months later the file shows an outcome and no reasoning.
At that point the organization has two bad options: reconstruct the reasoning from memory, which is weak and can be impeached, or concede that the basis cannot be shown. Neither is a model problem. Both are records problems that were solvable at design time and are nearly unsolvable afterward.
Build the record before you need it
Three actions, in order of leverage:
- Pick the decisions most likely to be challenged. Not every workflow needs this. The ones where an external party has standing to object do.
- Instrument those workflows to capture the five elements above — logging retrieved passages with stable identifiers is usually the missing piece, and it is an engineering change, not a policy change.
- Run a reconstruction drill. Pull a decision from three months ago at random and try to rebuild the file. Whatever you cannot find is the specification for what to start capturing.
The organizations that handle a regulatory challenge well are almost never the ones with the best model. They are the ones who can produce a clean file in an afternoon.
If you are mapping this before deployment, our AI Adoption Readiness Sprint produces the risk register, human-oversight boundary map, and source-traceability requirements this work depends on — and, for public-sector suppliers, a procurement readiness view of what an evaluator will expect to see.
SHARE