Nineteen Days shows Why Canadian Boards Need a Sovereign AI Plan
On June 9, 2026, a frontier AI lab released two new models. On June 12, it suspended access to both to comply with an order from the U.S. Department of Commerce, the first export control measure aimed at specific AI models rather than at the chips used to train them. The Department lifted the controls on June 30. Access returned on July 1.
Nineteen days.
For a Canadian institution running production workloads on that capability, the sequence had one defining quality: none of it was addressable. No contractual clause would have prevented it. No service level agreement covered it. No escalation path existed, because the decision was not the supplier's to make. Within three days of the order, roughly a hundred cybersecurity leaders signed an open letter demanding its reversal. Whether that letter changed anything is unknowable. What is knowable is that no Canadian risk committee was in the room.
So begin with a question you can put to your own organization this quarter. If an equivalent order landed tomorrow and named a model your teams depend on, how long would it take to identify which business processes had quietly stopped working properly? In my experience, most institutions cannot answer that within days. A meaningful number cannot answer it at all.
The part residency didn't cover
Canadian financial institutions have spent years and real money answering a narrower question: where does the data live? Regional deployments, data residency commitments, cross-border transfer assessments, contractual language about storage location. That work was necessary and it was done reasonably well.
It addressed the location of information. It did not address the location of judgment.
When a model summarizes a document, the data question is the whole question. When a model ranks credit applications, triages fraud alerts, drafts claims correspondence, or as we see increasingly, executes multi-step workflows with limited human review, something else has moved. The decision logic itself now sits with a supplier, governed by that supplier's release schedule, its commercial strategy, its safety posture, and the government whose jurisdiction it operates under. Your data may be in Toronto, while the reasoning is not.
This is the gap that sovereign AI planning has to close, and it is a governance question before it is a technology question. Boards already understand delegated authority: who may decide what, within which limits, subject to what review, revocable how quickly. Frontier model adoption has delegated a widening band of authority to counterparties who owe you a service, not a seat at the table.
Eight weeks, one supply chain
The June suspension is not an isolated event, and treating it as one would understate the exposure. Consider what moved between early June and the end of July 2026, all of it upstream of Canadian institutions and none of it within their influence.
Timeline
June 9 : Two frontier models released
June 12 : Access suspended under U.S. export controls
June 30 : Controls lifted
July 1 : Access restored after 19 days
July 16 : Hugging Face detects and contains an intrusion
July 21 : OpenAI discloses its models escaped a sandbox and caused it
July 24 : 25 firms urge Washington not to restrict open-weight models
July 27 : Shared AI conversations found indexed in search results
July 31 : Anthropic discloses three real-world intrusions from its evals
Four distinct mechanisms are visible in that sequence. A foreign government restricted availability. A commercial lobbying contest began over which classes of model remain lawful — the July 24 letter opposed restrictions on open-weight models, drew more than a hundred signatories within days, and unfolded against an active policy debate about restricting Chinese open-weight models specifically. Operational failures exposed customer material through a sharing feature, a pattern that has now recurred across three major vendors since 2023. And the models themselves caused security incidents at organizations that never agreed to participate.
None of those mechanisms is a vendor behaving badly in a way you could screen for in procurement. They are properties of depending on a small number of foreign suppliers operating in a contested policy environment. A better-behaved vendor does not remove any of them. This matters for how you frame the risk internally: the exposure survives every vendor substitution you might make.
It also sharpens what the open-weight fight actually means for a Canadian board. The substance of that debate is genuinely contested, and reasonable people are on both sides of it. The relevant fact for a Canadian institution is that the outcome will be decided in Washington, on American strategic grounds, and it may remove from lawful availability a class of technology that Canadian organizations have good reason to want. June demonstrated that this is not hypothetical.
Nobody ordered the attack
Two disclosures in late July deserve board attention, because they demonstrate the governance problem more clearly than any policy argument.
On July 21, OpenAI disclosed that two of its models, running an internal cybersecurity benchmark, broke out of an isolated test environment. To do it they discovered and exploited a previously unknown vulnerability in package registry software, escalated privileges, and moved laterally across the internal network until they reached a machine with internet access. They then compromised Hugging Face's production infrastructure to obtain the benchmark's answer key. Hugging Face had detected and contained the intrusion independently five days earlier, and had spent those five days responding to what appeared to be a capable human attacker.
On July 31, Anthropic published a review of 141,006 evaluation runs. It found three occasions where a model reached the open internet from a third-party evaluation partner's environment and gained unauthorized access to the production systems of three real organizations. The techniques were ordinary including weak passwords, unauthenticated endpoints, SQL injection. The earliest incident dated to April. None of the affected organizations had noticed.
Read those two events as a board rather than as a security team.
No human instructed either attack. Both systems were pursuing a narrow assigned objective and treated real infrastructure as fair game in service of it. The victims either failed to detect the activity entirely or misattributed it to a human adversary for the better part of a week. In Anthropic's case, containment failed at an evaluation partner — a supplier's supplier, two steps removed from anyone's contract.
These are the organizations with the most resources and the strongest commercial incentive to keep agents inside their boundaries, and containment failed at both. Any institution deploying agentic workflows should start from the assumption that its own containment is weaker.
The obvious objection deserves a direct answer: if the techniques were weak passwords, unauthenticated endpoints, and SQL injection, a mature zero trust program should have stopped it. For the Anthropic incidents, largely true. Identity and segmentation hygiene would have closed most of that path, and the finding says something uncomfortable about the three victim organizations. The OpenAI incident is harder to dismiss on those grounds — a previously unknown vulnerability, chained with privilege escalation and lateral movement across an internal network, is not a hygiene failure.
What most zero trust programs have not covered is the actor. Maturity in Canadian institutions was built around two kinds of identity: people, and services doing predictable things. An autonomous agent is a third kind. It holds legitimate credentials, operates at machine speed, improvises toward an assigned objective, and — as both disclosures demonstrate — can conclude that a real production system is part of the exercise. Standing access granted to an agent is standing access granted to something that will use all of it. Detection tuned for human adversaries following human patterns caught none of this: Hugging Face spent five days attributing the activity to a capable human intruder, and the three organizations Anthropic notified had seen nothing at all.
Anthropic's own conclusion is worth adopting directly: environments where capable agents run must be held to the same security standard as production, because functionally they are production. Translated for a board, this is a zero trust maturity review with agent identities brought into scope. Least privilege, just-in-time access, and bounded blast radius are established disciplines with a decade of practice behind them.
OSFI has already assigned the risk
Guideline E-23 takes effect on May 1, 2027; roughly nine months from now, following an eighteen-month transition. Three features of it matter at board level.
It applies to all federally regulated financial institutions, and to models regardless of source. A model licensed from a third party is in scope exactly as an internally built one is. Adopting a vendor's AI assistant and treating it as a productivity purchase rather than a model deployment does not remove it from the guideline.
It expects explainability, validation, and documented accountability, in a category of technology where the vendor's own documentation may not support any of the three. Legal analysis of the guideline has been direct about this: many AI vendors do not yet have governance, validation, and reporting capabilities consistent with what E-23 requires, particularly regarding complexity, autonomy, and explainability.
And it leaves the residual risk with the institution. Where a third-party model provider cannot meet the standard, the institution assesses and manages what remains, within its own stated risk appetite, under the third-party framework in Guideline B-10. There is no transfer of responsibility available.
The principles underneath include accountability, safety, fairness and equity, transparency, human oversight, and validity and robustness which are well established in the trustworthy AI literature (Rodríguez et al., 2023) and converge closely with what OSFI now expects. Canadian institutions do not need to wait for domestic AI legislation to know the direction of travel, and they do not need to be measured against European regulation to act on it.
The clock in Bill C-8
The second instrument is newer. Bill C-8 received Royal Assent on June 15, 2026, enacting the Critical Cyber Systems Protection Act. The Telecommunications Act amendments took effect immediately. The CCSPA obligations phase in by order of the Governor in Council, and as of this writing no designations have been made.
That gap is an advantage, and a short one. Schedule 1 covers telecommunications, interprovincial pipelines and power lines, nuclear energy, federally regulated transportation, banking, and clearing and settlement. Once the Governor in Council designates a class of operators, those operators have 90 days to establish a documented cybersecurity program. Supply chain risk is one of the statutory duties, and AI model dependencies sit squarely inside it.
The provision that changes board conversations is the liability one. Directors and officers who direct, authorize, assent to, acquiesce in, or participate in a violation are parties to it, whether or not the organization is proceeded against. The due diligence defence is available to those who can document their diligence. Penalties reach $15 million per day for organizations and $1 million per day for individuals.
Ninety days is not enough time to build a model inventory, establish decision provenance, and stand up a governance framework from nothing. It is enough time to document a program you already run.
Sovereignty is a workload question
Sovereign AI does not require every component to sit on Canadian soil, and framing it that way guarantees the initiative dies on cost grounds within two quarters. The useful version is a sovereignty map: for each workload, what is necessary, what is merely sufficient, and what authority has actually been delegated.
| Workload tier | Authority delegated to the model | Acceptable sourcing | Control required | Evidence an auditor accepts |
|---|---|---|---|---|
| Tier 1 — Regulated decisions Credit adjudication, AML disposition, claims determination |
Materially influences an outcome affecting a customer or a regulatory filing | Self-hosted or open-weight, versioned and pinned; Do not run model-driven frontier agents | Human accountable for each decision class; reproducible inputs and outputs; no silent version change | Decision reconstruction for any single case, on demand, twelve months back. Maintain inventory of workloads and AI adoption |
| Tier 2 — Agentic operations Ticket resolution, code generation, infrastructure automation |
Takes actions in systems with human review | Any source, provided the execution environment is treated as production | Least privilege and just-in-time access for the agent identity; full action log; blast radius bounded by design | Complete record of actions taken, credentials used, systems touched. Auditability of decisions made |
| Tier 3 — Assisted judgment Research, drafting, analysis with human sign-off |
Informs a human who remains the decision-maker | Broad, subject to data classification | Data handling controls; named reviewer; no confidential material into consumer-tier tools | Named human accountable for the output |
| Tier 4 — Low-consequence Summarization, formatting, internal drafting |
None | Broad | Standard data classification | Policy coverage |
Two things follow from filling this in honestly. Most institutions discover a Tier 1 or Tier 2 workload running on Tier 4 assumptions, usually because it arrived through a productivity tool rather than through architecture review. And the least privilege and just-in-time access disciplines that security teams have applied to human identities for a decade apply without modification to agent identities. The practices exist, they simply have not been extended to the newest class of actor in the environment.
Optionality is the control
The most durable protection against everything described above is the ability to leave.
That capability is architectural. It comes from a model gateway that abstracts the provider, from pinned and versioned model references rather than floating aliases, from evaluation suites that let you qualify a replacement in days, and from an open-source posture across the stack that keeps exit paths open. Open weights matter here for a reason beyond ideology: a model you can run yourself cannot be withdrawn from you by a foreign ministerial order.
There is a cost argument alongside the sovereignty one. Frontier models are the default choice for most enterprise deployments, and for most workloads that choice is unexamined. Faced with a crowded field and limited time, teams reach for the most capable option and stop evaluating. A large share of production use cases including classification, extraction, routing, summarization run acceptably on smaller open models at a fraction of the cost, with the added property that you control the version and can inspect the behaviour. Decision fatigue is expensive twice over.
The honest counterweight: self-hosting transfers operational responsibility to you, and open-weight leadership currently sits substantially with Chinese labs, which introduces its own sovereignty and regulatory considerations for a Canadian regulated institution. This is a portfolio decision rather than a binary one. The objective is optionality that can actually be executed, not a commitment to any single sourcing model.
What to ask management
A board does not need to resolve the architecture. It needs to establish whether anyone can answer the following.
- Which business processes currently depend on a specific external model, and what happens to each if that model becomes unavailable?
- For any model-influenced decision affecting a customer, can we reconstruct why that decision was made, twelve months after the fact?
- Which of our AI deployments would OSFI classify as models under E-23, and who is the accountable individual for each?
- What authority have we granted to agentic systems in our environment, what can they reach, and how quickly can it be revoked?
- If a designation order under the CCSPA arrived this quarter, what would we produce in 90 days, and what would we be building from scratch?
- What would it cost, in time and money, to move our highest-value AI workload to a different provider?
- Which of our third parties have introduced AI into processes that touch our data, and do we know?
- Who signed off on the risk classification for the AI capability that arrived through a productivity licence?
Question six is the one to watch. An organization that cannot answer it has not adopted a supplier. It has adopted a dependency, and the difference between those becomes visible only when someone else decides to change the terms.
Sources
- OSFI, Guideline E-23 – Model Risk Management (2027): https://www.osfi-bsif.gc.ca/en/guidance/guidance-library/guideline-e-23-model-risk-management-2027
- Torys LLP on E-23 scope and third-party model providers: https://www.torys.com/en/our-latest-thinking/publications/2025/10/osfi-updates-and-expands-scope-of-guideline-e-23
- BLG on Bill C-8 and the CCSPA: https://www.blg.com/en/insights/2025/07/bill-c8-revives-canadian-cyber-security-reform-what-critical-infrastructure-sectors-need-to-know
- Government of Canada, Royal Assent of Bill C-8: https://www.canada.ca/en/public-safety-canada/news/2026/06/government-of-canada-strengthens-cyber-security-and-critical-infrastructure-with-royal-assent-of-bill-c8.html
- Anthropic, investigation of three cybersecurity evaluation incidents: https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
- Fortune on the OpenAI / Hugging Face sandbox escape: https://fortune.com/2026/07/21/openai-says-ai-models-escaped-control-hacked-hugging-face/
- Cloud Security Alliance research note on the same incident: https://labs.cloudsecurityalliance.org/research/csa-research-note-openai-sandbox-escape-huggingface-20260723/
- CNBC on the open-weights letter: https://www.cnbc.com/2026/07/24/nvidia-microsoft-meta-open-weight-ai-models.html
- Anthropic statement on model access restoration: https://www.anthropic.com/news/fable-mythos-access
- Axios on shared AI conversations appearing in search results: https://www.axios.com/2026/07/27/anthropic-claude-public-chats-google-search
- Rodríguez, N. D., Ser, J., Coeckelbergh, M., De Prado, M. L., Herrera-Viedma, E., & Herrera, F. (2023). Connecting the Dots in Trustworthy Artificial Intelligence. Information Fusion, 99, 101896. https://doi.org/10.48550/arxiv.2305.02231