If an AI provider sits inside an important business service, its outage is your outage, and your tolerance must reflect that. Under SS1/21 and PS21/3 you map the service end to end, including the AI dependency, and test the failure. Running with no internet connection removes that provider, though your hardware and people remain in the map.
What counts as an important business service?
An important business service is something you deliver to an external end user where disruption could cause intolerable harm to clients or market participants, or threaten the soundness of the firm or the wider UK financial system. It is a service, not a system: making a payment, settling a trade, paying a claim.
That distinction matters more once AI is in the estate. Most firms inventory AI by tool. The PRA and the FCA ask you to work the other way round. Start at the service the customer depends on, then trace every resource that has to work for the service to be delivered: people, processes, technology, facilities, information, third parties. If a model sits anywhere on that chain, it is in scope, whether or not anybody internally has called it critical.
The transitional period closed on 31 March 2025, so this is a live supervisory expectation rather than a programme with runway. If you also run EU operations, DORA is the equivalent frame there and the mapping work is broadly similar.
Where does AI sit in the service map?
Wherever it touches a step that has to happen. In practice there are three positions: a decision inside the service, a control wrapped around the service, or a support function whose failure delays the service.
The third one catches people out. A model that drafts suitability letters is obviously inside advice. A model that triages complaints, summarises onboarding packs or helps the engineers who fix the service is further back, but the delay still lands on the customer. Map it by the party who can change it, not by the logo on the invoice. If your case management vendor calls an inference provider, that provider is on your map as a fourth party, and a model version replaced without notice is a change to your service that you did not authorise. Record the provider, the model version and who holds the account.
Does an AI outage count as disruption under SS1/21?
Yes, if the service degrades past your tolerance. The rules are outcome-based and indifferent to cause. A cable, a data centre and an inference provider having a bad afternoon are the same category of event: the service stopped being delivered inside a duration you called tolerable.
AI fails in ways more awkward than being down. Rate limits bite at peak. Latency triples while everything still returns a 200. A safety filter changes and starts refusing a prompt that has worked for a year. Output quality shifts quietly and nobody notices for a fortnight. Available but wrong is harder to detect than offline, and it is the failure mode most AI monitoring misses. Build detection for quality and refusal rates, not only for availability, or you will breach a tolerance without an alert firing. The UK critical third parties regime gives regulators supervisory reach over designated providers, but it does not move the obligation: the service is still yours.
How do you set a tolerance around an AI dependency?
Set the tolerance on the service first, then test whether the AI dependency lets you stay inside it. The other way round gives you a tolerance shaped by whatever your vendor would commit to, which is not the point at which customers start being harmed.
Three pieces of work follow. State the maximum tolerable duration and the metric you will measure it with. Establish what the service actually does when the model is absent, at real throughput, with the staff you have on a Tuesday in August. Then count the queue. As arithmetic rather than an industry figure: if a team clears forty reviews a day by hand and four hundred with assistance, a two-day outage leaves around seven hundred items behind, and your recovery window is the time to clear that backlog, not the time the provider takes to come back. Tolerances get breached in the catch-up, not the incident.
What does severe but plausible scenario testing look like for AI?
It looks like removing the provider, not slowing it down. Severe but plausible means the dependency is gone for a period you would not have chosen, and you still have to show you stayed inside tolerance.
Scenarios worth writing down: the inference provider is unavailable for forty-eight hours across the region you use; the model version you validated is withdrawn and the replacement behaves differently; the account is suspended over a commercial or compliance dispute; the network path fails while the provider is perfectly healthy; a policy change starts refusing a class of your inputs. Run each with the people who do the work, not as a tabletop reading of the supplier's runbook. Record what broke, what the manual path really achieved per hour, and what decisions were made without evidence because the usual summary was missing. The lessons learned step is a rule, not a courtesy.
What does a credible substitute for an AI provider look like?
Substitutability is the standard applied to third parties, and for AI it usually fails on the quiet details. A second provider is not a substitute if your prompts, thresholds, validation evidence and audit trail do not port to it.
Ask three things. Can you switch inside your tolerance, using a procedure somebody has actually executed? Does the alternative hold its own validation, or would you be running an unevaluated model in a regulated process mid-incident, which is worse than a graceful manual fallback? And is it genuinely independent, given how many vendors rest on a small number of inference estates. Two suppliers on one estate is one supplier with two invoices. If the honest answer is that no substitute exists, say so in the self-assessment and rely on the manual path. A tested manual fallback is better evidence than an untested second vendor.
How does running AI offline change the mapping?
It removes one party from the chain. The Mickai Sovereign Intelligence Operating System runs on hardware you own, is offline capable and sends no data out, so there is no inference provider to lose, no account to suspend and no model swapped under you. We gate deployment to hardware we have qualified, so the sizing conversation comes before any commitment.
What it does not do is empty the map. Your own estate moves onto it: nodes, site, power, cooling, spares, patching, key custody, and the handful of people who understand the deployment. Model updates become your change control. The Open Audit Record seals every consequential action under ML-DSA-65, the post-quantum signature scheme NIST published as FIPS 204 in 2024, and an auditor verifies an exported record offline with a public key, using tools that are not ours. That is tamper-evident, which is a different claim from tamper-proof. Nothing stops somebody altering a record. An altered record fails verification, and the failure is visible to whoever checks. For resilience evidence that is the useful property, because the record outlives the supplier relationship. Consequential actions also wait for a named person to approve them, which keeps a human in the record.
This is not an argument against the companies building the compute or the cloud layer. It is an argument against the assumption that a regulated organisation must rent its intelligence, ship the data offsite and take a vendor's word for what happened to it. Cloud stays valuable for work that is not load-bearing for a regulated service.
What still has to be tested when the AI runs on your own hardware?
Everything you no longer get to outsource. Node loss and site loss as separate scenarios. Restore from backup with model weights and configuration included, timed, carried out by somebody who did not build it. Degraded mode, where the system runs but one part of it is unavailable. Key rotation. Verification of an exported audit record by a party outside the team that produced it. And the approval path when the named approver is on leave, because a control that stalls a service is itself a tolerance risk.
Be specific about maturity in your documentation, including ours. There are 63 studios in the platform, 14 production-ready and 49 in development, and the closed beta currently has one regulated company onboarding as a design partner, whose sector and identity we do not discuss. Document reading shows where the line sits: a local OCR runtime has read scanned PDFs in controlled tests, and the extraction and ingestion integration into SIOS is still being completed. If one of your services depends on reading documents, map that as work in progress, not as a finished control.
The test of all this is not whether the AI is impressive. It is whether you can say, on one page, what your service does when the model is not there, for how long, and who checked.
Related reading: AI for UK financial services, sovereign AI, offline AI, DORA and the ICT third-party test, AI vendors as critical third parties.
Frequently asked questions
Is using AI in a back-office task inside an important business service?
It can be. The test is not where the work sits in your org chart but whether the service to the end user degrades when that task stops. If a back-office AI step gates settlement, claims payment or client onboarding, it belongs on the map. If the service continues unaffected at normal throughput, record why you concluded that and move on.
Does an on-premise AI system remove concentration risk?
It removes the shared inference provider that many firms depend on at the same time, which is the concentration regulators worry about. It does not remove your own concentration: one site, one hardware supplier, one small team who understand the deployment. Map those, and test node loss and site loss as separate severe but plausible scenarios.
How often should we test the AI failure scenario?
The rules require regular testing rather than naming a frequency, and most firms settle on at least once a year. Test again whenever the dependency materially changes: a new model version, a new provider, a change to the fallback, or a change to the people who operate it. Testing under PS21/3 is not a one-off exercise. Treat an unannounced model change by a provider as a trigger for a fresh test.
Who signs off impact tolerances for AI-supported services?
The board, on the same basis as any other important business service. Setting a tolerance is a board judgement about tolerable harm to customers and markets, not a technical estimate, and the self-assessment supporting it goes to the board as well. AI does not create a separate approval route. It changes the evidence the board should ask to see.
Does choosing a UK data region help with operational resilience?
It helps with data protection questions and sometimes with contractual clarity. It does less for resilience, because the control plane, the model and the account can still sit elsewhere, and an upstream failure reaches your UK region anyway. Region choice is a data location control, not a substitute for a fallback you have actually exercised.
Related briefings
Financial services
- Notify the FCA Before Using AI? What UK Firms Must Do
- PRA SS2/21 and AI Outsourcing: Is Your Supplier Material?
- AI Complaints Triage and Vulnerable Customers: FCA Rules
Governance, audit and oversight
- Tamper-Evident vs Immutable Log: The Real Difference
- Human on the Loop vs In the Loop: AI Oversight Explained
Part of a series of 60 briefings on deploying and governing AI in UK regulated organisations, archived with a DOI at 10.5281/zenodo.22975756.
Evaluating AI for a regulated organisation? Mickai runs on hardware you own, offline. Consequential actions wait for a named person to approve them, and what the AI did is sealed into a signed record an auditor can check without us. Applications for the invitation-only closed beta are open. Apply for the closed beta.
Written by Micky Irons, founder and chief executive of Mickai LTD.
Top comments (0)