You keep oversight by making it structural, not procedural. Every consequential action waits for a named person to approve it, every approval and refusal is written to a signed record, and permissions decide who may approve what. Sampling reviews fail at scale. A complete, independently verifiable record does not.
Why does oversight break down when AI moves from one team to the whole organisation?
Oversight breaks because the method that worked in a pilot does not survive multiplication. In a pilot, one team watches one system and a person reads most of what it produces. That is genuine supervision. Run the same method across forty teams and several thousand actions a day and it becomes a spot check wearing the clothes of control.
I have watched this sequence run in the same order more than once. The pilot succeeds. The organisation approves a wider rollout. Review meetings that once covered every case start covering a sample. The sample shrinks, because the people running it are measured on throughput. Within two quarters nobody can answer a plain question: who approved this, and what did they see when they approved it?
The root cause is that the oversight was procedural. It lived in a meeting, a rota, a spreadsheet of reviewed cases. Procedures degrade under load. Structure does not, because the system will not proceed without it.
So the question is not how to supervise more. It is what to make structural, so supervision does not depend on one person's diligence on a busy Thursday.
What does per-action approval actually mean in practice?
It means the system stops before a consequential action and waits for a named person. Not a role, not a queue, not a service account: an identified individual who saw the specific action and said yes.
In the Mickai Sovereign Intelligence Operating System, every action is classified before it runs. Drafting a summary, retrieving a document, comparing two versions of a contract: reversible and internal, so they run. Sending an external communication, writing to a system of record, releasing a payment instruction, changing a permission, moving data across a boundary: consequential, so they hold.
When an action holds, the approver sees three things. What the system intends to do. What it relied on to decide that. What follows if it proceeds. The approval then carries the approver's identity, the time, and the exact content approved. If that content changes afterwards, the approval is void and the action holds again.
That rule matters more than it sounds. Most oversight theatre comes from approving something in general and letting the specifics drift afterwards.
Per-action approval has a cost. It puts a person in the path of work that would otherwise be instant. The honest position is that you choose where to pay that cost. Which actions count as consequential is a business decision, set by the organisation rather than by us, and the threshold itself is an auditable setting. Change it, and the change is recorded with the name of whoever made it.
Is sampling enough, or do you need a complete record?
Sampling tells you about the sample. A complete record tells you about the population. At pilot volumes the difference is academic. At organisational scale it is the entire argument.
Think about what a regulator, an auditor or a claimant's solicitor actually asks for. They do not ask whether your process is sound on average. They ask about one decision, on one date, affecting one person or one counterparty. If you sampled, and that decision fell outside the sample, you do not have a weak answer. You have no answer.
This is why the Open Audit Record in SIOS is complete rather than representative. Every consequential action produces an entry: the action, the inputs it relied on, the studio that produced it, the person who approved or refused it, and the time. Nothing is summarised away, and nothing is dropped when volume rises.
Sampling still has a job. It is how you spot quality drift and workflows that have quietly stopped working. It is a management tool, not an evidential one.
The ICO's guidance on AI and data protection is direct about accountability having to be demonstrable, with records of processing and of decision logic forming part of that demonstration (ICO). Where a decision produces legal or similarly significant effects, UK GDPR rights on automated decision-making call for meaningful human involvement rather than a nominal sign-off (ICO). "We reviewed ten per cent" says nothing at all about the other ninety.
How do roles and permissions stop approval becoming a rubber stamp?
They stop it by fixing who may approve what inside the system, rather than inside a training deck. A person can approve only actions that fall within their own authority. Outside it, the action escalates or fails. It is not designed to proceed silently, and any escalation is written to the record with the name of whoever handled it.
Two separations do most of the work. The person who configures a workflow should not be the person who approves its outputs. The person who approves an action should not be the person able to edit the record of that approval. Neither is new. Both are ordinary segregation of duties, applied to a system working at a pace no rota can follow.
For firms in scope of the FCA's Senior Managers and Certification Regime, accountability is already mapped to named individuals with documented responsibilities (FCA). Approval rights should mirror that map rather than invent a parallel one that nobody has signed.
The practical cause of rubber stamping is volume plus ambiguity. An approver facing four hundred near-identical requests with no context will click through them, and would be irrational not to. The fix is to hold only what genuinely warrants holding, and to give the approver enough context to make a real judgement in seconds. Fewer, richer decisions produce better oversight than many empty ones.
How do you prove oversight happened months later, to someone who does not trust you?
You hand them a record they can verify without you. That is the only form of proof that survives an adversarial setting: a dispute, an investigation, a claim, a change of supplier.
Every Open Audit Record entry is sealed with ML-DSA-65, the post-quantum signature scheme published by NIST as FIPS 204 in 2024 (NIST). Entries are chained, so an entry's position in the sequence is covered by the seal along with its contents. Export the record, hand over the public key, and the other side verifies it offline using standard tooling that is not ours and that we cannot influence.
Be precise about what that gives you. The record is tamper-evident, not tamper-proof. Nobody can stop a sufficiently privileged person from altering bytes on a disk. What the seal provides is that any alteration makes verification fail, visibly, at an identifiable point in the chain. That is the property worth having. A record that claims it cannot be changed is asking you to trust the claim. A record that reveals when it has been changed asks you to trust nothing.
Traceability of this kind is treated as part of secure operation rather than an optional extra in the NCSC's guidelines for secure AI system development (NCSC), and the accountability duties behind it sit in the Data Protection Act 2018 (legislation.gov.uk).
All of it runs on hardware the organisation owns. The system is offline capable, with no data egress, so the evidence of oversight lives in the same building as the decisions it describes.
What should you ask a vendor before scaling AI across the business?
Ask five questions, and insist on demonstrations rather than assurances.
- Which actions stop for a human, and who decides that list? If the vendor decides, you have outsourced your own risk appetite.
- Is the audit record complete or sampled? Name a single action on a single day and watch how it is retrieved.
- Can a third party verify an exported record with the vendor absent and the system switched off?
- Do approval rights map to your existing accountability map, and can anyone approve their own configuration?
- Where does the data physically sit while all of this happens?
A vendor that answers all five without hedging is offering oversight. One that answers with a dashboard is offering visibility, which is a weaker and quite different thing.
None of this is an argument against the cloud. For work that is not regulated, renting compute remains a sensible way to buy it, and the companies building that layer are doing hard engineering well. The argument is narrower. Where a decision has to be defended years later, oversight cannot be a report you receive. It has to be a property of the system, and the record of it has to be yours.
Mickai LTD is a UK company, Companies House 17166618, held privately by its founder. SIOS ships with 63 studios, 14 production-ready at launch and 49 in development, supported by 50 specialised models. The architecture behind the approval and audit design sits inside 104 filed UK patent applications carrying 2,340 claims: filed, not granted. The closed beta is open, with one regulated company onboarding as a design partner.
Frequently asked questions
What counts as a consequential action?
A consequential action is one that changes something outside the system or is hard to reverse: sending an external communication, writing to a system of record, releasing a payment instruction, changing a permission, or moving data across a boundary. Your organisation sets that list, not the vendor, and any change to the list is itself recorded against a named person.
Does per-action approval slow the business down?
It slows the actions you chose to slow, and nothing else. Reversible internal work runs without interruption. The cost is real, so the answer is to hold only what genuinely warrants a decision and give the approver enough context to judge it in seconds. Holding everything produces rubber stamping, which is slower and worth considerably less.
Can you keep human oversight if the system runs offline?
Yes, and offline operation makes it easier. SIOS runs on hardware the organisation owns, with no data egress, so approvals, refusals and the audit record all stay inside your own boundary. Nothing depends on a supplier's availability or retention policy. The evidence of oversight sits in the same place as the decisions it describes.
Is a complete audit record the same as application logging?
No. Logs are written for engineers, are usually rotated or sampled, and can be edited by anyone with sufficient access. The Open Audit Record is written for evidence: complete, chained, and sealed with ML-DSA-65, published by NIST as FIPS 204 in 2024. It is tamper-evident, meaning alteration makes verification fail rather than being prevented.
Who verifies the record, you or us?
You do, or a third party you choose. Export the record, take the public key, and verify the signatures offline using standard tooling that is not ours. We cannot influence the result. That matters most in exactly the situation where our word would count for least: a dispute, an investigation, or a change of supplier.
Written by Micky Irons, founder and chief executive of Mickai LTD, which builds a sovereign AI operating system for regulated organisations. More at mickai.co.uk.
Top comments (0)