AI can support a public sector evaluation, but it cannot make the award decision. The Procurement Act 2023 requires assessment against published award criteria and an assessment summary giving each supplier reasons, so a named officer must own every score. Record what the AI produced and what the panel decided: a challenge will test exactly that.
Can public sector buyers use AI to score bids?
Yes, with one hard limit: AI can help a panel work through tender text, but it cannot be the thing that decides who wins. Nothing in the Procurement Act 2023 mentions artificial intelligence, and nothing in it bans the use of software in an evaluation. What the Act does is place duties on the contracting authority that a fully automated score cannot discharge on its own.
Buyers keep asking me whether there is a rule against AI scoring. There is no prohibition. There is something more awkward: a set of obligations about published criteria, reasons and record-keeping that only a person can carry. Once you read the Act that way, the design question stops being "is this allowed" and becomes "which part of this work can the machine do without taking the decision away from the officer who has to sign it".
What does the Procurement Act 2023 require on award criteria?
Section 23 of the Act sets the frame. Award criteria must relate to the subject matter of the contract, be sufficient to allow the authority to assess tenders, and be proportionate. The authority also has to set out how tenders will be assessed against those criteria and the relative importance of each one, and it has to do that before bids come in. Section 24 allows criteria to be refined during a competitive flexible procedure, but only inside the limits that section sets, and only where the authority reserved that possibility in advance.
The practical consequence for anyone putting AI near an evaluation is blunt. Every mark has to be traceable back to a criterion you published. If your tool has learned that longer answers read better, or that a tidy management summary signals competence, it is applying a criterion nobody published and nobody can see. That is the failure mode I would look for first if I were acting for a losing bidder. It does not announce itself in the output. It shows up as a pattern across marks that the published methodology cannot explain.
Can an AI score count towards the award decision?
It can inform the decision. It cannot be the decision. The award has to be made on the basis of the most advantageous tender, assessed against the published criteria, and the authority has to be able to say why each tender scored what it scored. The UK government's AI Playbook puts the same point as a principle: you have meaningful human control at the right stage. A panel that accepts a generated mark without reading the answer behind it has human involvement, but not at the stage that matters.
There is a simple test I give buyers. Take any mark and ask the evaluator to point to the passage in the bid that justifies it, and to say why it is not a grade lower. If they can do that from the bid itself, the score is theirs and the tool was an assistant. If the answer comes back as a version of "the system gave it a four", the score belongs to the software, and the authority has no defensible reasons to put in front of a supplier or a judge.
What must an assessment summary contain?
Section 50 is where this becomes concrete. Before entering into a public contract the authority publishes a contract award notice, and for a competitively tendered contract it must provide an assessment summary to each supplier that submitted an assessable tender. That summary covers the authority's assessment of that supplier's tender against the award criteria, together with its assessment of the winning tender. The detail of what goes into it sits in regulations made under the Act.
So the reasons are not optional and they are not internal. Each losing bidder receives a written account of how its tender was judged and how the winner's compared. Those reasons have to be the panel's reasons, in the panel's words, matched to the marks actually given. Generated narrative is dangerous here for a specific reason: it tends to be fluent and slightly generic, and if the wording drifts from the mark, you have created two documents that disagree with each other. A challenge is often built from exactly that gap.
How does AI-assisted evaluation survive a challenge?
A challenge tests process, not taste. You survive one by answering from a contemporaneous record of what was scored, by whom and why, rather than reconstructing the story months later. During and after the standstill period the questions land in a predictable order. What criteria did you publish. Who scored each response. What did they say. What changed at moderation, and why. If AI was involved, three more arrive: what were you given as input, what did you return, and what did the evaluator do with it.
You survive that by answering from a contemporaneous record rather than reconstructing the story months later out of inboxes and spreadsheet versions. Reconstruction is where authorities lose ground, because the reconstruction itself becomes the thing under argument. If you cannot show that the model's output was not the decision, you end up arguing the one point you least want to argue.
What record should the authority keep?
Keep more than most people expect, and keep it as it happens. For each bid and each question: the criteria version that was live, the exact input handed to the tool, the exact output it returned, the named evaluator, their mark, their written reason, and every moderation change with the person and the rationale behind it. Add the model version and configuration in use, because "the tool" in March is not the tool in September.
Then there is the integrity question, which is the part that gets skipped. A log the authority can quietly edit is worth much less in a dispute than one it cannot. That is the whole point of a verifiable AI audit trail, and it is the problem the Mickai Sovereign Intelligence Operating System was built around. Our Open Audit Record seals every consequential action under ML-DSA-65, the post-quantum signature scheme NIST published as FIPS 204 in 2024. It is worth being precise about what that buys you. The record is tamper-evident, not tamper-proof. Nothing prevents someone altering a stored file. Sealing means the alteration makes verification fail, and an auditor can establish that offline, from an exported record and a public key, using tools that are not ours: any ML-DSA-65 implementation that conforms to FIPS 204 will do. Our record format notes set out what to point one at. Consequential actions also wait for a named person to approve them, which is the control the Act already implies.
Where can AI safely help an evaluation panel?
The useful work sits upstream and sideways of the mark. Retrieval across long submissions. Cross-referencing an answer against the method statement and the pricing schedule. Completeness and arithmetic checks. Drafting clarification questions for a human to send. Flagging where two evaluators have diverged, so moderation starts in the right place. None of that produces a score or a reason, and all of it saves a panel real hours.
Two boundaries. AI should not run the moderation stage, because moderation is where disagreement between named evaluators gets resolved, and software cannot hold the view it is meant to be defending. It can prepare the meeting. It cannot chair it. And confidentiality is live, because tender responses are commercially sensitive, so pasting them into a hosted service is a disclosure you have to be able to account for. That is why SIOS runs on hardware the authority owns, offline capable, with no data egress. Cloud stays valuable for plenty of non-regulated work. An open tender box is not that work.
One caveat on documents, since bids arrive as PDFs. A local OCR runtime has read scanned PDFs in controlled tests here, and the extraction and ingestion integration into SIOS is still being completed. I would rather say that than imply it is finished. The closed beta is running now, with one regulated company onboarding as a design partner.
None of this is an argument against the companies building the compute or the cloud layer. It is an argument against the assumption that a regulated body has to rent its intelligence, send its evidence somewhere else and take a supplier's word about what happened to it. In procurement that assumption is harder to hold than anywhere else, because the award decision has to stand up in public, attached to a name.
Frequently asked questions
Can AI score a tender for us?
Not in any way that ends the matter. A tool can summarise a response, check it against your method statement and flag gaps, but a named evaluator has to read the answer, set the mark and write the reason behind it. The Procurement Act 2023 makes the authority accountable for the assessment, never the software it happened to use.
Does the Procurement Act 2023 allow AI in evaluation?
The Act does not mention AI, so it neither authorises nor forbids it. It regulates the decision instead. Sections 23 and 24 require assessment against published award criteria, and section 50 requires an assessment summary giving each bidder reasons. Software can help a panel meet those duties faster. It cannot discharge them, because it cannot be accountable for them.
Who signs off an AI-assisted score?
The named evaluator who set the mark, and then the panel or moderation chair who confirms it. Sign-off means that person can point to the passage in the bid behind the score and say why a lower grade would be wrong, without referring to the tool. If the only available reason is the output, nobody has really signed anything off.
What do we put in the assessment summary if AI helped?
Write the panel's reasons, matched to the marks actually awarded, in the panel's own language. Do not paste generated narrative: fluent text that drifts from the mark leaves you with two documents that disagree. Whether you also name the tool is a transparency judgement for the authority, but the reasons themselves must be the human assessment.
How likely is a legal challenge if we use AI in evaluation?
The risk sits in being unable to explain a score, not in the tool itself. Challenges test process: what criteria you published, who scored, what changed at moderation and why. An authority holding a contemporaneous, verifiable record of input, output and human decision is in a strong position. One reconstructing the story after standstill is not.
Related briefings
Procurement and defence
- Buy AI Through G-Cloud and Government Frameworks: Guide
- AI Disclosure in Public Sector Tenders: What's Required?
- AI for Procurement Contract Review: Is It Allowed?
- AI Supplier Contract Clauses: A UK Buyer's Checklist
- JSP 936 for AI Suppliers: What the MOD Now Requires
UK AI regulation
- Is There a UK AI Act? How the UK Regulates AI Today
- AI Cyber Security Code of Practice: Who It Applies To
Part of a series of 60 briefings on deploying and governing AI in UK regulated organisations, archived with a DOI at 10.5281/zenodo.22975756.
Evaluating AI for a regulated organisation? Mickai runs on hardware you own, offline. Consequential actions wait for a named person to approve them, and what the AI did is sealed into a signed record an auditor can check without us. Applications for the invitation-only closed beta are open. Apply for the closed beta.
Written by Micky Irons, founder and chief executive of Mickai LTD.
Top comments (0)