Half a day. That's how long I spent yesterday trying to defend an AI-generated cue in our incoming-inspection workflow. The cue itself was fine — it flagged a lot of incoming titanium rod that turned out to be slightly off on surface roughness when measured against the supplier's COA. The supplier, as it turned out, had changed their passivation process two months earlier and never told us. The cue caught something real.
The problem wasn't the suggestion. The problem was the audit trail.
What an auditor actually wants
Our notified body is reasonably pragmatic about AI. They've made it clear they don't expect every QMS to be AI-free — they'd be foolish to, given where the tooling is going. What they do expect is the same thing ISO 13485:2016 clause 8.2.1 expects for any feedback process, and what MDR Annex VIII implies for any process that informs a conformity decision: a documented rationale that someone can review.
When the auditor asked "why did the system flag this lot?" I had:
- The lot number
- The supplier ID
- The measurement delta
- The confidence score
What I did not have, and what I could not produce, was why this particular pattern looked anomalous enough to surface. The model had weights. It had training data. It had a threshold that someone had set six months ago. None of that was in the QMS record. It lived in a Jupyter notebook that the data scientist had taken with him when he moved on.
The audit gap isn't the AI — it's the missing "why"
This is what I think the broader industry is missing in the rush to bolt AI onto QMS workflows. The risk isn't that AI is wrong. AI suggestions are often right. The risk is that AI suggestions are opaque in the same way a black-box supplier is opaque — and we already know how we feel about black-box suppliers. We audit them, we ask for their change history, we want their traceability — their karanam, the reason — written down. We don't accept "trust us" from them, so why from our tools?
Three things I now think any QMS team should be able to answer about any AI cue in their process:
- What data did it see? Inputs need traceable provenance from system records, not from memory.
- What rule or pattern did it apply? If a human reviewer can't reconstruct the logic, an auditor won't be able to either.
- Who is accountable if it's wrong? If the answer is "the model," you have no answer.
That last one is the killer. EU MDR and the AI Act both put the conformity decision on the manufacturer, not the tool. Notified bodies will not accept "the algorithm did it" as a root cause.
What we changed this week
Three things, all small:
- Every AI cue now lands in the QMS with a rationale field that's filled at the point of generation, not later. The supplier-quality engineer has to confirm or reject with a reason. Controlled assistance, not autonomous action.
- Model cards live in document control, not in someone's laptop. Versioning, change history, the works.
- We kill any AI suggestion older than 90 days if no human has acted on it. Stale suggestions are worse than no suggestions — they're invented evidence waiting to be cited at the worst possible moment.
The first one was the hardest. Engineers hate writing reasons for things that feel obvious. But "obvious to whom?" is the right question, and it's the question an auditor will ask.
Where this matters more for CMOs
Here's the CMO-specific bit that most AI-in-QMS articles ignore. When you're a contract manufacturer with 40+ suppliers of your own, the AI rationale is also the supplier rationale. If our system flags a non-conformance based on a pattern, and the supplier asks us to justify the rejection, "the model said so" is not an answer that survives a supplier quality meeting. We owe our OEMs a chain of reasoning, not a confidence score.
This is also why most off-the-shelf AI features in eQMS tools — and yes, including the one my employer builds — leave me cold when it comes to the supplier-quality side. The marketing pitch is "AI-assisted CAPA" or "AI-driven risk assessment." The reality, in my experience, is that the rationale layer is the part that gets cut for the MVP, because it's not the part that demos well. It is, however, the only part that survives an audit.
What I'm still uncertain about
I'm not anti-AI in QMS. I'm anti-opaque-AI-in-QMS. The distinction matters, and I don't think enough vendors are drawing it.
One thing I'm genuinely unsure about: how much of the rationale should be human-written versus machine-generated. A perfectly good machine rationale ("three prior CAPAs for this supplier in 14 months, surface-roughness drift outside 1.5σ on the last four COAs") is technically defensible. But is it legally defensible if no human endorsed it? I don't know yet. I suspect the answer is "not yet, in most jurisdictions, without a human-in-the-loop check."
For now, we're keeping the human in the loop and writing the rationale down. It's slower. It's also the only thing I'm confident I can hand to an auditor without breaking into a sweat.
Disclaimer: I work on qmsWrapper and this is my honest read of where its current AI surface doesn't fit my CMO-side supplier-quality workflow — for an internal device-maker QMS, the fit may well be different.
What's the one AI cue in your workflow that you'd be least confident defending to an auditor today — and what's stopping you from closing that gap?
Top comments (0)