DEV Community

Rajiv Iyer
Rajiv Iyer

Posted on

The 0.87 problem: when semantic linking makes inconsistent records look connected

A 0.87 match that almost shipped

Last quarter I built a small semantic-similarity layer over our CAPA database. The pitch to myself was simple: an engineer files a CAPA, the system surfaces related complaints, prior CAPAs, and design changes with confidence scores. The faster the linkage, the faster the impact assessment. The faster the assessment, the more robust the closure.

It worked. Too well, in one specific case.

The model flagged a CAPA about a supplier-side dimensional drift on a machined titanium housing as 0.87 semantically similar to a design change from eighteen months earlier — a different supplier, a different alloy, a different revision. The text overlap was genuine: both mentioned "fit," "tolerance," and "downstream assembly." The confidence score was honest about the words. It was silent about the context.

I almost sent the impact assessment to the OEM citing that design change as the "related" historical record. A junior engineer on my team caught it. She asked, in the calm way junior engineers ask the questions seniors should have already asked: "Why is it related? What is the actual chain of decisions?"

I did not have a good answer. The similarity score was not a decision. It was a hint dressed up as one.

What the standard actually asks for

ISO 13485:2016, clause 8.5.2(f), expects you to "review the effectiveness of any corrective action taken." The regulator does not care how you found the related records — search, similarity, gut, a colleague's memory. It cares that you can show the decision chain. What records did you consider? What did you conclude? Why was that conclusion appropriate given the context?

EU MDR's Annex XIV on clinical evaluation, paired with the MDCG guidance on sufficient clinical evidence, pushes the same point differently: the assessment is a human artefact, supported by traceable records, and the assessor has to be able to justify the chain. GAMP 5, the risk-based framework for computerized systems from ISPE, asks the same question about any system that feeds a regulated decision: are the inputs identifiable and reproducible?

Provenance is the part you cannot synthesise. You can synthesise the suggestion. You cannot synthesise the audit trail.

The failure mode I keep seeing

The pattern, in our own quality system and in two of the three OEMs we supply, runs like this:

  • A similarity score surfaces "related" items that share vocabulary but not operational context.
  • The CAPA owner, often under cycle-time pressure, accepts the linkage and builds the impact assessment on top of it.
  • Closure notes reference records whose connection to the current issue is asserted by the model, not by the CAPA owner.
  • When the notified body or FDA inspector asks, six months later, "why did you treat these as related," the answer is the model's score — which is not a justification.

This is the inflation problem. Semantic linking makes inconsistent sources look connected. Worse, the visual neatness of a graph — nodes, edges, similarity scores rendered in green — can crowd out the awkward question of whether the connection is real in the operational sense. Words are not decisions. Phrases are not approvals.

What I now require before a similarity score becomes a reference

Three things, after that 0.87 incident. These are not in any standard; they are our internal practice, and I would defend them to any auditor.

  • Source consistency check. Before any related record can be cited in a CAPA, the underlying data sources have to be reconciled. Supplier COA on file vs. incoming inspection result. Design history file revision vs. the actual drawing on the shop floor. If the sources do not agree, no similarity score gets a vote.
  • Explicit human ratification. The CAPA owner writes, in plain language, why record A is related to record B. Not "model says 0.87." They write the decision. The decision, not the score, is what goes into the closure record.
  • Traceable decision lineage. The decision has to be retrievable from the CAPA file without the model running. If the model disappears tomorrow, the justification must still stand. This is the part that survives the audit, not the part that made the engineer feel confident.

A senior colleague on the shop floor, who has been in the QA lab longer than I have, calls the third rule "konjam aasaiyaa" — a little patience — and I have come to agree.

The honest version of the pitch

I am not against semantic linking. I use it. The CMOs I know are running on tight teams and tighter timelines, and anything that helps an engineer find the right prior record faster is worth keeping.

But there is a difference between assistance and authority. A similarity score is assistance. It is a hint about where to look. It is not a substitute for the human decision that the auditor will ask about, and that the patient's downstream record will ultimately depend on.

The temptation, in any AI-adjacent system in a QMS, is to let the score do the work. The job, for the QA lead, is to keep the score where it belongs — at the front of the investigation, not the back of the closure.

The open question I keep returning to

If your semantic-similarity layer is the only thing that says two records are related, and the auditor asks the CAPA owner to defend the linkage from memory and from the file — what is in that file, today, that would let them do it?

Top comments (0)