A customer uploads an identity document.
OCR extracts the name and document number. A model classifies the image. Another component detects suspicious regions. AI compares extracted attributes with application data, and eventually the customer becomes KYC-approved.
From the application perspective, this can look simple:
Document Upload
↓
OCR
↓
AI
↓
KYC Approved
Architecturally, that hides an important question:
Can we reconstruct exactly what evidence, transformation, model, rule, and human decision produced the final result?
That is where KYC starts becoming an evidence provenance problem rather than only a document-processing problem.
The Problem
A document is evidence.
An OCR result is a derived interpretation of that evidence.
An AI-generated conclusion is another derived claim.
A final KYC status is a policy or human decision based on some combination of those artifacts.
These should not be treated as equivalent.
For example, OCR may correctly extract:
Name: Rahul Sharma
Document Number: ABC12345
That does not prove the document is genuine.
Similarly, a cryptographic hash can help establish that stored bytes have not changed, but it does not establish that the underlying document is authentic.
The architecture needs to preserve these distinctions.
Where It Gets Difficult
1. Derived data can silently replace evidence
Suppose OCR initially extracts the wrong document number and a later model produces the correct value.
If the system simply updates:
document_number = corrected_value
the historical reasoning disappears.
A better model keeps the lineage:
SOURCE_DOCUMENT:v1
↓
OCR_RESULT:v3
↓
ATTRIBUTE_SET:v2
↓
POLICY_EVALUATION:v7
↓
KYC_DECISION:v1
If OCR is rerun later using another model version, it should create another branch rather than rewrite the evidence used by the original decision.
2. AI inference can become accidental truth
AI may be useful for:
- comparing names,
- identifying inconsistent fields,
- classifying unusual layouts,
- summarizing discrepancies,
- prioritizing manual reviews.
But an AI result should remain a derived claim.
If one document contains:
VAIBHAV S SHAKYA
and another contains:
VAIBHAV SHAKYA
a model may conclude:
Likely same identity
That conclusion can be useful without rewriting either source.
The system should preserve:
- both source values,
- the comparison operation,
- model or algorithm version,
- resulting interpretation,
- downstream policy decision.
3. Retries can create competing evidence chains
Imagine an upload reaches the backend, but the client times out before receiving the response.
The client retries.
Without idempotency, the platform may create two evidence records, trigger two OCR jobs, and run two policy evaluations for what the user considers one submission.
One branch may pass automatically while another reaches manual review.
For evidence-processing workflows, idempotency is not only a reliability concern.
It protects the coherence of the provenance chain.
4. Current state is not historical reasoning
A customer profile might contain:
kyc_status = VERIFIED
verified_name = "Ananya Rao"
risk_level = LOW
That is useful operational state.
It cannot answer:
- Which artifact introduced the name?
- Which validator confirmed it?
- Which model or normalization version processed it?
- Which policy approved the customer?
- Did a reviewer override an earlier decision?
Current state and historical evidence usually need different models.
Architectural Direction
A stronger design treats every material artifact as part of a provenance graph.
The raw uploaded document becomes a source node.
Every processor creates a derived artifact that references its input.
Conceptually:
Raw Evidence
↓
OCR Extraction
↓
Document Validation
↓
Attribute Verification
↓
AI Interpretation
↓
Policy Evaluation
↓
Human Review
↓
KYC Decision
A few principles make this architecture much easier to reason about.
Preserve source evidence separately
Do not retain only the resized, enhanced, cropped, or OCR-optimized representation.
Where retention policy permits, preserve the original source artifact separately from derived versions.
Build around evidence IDs
Derived artifacts should reference their parent evidence instead of overwriting it.
For example:
Evidence derived = evidenceRepository.append(
Evidence.create(
EvidenceType.OCR_RESULT,
rawDocument.id(),
sha256(ocrJson),
ocrEngine.version(),
clock.instant()
)
);
The important idea is append.
The OCR result becomes another artifact in the chain.
Keep inference separate from policy
Avoid allowing a probabilistic component to directly mutate customer state:
model.score > 0.91
→ VERIFIED
Prefer:
AI inference
↓
Derived signals
↓
Versioned policy
↓
APPROVE / REVIEW / REJECT
The model produces signals.
The policy decides what the organization is allowed to do with those signals.
Version anything that can influence a decision
That may include:
- OCR models,
- AI models,
- normalization algorithms,
- prompt or task definitions,
- schemas,
- policy rules,
- decision thresholds.
Without versions, historical reconstruction becomes guesswork.
A Practical Example
Suppose rejection rates suddenly increase after a deployment.
The current database only says:
status = REJECTED
reason = NAME_MISMATCH
Without provenance, engineers start reproducing uploads and searching logs.
With lineage, they can follow the decision.
The rejection used policy version 17.
Policy version 17 evaluated:
RAHUL K SHARMA
That value came from normalization service version 6.
The parent OCR value was:
RAHUL KUMAR SHARMA
Another evidence source contained:
RAHUL SHARMA
The customer's documents did not change.
The transformation in the middle of the chain did.
That is exactly what evidence provenance should make visible.
The Takeaway
AI-assisted KYC should not treat the uploaded document, AI output, and final KYC status as interchangeable truth.
They represent different layers:
Evidence records what existed.
Interpretation records what the system inferred.
Policy records what the organization decided to do with that inference.
The goal is not to log everything forever.
The goal is to make consequential decisions reconstructable without silently destroying the evidence and reasoning that produced them.
Want the deeper architectural breakdown, implementation considerations, failure scenarios, privacy trade-offs, cryptographic chaining discussion, and full reasoning?
Read the full article on Medium
#SoftwareArchitecture #ArtificialIntelligence #KYC #SecurityArchitecture
Top comments (0)