My RAG Eval Passed. The Citations Were Still Wrong
Consider this small regression test:
Question:
What is the current refund period?
Retrieved:
POLICY-2025: Customers may request refunds within 30 days.
POLICY-2026: Customers may request refunds within 14 days.
Generated answer:
Customers may request refunds within 14 days. [POLICY-2025]
The answer contains the current policy.
The expected document was retrieved.
A citation was present.
Several common RAG checks could mark this response as successful. Yet the citation points to evidence that contradicts the answer.
That is a dangerous kind of failure because the response looks verifiable.
The Green Dashboard Missed a Separate Contract
RAG evaluation often combines retrieval and generation into a few broad scores:
- Did search find a relevant document?
- Was the response related to the question?
- Was the answer grounded in the supplied context?
- Did the model include a citation?
Those questions matter, but none proves that a specific citation supports the exact claim beside it.
The TREC 2025 RAG Track treats relevance, response completeness, attribution verification and agreement as separate evaluation layers. That separation is useful because citation quality can fail even when retrieval and generation appear healthy.
A citation is more than formatting. It is a contract between one claim and one piece of evidence.
Break the Answer Into Atomic Claims
Suppose a chatbot produces this response:
The Pro plan allows 20 users and includes priority support. [DOC-7]
That sentence contains at least two claims:
Claim 1: The Pro plan allows 20 users.
Claim 2: The Pro plan includes priority support.
DOC-7 might support the user limit while saying nothing about priority support. A sentence-level evaluator may give the entire sentence a passing score because part of it is supported.
A claim-level evaluator should test each statement independently.
Recent Q-CARE research follows a similar direction by decomposing answers into atomic claims and measuring whether each claim can be verified. Its evaluation correlated more closely with human judgments than several existing RAG metrics across the tested datasets.
Add a Citation Gate After Generation
The generation step should return structured claims instead of only formatted prose.
claims = [
{
"text": "The refund period is 14 days.",
"source_id": "POLICY-2026",
"source_version": "2026-09",
}
]
A separate citation gate can then inspect each claim:
def evaluate_citations(claims, sources, supports):
failures = []
for claim in claims:
source = sources.get(claim["source_id"])
if source is None:
failures.append({
"claim": claim["text"],
"reason": "source_not_found",
})
continue
if source["version"] != claim["source_version"]:
failures.append({
"claim": claim["text"],
"reason": "version_mismatch",
})
continue
if not supports(source["content"], claim["text"]):
failures.append({
"claim": claim["text"],
"reason": "claim_not_supported",
})
return failures
The supports function could use deterministic rules for exact values, a natural-language-inference model for descriptive claims or a calibrated LLM judge.
I would avoid allowing the same model to write the answer and approve its own citations without another check. Research on citation attribution has found that LLM judges can also reject valid evidence, especially when terminology varies.
High-risk claims may still need sampled human review.
Test Evidence Strength, Not Keyword Overlap
A source can be relevant while remaining too weak for the wording used.
Consider these two claims:
Evidence: The update may reduce processing time.
Claim A: The update may reduce processing time.
Claim B: The update will reduce processing time by 40%.
Both discuss processing time. Only Claim A is supported.
This failure is sometimes described as citation laundering: a related source makes an overstated claim appear trustworthy. FORCEBENCH research found that generic support prompts struggled to distinguish appropriately qualified claims from stronger unsupported versions.
Your evaluator should check numbers, dates, conditions, scope and certainty. Text similarity alone cannot do that reliably.
Put the Failure Into CI
Every citation bug found in production should become a regression case.
Useful tests include:
- Two versions of the same policy
- One source supporting only half a sentence
- A correct fact attached to the wrong document
- A citation pointing to a deleted or inaccessible source
- Evidence using “may” while the answer says “will”
- Two documents that disagree about the same value
For teams evaluating an AI chatbot development company, asking how citations are tested is more revealing than asking whether the chatbot “uses RAG.”
A healthy RAG system should prove four things independently: it retrieved useful evidence, produced a correct answer, covered the important claims and attached each claim to evidence that actually supports it.
My evaluation passed because it tested whether citations existed.
The better test asks whether every citation deserves to be there.
What citation failure would you add to this regression suite?

Top comments (0)