A citation under an answer is only as good as whoever wrote it. If the model writes it, it's exactly as reliable as everything else the model says.
When we built yOGI Neural Grid DocQA for regulated teams, the brief came down to one line: an answer is either backed by a quote that has been mechanically checked against the document it cites, or it isn't shown. This is how that works, and the things we got wrong on the way.
A prompt is a request
The prompt does ask the model to quote its sources and to say when the passages don't answer the question. Models mostly comply. "Mostly" is the problem when the reader is a pharmacist checking a dose. So the prompt is treated as a request and the guarantee lives in a verification step that runs after the model replies.
The model returns JSON: the answer, the quotes it relied on with a marker for each passage, and a flag saying whether it thinks the passages were sufficient. Then three rules apply.
- Every quote has to appear in the passage it's attributed to.
- A sentence in the answer only counts as supported if it carries a marker for a passage whose quote passed rule 1.
- Document name, section and page are looked up from our own index using the passage ID. The model never gets to write them.
If too little of the answer survives, it's withheld, and the user is told that rather than shown a weaker answer.
Character matching was the wrong idea
Our first matcher compared quotes to passages character by character. It looked fine until we fed it deliberately invented quotes, and found it gave them a lot of credit for incidental shared letters. A paraphrase scored far too close to a real copy.
Comparing token by token fixed it. A dropped "the" or a corrected typo still passes, a paraphrase doesn't. On our test cases genuine copies and paraphrases now land far enough apart that the threshold between them isn't balanced on a knife edge.
To test this properly we needed a model that fabricates quotes on demand, and real models only do it occasionally. So the test suite has a stand-in model whose whole job is to produce the bad output: invented quotes, quotes attached to the wrong passage, sentences with no marker.
Two places to say nothing
There's a gate before the model too. Retrieval is hybrid (dense and sparse vectors, fused by rank, then a cross-encoder re-ranks the top few dozen). If the best passage still scores below a threshold, DocQA declines without calling the model at all. Asking a model to answer from passages that don't contain the answer is how you get a fluent invention.
Our evaluation set deliberately includes questions with no answer in the corpus, and for those the only passing response is to decline.
Permissions go inside the query
If a user isn't cleared for a document, the filter is applied inside the vector search. We considered filtering afterwards and rejected it: by then the passage is already in the model's context, and fragments of it can leak into the answer text even if the UI hides the source.
An audit log you can check
Every question, answer and admin action goes into an append-only log where each row's hash covers its own content and the previous row's hash. Edit or delete a row and the chain breaks from that point; the verification endpoint names the first bad row. One thing that bit us: timestamps have to be set explicitly and normalised to UTC before hashing, or the chain fails to verify against itself.
What we'd tell someone building one
Put the source metadata out of the model's reach. Build the test set with unanswerable questions in it from day one. And write down the tests you fail. Ours needs a GPU for conversational response times, and we found that in our own acceptance run.
The longer, less technical version is on our site: https://www.2sdtechnologies.com/insights/private-document-qa-2026/
Top comments (0)