DEV Community

Cover image for Building an Intelligent Document Processing Pipeline with MongoDB Atlas
MongoDB Guests for MongoDB

Posted on

Building an Intelligent Document Processing Pipeline with MongoDB Atlas

This tutorial was written by Matteo Rossi.

Imagine a supplier changes the layout of an invoice. The filename still says "invoice," so the file goes to the right parser, but the total has moved to another part of the page. The parser puts a number in the wrong field. Later, someone looking at the saved result has no easy way to see where that number came from.

I want the next system to see where each answer came from, not just the answer itself. The pipeline reads the document to decide what type it is and which values to extract. When a model fills in a field, we keep the original text and its position on the page, along with the model used and the source of any score.
MongoDB Atlas lets us store that information together. For files with a fixed, machine-readable format, a rule-based parser is still the better choice.

Why Keep the Evidence?

When a model says, "This is an invoice," it might be wrong. We can compare it with known invoices and check how close the next possible type is to it. The same goes for "the total is 1525.00": we need to see the text it came from and where it appears on the page. An OCR score can help, but only if we know what it measures.

If we save only the value, we can't tell a clear scan from a poor one. The person reviewing the result doesn't know where to look, and the team can't easily choose which documents to check first.

The filename, folder, and sender address can still help, but they don't tell us what is on the page. If people change the way they name or send files, we need to check the content.

Templates still make sense when you have a few formats that rarely change. A rule-based parser can be cheaper and easier to maintain, although it still needs checks: a layout may match a rule and put a value in the wrong field.

Where is The Risk

Think of a common setup. Optical character recognition (OCR) reads the scan, a rule based on the filename picks a parser, and a template finds the fields on the page. The result goes into database tables. This works well for layouts that stay the same, but it gets harder when you have many formats or when they change without warning.

  • The router trusts the filename. A supplier can change the way it names files without telling you. A filename rule can't check whether the file contains the right type of document, so a bad match may reach the parser without an error. You can measure how often the rule is right if you have documents with known types, but the rule itself cannot say, "I'm not sure." A classifier that reads the content can send uncertain cases to review.
  • Different types need different fields. Invoices, delivery notes, and contracts have little in common beyond basic file information. One wide table leaves many empty columns. Separate tables can work, but they make queries across document types harder. A relational database can still store the source and score for each value. A document database simply lets us keep them next to the value, while each type has its own set of fields.
  • The write step loses useful details. If extraction returns a score or a place on the page, but we save only the value, we lose those details. Still, a score given by a generative model is not always a reliable measure of how often it is right. Saving it doesn't make it trustworthy.

For each field filled in by a model, keep the value and the place it came from. If there is a score, record how it was made, along with the model version. If you don't have a score you can trust, leave it empty and use other checks to decide whether someone should review the field.

Keeping the Evidence with the Document

The page images can stay in object storage. In Atlas, I'd keep the OCR text, document type, extracted values, and review status together, with a link to the original file.

Start with the document. One documents collection holds a record for each file. The example leaves out the OCR text for each part of the page to keep it short. In a real system, you need that text and its position so a reviewer can check each value against the scan. The fields in extraction change with the document type.

{
  "_id": "6f1c9a2e",
  "source": { "uri": "s3://ingest/2026-09-14/INV-88213.pdf", "receivedAt": "2026-09-14T02:11:04Z" },
  "status": "needs_review",
  "classification": { "type": "invoice", "score": 0.81, "margin": 0.34, "exemplarSetVersion": 7 },
  "extraction": {
    "invoiceTotal": {
      "value": "1525.00",
      "confidence": { "score": 0.93, "source": "ocr", "method": "token-aggregation" },
      "provenance": { "page": 1, "boundingBox": [0.61, 0.72, 0.84, 0.78], "text": "Total EUR 1.525,00" }
    },
    "vatAmount": {
      "value": "275.00",
      "confidence": { "score": 0.46, "source": "model", "method": "self-reported-uncalibrated" },
      "provenance": { "page": 1, "boundingBox": [0.61, 0.66, 0.84, 0.71], "text": "VAT 22% 275,00" }
    }
  },
  "review": {
    "required": true,
    "priority": 87,
    "reasons": [
      { "code": "LOW_FIELD_CONFIDENCE", "field": "vatAmount" },
      { "code": "HIGH_VALUE_DOCUMENT" }
    ]
  },
  "pipeline": {
    "embeddingModel": "voyage-4",
    "extractionModel": "<model-id>",
    "contractVersion": 3,
    "classifiedAt": "2026-09-14T02:11:09Z"
  }
}
Enter fullscreen mode Exit fullscreen mode

The two scores in the example mean different things. The OCR score of 0.93 tells us how clearly the system reads the text; it doesn't tell us if the value belongs in the right field. The 0.46 comes from the model itself and hasn't been tested against real results. Keeping the source and method with each score helps us use them in different ways. Sometimes, a check against a business rule and a link to the page tell us more than either score.

The page number, the box around the text, and the original words let a reviewer find the value on the scan. We have to keep that link from the OCR step onward. If we pass only plain text to the extraction step, we can't recover the page position later.

One important thing is to decide what needs to be reviewed when you save the result. One score limit for the whole document treats every field the same. A doubtful note may matter less than a total of 10,000 that looks right but isn't. The worker checks the fields and saves their decision in review, so the queue can show the most important cases first.

It’s important to try finding the type by example. The worker turns the first page into a vector, then uses Atlas Vector Search to find similar files in exemplars, a collection labelled by people. It groups the closest results by type instead of relying on the filename. You still need to test this with your own files: the first page won't always be enough.

db.exemplars.aggregate([
  { $vectorSearch: {
      index: "exemplar_vectors",
      path: "embedding",
      queryVector: firstPageEmbedding,
      numCandidates: 200,
      limit: 25
  }},
  { $project: { label: 1, score: { $meta: "vectorSearchScore" } } },
  { $group: {
      _id: "$label",
      neighbours: { $sum: 1 },
      maxScore: { $max: "$score" },
      avgScore: { $avg: "$score" }
  }},
  { $addFields: { rank: { $add: [
      { $multiply: ["$maxScore", 0.7] },
      { $multiply: ["$avgScore", 0.3] }
  ]}}},
  { $sort: { rank: -1 } },
  { $limit: 2 }
])
Enter fullscreen mode Exit fullscreen mode

Adding up the scores for each type can favour the type with more examples among the top results. The query uses the best score and the average instead, considering that the 0.7 and 0.3 weights are just an example. The query returns the top two types so the worker can compare them. If the best score is too low or the gap is too small, it sends the file to review. These scores are not probabilities, but use a separate set of documents labelled by people to choose the weights and limits, and check what happens when one type has more examples than another.

For this example, use Voyage AI's voyage-4 with its default vector size of 1024 and a cosine vector index. The model can read up to 32,000 tokens, but check the length of each page so no text gets cut off. An invoice and a credit note may still look almost the same. Test such cases with documents kept aside for testing before you try voyage-4-large or add another model to decide between the two.

Atlas can also create the vectors for you. Set the text field and model in an autoEmbed index, then pass query: { text: ... } to $vectorSearch instead of queryVector. Here, the worker creates the vector itself, so it can record the model used for each run. It’s important to highlight that MongoDB currently lists Automated Embedding as a Preview feature.

To add a type, provide examples labelled by people and a list of fields to extract. If the match is weak, send the file to review before extraction; once a person confirms its type, it can join exemplars. Limit the number of examples per type and give each set a version number so you can track the effect of changes. Changing the vector model may mean creating new vectors and a new index. Models in the voyage-4 family produce compatible vectors, but you still need to test their results.

Lastly, define what to extract. The document_types collection holds the expected fields and format for each type. The worker asks the model to follow that format and checks the answer before saving it. When you deploy the system, you also use these definitions to update the database rules that check new records.

// document_types: one entry per type
{
  _id: "invoice",
  version: 3,
  reviewPolicy: { minFieldConfidence: 0.75, highValueThreshold: 10000 },
  extractionSchema: {
    bsonType: "object",
    required: ["invoiceNumber", "invoiceTotal", "issueDate"],
    properties: { /* one entry per field, each a value/confidence/provenance object */ }
  }
}

// deployment step: compile the registered schemas into the collection validator
db.runCommand({
  collMod: "documents",
  validator: { $jsonSchema: {
    bsonType: "object",
    required: ["source", "status", "classification", "pipeline"],
    oneOf: [
      { properties: { classification: { properties: { type: { enum: ["invoice"] } } },
                      extraction: invoiceExtractionSchema } },
      { properties: { classification: { properties: { type: { enum: ["delivery_note"] } } },
                      extraction: deliveryNoteExtractionSchema } }
    ]
  }}
})
Enter fullscreen mode Exit fullscreen mode

This snippet shows the idea, not a command you can run as it is: invoiceExtractionSchema and deliveryNoteExtractionSchema stand for schema objects created during deployment. The full database rule also needs to accept files sent to review before extraction, and to require extraction and classification.type for files that were extracted. MongoDB does not read document_types every time it saves a file. These rules can catch missing fields or the wrong format, but not a total that looks valid and is wrong.

 Keeping a Review Queue

The query below lists files that need a person to check them, starting with the highest priority. If the classifier cannot choose a type, or a result fails a format check, the worker also sets status: "needs_review" and saves the reason in review.reasons.

db.documents.createIndex({ status: 1, "review.required": 1, "review.priority": -1 })

db.documents.aggregate([
  { $match: { status: "needs_review", "review.required": true } },
  { $sort: { "review.priority": -1 } },
  { $limit: 50 },
  { $project: { "source.uri": 1, "classification.type": 1, "review.reasons": 1 } }
])
Enter fullscreen mode Exit fullscreen mode

A worker watches Change Streams and moves each file to the next step. It may receive the same event twice, so it must check the file's current status and pipeline version before making a change. Atlas Triggers can handle simple steps. A separate worker gives you more control when you need to retry failed calls, limit how many files run at once, or keep a queue of failures.

A file usually moves from received to classified, then extracted, and finally to confirmed or needs_review. A file also goes to review if the classifier cannot choose a type or if the result fails a check. In the first case, the reviewer starts by choosing the type.

Sometimes a Parser is Enough

An electronic invoice in a required XML format already names its fields. Read it with a parser and check that it follows the format; a model would add cost and another way to fail. Keep the original file and handle files that don't follow the rules, but you don't need a model to find the total.

The same applies when most files use one familiar layout. Try the parser first and send a file to the model if it fails a check or uses a new layout. Even a parser that reports success can put a value in the wrong field, so keep track of later corrections.

The model is most useful for changing layouts and document types that have no good template. For the rest, start with the parser.

Migration Path

You can add these parts while the old pipeline keeps running. You may still need to change how OCR text, page positions, and review decisions are saved.

  1. Keep the original and the OCR text. Save text and page positions, but leave the old routing in place for now. Reviewers will need both later.
  2. Run the new classifier without using its answers yet. Compare its choices with the old router and check the files where they disagree. Agreement does not mean either one is right.
  3. Collect examples with checked labels. Ask people to confirm the types of files from the old pipeline. Use some as examples and keep the rest separate to test error rates and choose score limits.
  4. Change routing for one type at a time. Start where the old router makes expensive mistakes. Send uncertain files to review unless the old route has passed a check based on the file's content.
  5. Try extraction for types without templates. Compare the results with the original pages before using them for payments or approvals.
  6. Set review rules for each field. A wrong total costs more than a missing optional note. Use real corrections to test the rules, not just model scores.
  7. Check results after launch. Track how often the classifier is right on checked files, how often it asks for review, how many files go through without review, and how many errors people find later. Late corrections show mistakes that the other checks missed.

What Survives is the Pipeline

Go back to the invoice with the moved total. A classifier that reads the page may still call it an invoice. But did the extraction step put the amount in the right field? If we keep the original text and its place on the page, a reviewer can check it when a rule finds something odd. The model can still make mistakes; the difference is that we have a way to find them before another system uses the value.

Top comments (0)