DEV Community

Cover image for From JSON to Word: Building a Data-Driven Document Pipeline in C#
Chloe
Chloe

Posted on

From JSON to Word: Building a Data-Driven Document Pipeline in C#

Almost every C# application eventually needs to turn structured data into a Word document — an order confirmation, a nightly report, a generated letter. The data itself is rarely the hard part; it's usually already sitting there as JSON, from an API response or a database query. The hard part is turning that JSON into a document that looks right every time, without the code becoming difficult to maintain when the schema changes or additional document types are introduced.

It's tempting to treat this as a small problem: deserialize, loop over some fields, write them into a document, done. That works fine as a one-off script. This article looks at what a slightly more deliberate approach looks like — a small C# pipeline that turns a JSON payload into a maintainable document generation workflow, with Word rendering as the final output step.

1. Why the Naive Approach Breaks Down

The most direct way to solve this problem looks something like: deserialize the JSON data and immediately mix data processing with document generation in one long method — pulling values straight out of the model and calling the Word library's API in the same method that also handles calculations, formatting decisions, and data preparation. For a fixed, simple document, this works — until one of a few predictable things happens:

  • The JSON schema changes. A field gets renamed or a new optional one is added upstream. Nothing enforces the shape of the data at compile time, so these changes surface as runtime exceptions or silently missing content.
  • A second document type is needed. Now there are two long methods with overlapping logic — date formatting, table building, heading styles — copy-pasted and slowly diverging.
  • Business logic and rendering logic are tangled together. Calculating a total or deciding whether to show a discount line has nothing to do with Word specifically, but it's interleaved with document API calls, so testing it means instantiating the Word engine just to check a number.

None of this comes from JSON or from any particular library. It comes from collapsing "understanding the data," "deciding what the document should contain," and "producing the file" into a single method — so every change has to fight through all three at once. The next section lays out what a lighter separation between those three responsibilities looks like.

2. Designing the Pipeline: Three Layers

The fix isn't to build something elaborate — it's to give each of the three responsibilities from the last section its own place to live:

  • Data layer. Takes the raw JSON and turns it into strongly typed C# objects. Its only job is to answer "what data do we have, and is it shaped the way we expect?"
  • Transformation layer. Takes those objects and prepares them for rendering — performing calculations, reshaping nested structures, and applying business rules that determine what data should be included or how it should be structured. This is where business logic lives, separate from any Word-specific API calls.
  • Rendering layer. Takes the prepared data and calls the document library to actually produce the .docx file. This is the only layer that knows anything about Word specifically.

Data flows one direction through these three layers: JSON in, prepared data through the middle, a Word file out. Each layer only needs to know about the shape of its own input and output — the data layer doesn't know Word exists, and the rendering layer doesn't know or care where the data originally came from.

The value of this split shows up the moment something changes. A new field in the JSON schema is a data layer change. A new business rule about when to show a line item is a transformation layer change. A tweak to table styling is a rendering layer change. None of these require touching the other two layers, and each one is small enough to reason about — and test — on its own. That's the whole goal here: not a framework, just enough separation that the code doesn't have to be read top to bottom to make a safe change.

With the shape of the pipeline settled, the rest of this article walks through each layer in turn, starting with the data layer.

3. Data Layer: From JSON to Strongly Typed Models

The data layer's job is narrow: take a JSON payload and produce C# objects that the rest of the pipeline can work with confidently. That means resisting the pull to skip this step and operate directly on a JsonDocument or a dynamic object further downstream.

Working with loosely typed JSON representations can feel faster at first — no model classes to write, no mapping to keep in sync. But it pushes every assumption about the data's shape into whichever code happens to touch it first, and those assumptions are invisible until something is missing at runtime. A strongly typed model makes the expected shape explicit and lets the compiler catch a renamed or missing property before the code ever runs against real data.

Take a simple order confirmation payload as an example:

{
  "orderId": "ORD-10492",
  "customerName": "Jordan Ellis",
  "orderDate": "2026-09-10",
  "items": [
    { "name": "Wireless Mouse", "quantity": 2, "unitPrice": 24.99 },
    { "name": "USB-C Hub", "quantity": 1, "unitPrice": 39.50 }
  ],
  "totalAmount": 89.48
}
Enter fullscreen mode Exit fullscreen mode

The corresponding C# models are straightforward:

public class OrderConfirmation
{
    public string OrderId { get; set; }
    public string CustomerName { get; set; }
    public DateTime OrderDate { get; set; }
    public List<OrderItem> Items { get; set; }
    public decimal TotalAmount { get; set; }
}

public class OrderItem
{
    public string Name { get; set; }
    public int Quantity { get; set; }
    public decimal UnitPrice { get; set; }
}
Enter fullscreen mode Exit fullscreen mode

Deserializing with System.Text.Json is a single call:

var order = JsonSerializer.Deserialize<OrderConfirmation>(
    jsonPayload,
    new JsonSerializerOptions { PropertyNameCaseInsensitive = true }
);
Enter fullscreen mode Exit fullscreen mode

At this point, order is a fully typed object — no string keys, no casting, no risk of a typo in a field name slipping through unnoticed. Nested arrays like Items come through as regular C# collections, which the transformation layer in the next section can work with directly. This is a small step on its own, but it's what lets every layer after it assume the data is already in the shape it expects, rather than re-checking that assumption everywhere.

4. Transformation Layer: Preparing Data for Rendering

It's tempting, once you have a clean typed model, to start designing a generic document abstraction with reusable sections and tables. For a document-generation framework, that's a reasonable direction. For a single JSON-to-Word pipeline, it's usually more structure than the problem needs, and it risks turning an article about a pipeline into an article about designing a document engine.

What the transformation layer actually needs to do is smaller and more concrete: take the typed model from the data layer and shape it into whatever form is easiest for the rendering layer to loop over — nothing more. Three things typically happen here:

  • Deriving values the JSON didn't include directly, like a line total per item.
  • Aggregating or filtering data based on business rules — totals, excluded items, conditional content.
  • Flattening or reshaping nested data into a form the rendering loop can consume without extra branching.

Formatting — how a date or a currency amount actually gets displayed — stays out of this layer on purpose. It's presentation, not business logic, so it belongs in the rendering layer, next to the rest of the display decisions.

For the order confirmation example, that looks like a small view model and a method that builds it:

public class OrderConfirmationView
{
    public string OrderId { get; set; }
    public string CustomerName { get; set; }
    public DateTime OrderDate { get; set; }
    public List<OrderLineView> Lines { get; set; }
    public decimal TotalAmount { get; set; }
}

public class OrderLineView
{
    public string Name { get; set; }
    public int Quantity { get; set; }
    public decimal UnitPrice { get; set; }
    public decimal LineTotal { get; set; }
}

public static class OrderConfirmationTransformer
{
    public static OrderConfirmationView Prepare(OrderConfirmation order)
    {
        return new OrderConfirmationView
        {
            OrderId = order.OrderId,
            CustomerName = order.CustomerName,
            OrderDate = order.OrderDate,
            Lines = order.Items.Select(item => new OrderLineView
            {
                Name = item.Name,
                Quantity = item.Quantity,
                UnitPrice = item.UnitPrice,
                LineTotal = item.Quantity * item.UnitPrice
            }).ToList(),
            TotalAmount = order.TotalAmount
        };
    }
}
Enter fullscreen mode Exit fullscreen mode

Nothing here knows anything about Word, and nothing here decides how a number or a date should look on the page — LineTotal is a computed decimal, not a formatted string. OrderConfirmationView could just as easily feed an HTML template or a PDF generator; it's a plain data shape carrying business-ready values, not a document structure and not a display format. That's the real benefit of keeping this layer small: it stays a translation step, not a second architecture bolted onto the first one. The rendering layer, next, decides how each of these values actually gets displayed.

5. Rendering Layer: Generating the Word Document with Spire.Doc

The rendering layer is the only place that knows Word exists. Its input is the prepared view model from the transformation layer; its output is a .docx file. This article uses Spire.Doc for .NET to handle that last step, but the code here maps to the same handful of operations any Word library exposes: create a document, add paragraphs, add a table, save the file. Since formatting decisions were deliberately kept out of the transformation layer, they show up here instead — this is where a date or an amount actually becomes display text.

using Spire.Doc;
using Spire.Doc.Documents;

public static class OrderConfirmationRenderer
{
    public static void Render(OrderConfirmationView view, string outputPath)
    {
        var document = new Document();
        var section = document.AddSection();

        var title = section.AddParagraph();
        title.AppendText($"Order Confirmation — {view.OrderId}");
        title.Format.HorizontalAlignment = HorizontalAlignment.Center;

        var meta = section.AddParagraph();
        meta.AppendText($"Customer: {view.CustomerName}");
        meta.AppendText($"    Date: {view.OrderDate:MMMM d, yyyy}");

        section.AddParagraph(); // spacing

        var table = section.AddTable(true);
        table.ResetCells(view.Lines.Count + 1, 4);

        var header = table.Rows[0];
        string[] headers = { "Item", "Qty", "Unit Price", "Line Total" };
        for (int i = 0; i < headers.Length; i++)
        {
            header.Cells[i].AddParagraph().AppendText(headers[i]);
        }

        for (int i = 0; i < view.Lines.Count; i++)
        {
            var line = view.Lines[i];
            var row = table.Rows[i + 1];
            row.Cells[0].AddParagraph().AppendText(line.Name);
            row.Cells[1].AddParagraph().AppendText(line.Quantity.ToString());
            row.Cells[2].AddParagraph().AppendText(line.UnitPrice.ToString("C"));
            row.Cells[3].AddParagraph().AppendText(line.LineTotal.ToString("C"));
        }

        var totalParagraph = section.AddParagraph();
        totalParagraph.AppendText($"Total: {view.TotalAmount:C}");
        totalParagraph.Format.HorizontalAlignment = HorizontalAlignment.Right;

        document.SaveToFile(outputPath, FileFormat.Docx2013);
    }
}
Enter fullscreen mode Exit fullscreen mode

The method reads almost like a direct translation of the view model into document elements — one row per line item, one paragraph per piece of metadata — because that's exactly what a rendering layer should be: a straightforward walk over already-prepared data, with no formatting decisions or business logic left to make along the way. Those were resolved in the transformation layer; this layer's job is turning prepared data into document elements, including the formatting and layout decisions required by the target output format.

Swapping the document library later — to Open XML SDK, or anything else — means rewriting this one method. Nothing in the data or transformation layers would need to change.

Project structure after splitting into three layers

6. A Few Things That Show Up in Real Projects

Two issues tend to surface once this pipeline moves from a sample into an actual service, and both are worth addressing at the layer where they belong rather than patching them into the rendering code.

Validating the data before it reaches the transformation layer. Deserialization succeeding doesn't mean the data is usable — a missing CustomerName or an empty Items array will deserialize fine and then produce a broken or empty document three layers later. It's worth validating the model right after deserialization, in the data layer, and failing early with a clear error rather than letting a malformed document quietly reach whoever's expecting it:

if (order.Items == null || order.Items.Count == 0)
{
    throw new InvalidOperationException(
        $"Order {order.OrderId} has no line items and cannot be rendered.");
}
Enter fullscreen mode Exit fullscreen mode

Testing the transformation layer without touching Word at all. Because OrderConfirmationTransformer.Prepare() takes a plain object and returns a plain object, it can be unit tested directly — no document engine, and no display formatting, involved:

[Fact]
public void Prepare_CalculatesLineTotalCorrectly()
{
    var order = new OrderConfirmation
    {
        Items = new List<OrderItem>
        {
            new OrderItem { Name = "Widget", Quantity = 3, UnitPrice = 10.00m }
        }
    };

    var view = OrderConfirmationTransformer.Prepare(order);

    Assert.Equal(30.00m, view.Lines[0].LineTotal);
}
Enter fullscreen mode Exit fullscreen mode

The assertion checks the computed decimal, not a formatted string — the test verifies the calculation, not how a particular machine's culture settings happen to render currency. This is the practical payoff of keeping formatting out of the transformation layer: the part of the pipeline most likely to contain bugs — calculations, aggregation, conditional logic — is also the part that's cheapest and most reliable to test.

7. Putting It Together: End to End

With all three layers in place, generating a document becomes a short pipeline call from a JSON payload to a saved .docx file:

string json = File.ReadAllText("order.json");

var order = JsonSerializer.Deserialize<OrderConfirmation>(
    json, new JsonSerializerOptions { PropertyNameCaseInsensitive = true });

if (order.Items == null || order.Items.Count == 0)
    throw new InvalidOperationException($"Order {order.OrderId} has no line items.");

var view = OrderConfirmationTransformer.Prepare(order);

OrderConfirmationRenderer.Render(view, "order-confirmation.docx");
Enter fullscreen mode Exit fullscreen mode

Given the sample JSON from Section 4, this produces a document with a centered title, the customer name and order date, a four-column table listing each item with its quantity, unit price, and line total, and a right-aligned total at the bottom — the same document a hand-written script could produce, but built from three pieces that can each be changed, tested, or replaced on their own. The benefit of this structure is not fewer lines of code, but clearer boundaries between responsibilities.

Generated order confirmation document

8. Beyond Word

The same three layers carry over cleanly if the output ever needs to change:

JSON → [ Data Layer ] → [ Transformation Layer ] → [ Rendering Layer ] → Word / PDF / HTML

Everything to the left of the rendering layer stays the same regardless of output format, because the data and transformation layers never referenced Word in the first place. A PDF or HTML version of the same document is a new rendering layer, not a redesign of the pipeline. The same reasoning applies to exposing this as an endpoint that returns a generated document on request: the pipeline itself doesn't change, only what calls it.

Conclusion

The pipeline here is small on purpose: a data layer that trusts nothing until it's typed and validated, a transformation layer that handles calculations and business logic without touching a document API, and a rendering layer that does nothing but translate prepared data into Word calls. None of the three pieces is difficult by itself — what they buy is a codebase where the next change, whatever it turns out to be, has an obvious place to go.

Spire.Doc did the actual file-writing in this article, but that's the one piece of this pipeline that's genuinely swappable. The layering is the part worth keeping.

Top comments (0)