DEV Community

Cover image for Word doesn't store list numbers in the text — my DOCX converter printed clauses as bare paragraphs
InApp
InApp

Posted on Originally published at imapp.blogspot.com

Word doesn't store list numbers in the text — my DOCX converter printed clauses as bare paragraphs

A user fed a 40-page contract into my DOCX converter and complained that every clause came out as a plain paragraph — no bullets, no numbers, just walls of text. The original Word file had nested numbered clauses (1.2.3 style) and my output flattened all of it.

I assumed I was reading malformed XML. I wasn't. OOXML simply doesn't store list markers in the document text. A numbered paragraph looks like:

<w:p><w:pPr><w:numPr><w:ilvl w:val="1"/><w:numId w:val="4"/></w:numPr></w:pPr><w:t>Deliverables</w:t></w:p>

There's no "1." anywhere in the file. Word generates the marker at render time by looking up numId 4 in a separate part of the zip archive, numbering.xml, which maps abstract list definitions to formats — decimal, lowerLetter, romanNumeral, bullet — each with its own start value and indent level. My first version read only document.xml, so every list item printed as bare text. Word itself rendered the file perfectly; nothing was objectively wrong with it.

Fixing it meant rebuilding a small slice of Word's list logic by hand: track counters per numId and level, reset child counters when a parent level increments, respect w:start offsets, and treat numId 0 as "no numbering." Bullets were a trap too — Word defines them as glyph characters from Wingdings or Symbol fonts, so a literal bullet came out as a random letter "l" or "8" until I mapped the common glyphs to plain - markers.

I tested against 30 documents sitting in my own archive. Before the fix, 19 of 30 lost all list structure. After, 28 kept it. The two remaining failures: one document whose numbering.xml pointed at an abstract definition that doesn't exist (Word tolerates this silently; I now fall back to plain paragraphs instead of crashing), and one using lvlRestart rules I still approximate.

The lesson: "clean" conversion isn't about the bytes you can read — it's about the parts a real renderer invents on the fly. I shipped the fixed pipeline as my DOCX-to-Markdown API at https://x402.freeq.one/tools/docx_to_markdown.html, but the takeaway generalizes: an Office file is a zip of loosely synchronized databases, and no single part tells you the truth.

Top comments (0)