DEV Community

Cover image for My EPUB converter alphabetized the chapters — the spine, not the zip, defines reading order
InApp
InApp

Posted on Originally published at imapp.blogspot.com

My EPUB converter alphabetized the chapters — the spine, not the zip, defines reading order

A user sent me a 40-chapter technical manual they'd run through my EPUB converter. All the content was there, but the preface sat at position 14, chapter 3 came after chapter 12, and the appendix landed mid-book. Clean text, nonsense order.

The bug was embarrassingly mine. An EPUB is a ZIP of individual XHTML files — one per chapter, usually named chapter1.xhtml, chapter2.xhtml. I processed the entries in alphabetical filename order because it felt deterministic. Lexicographic sort says chapter10.xhtml comes before chapter2.xhtml. And the preface was a one-off file the publisher had named fm1.xhtml, which sorted last. The reading order I emitted was a file-sorting accident, not the book.

The fix: reading order in an EPUB is not a filename convention — it's declared in content.opf. The <spine> element lists manifest IDs in the order the book must be read, and the manifest maps each ID to its href. EPUB2 books also carry a toc.ncx with a nested navMap; EPUB3 replaces it with nav.xhtml. The spine is authoritative for linear reading; the TOC adds hierarchy (parts containing chapters) that I now use to set top-level heading depth.

The current pipeline: parse the OPF, walk the spine idrefs in order, fall back to the NCX navMap for anything the spine omits, and only then stoop to sorted filenames for badly mangled files. Every test book has come out in order since, and chapter 10 finally stays after chapter 2.

Lesson: a container's file listing is storage order, not semantic order. Same trap as trusting object order in a PDF's internal tree. When a format ships a manifest, believe the manifest.

That converter lives at https://x402.freeq.one/tools/epub_to_markdown.html, spine-order fix included.

Top comments (0)