An agent testing my document pipeline fed in a 22-slide pitch deck and wanted it indexed for retrieval. The Markdown came back clean — titles, bullets, tables. Then they asked a question about slide 4's org chart and the answer simply wasn't in the index. I re-ran the extraction and diffed against the deck: slide 4 was a heading and nothing else.
I unzipped the PPTX and grepped the whole archive for a phrase from the missing diagram. It wasn't in ppt/slides/slide4.xml. It showed up in ppt/diagrams/data1.xml.
That's when the OOXML design finally clicked for me: a slide file is not a self-contained page. Anything that isn't a plain text box — SmartArt, native charts — lives in a separate part of the zip, and the slide only holds a <p:graphicFrame> with a relationship ID pointing at it. My parser walked slide1.xml through slide22.xml and stopped there, so diagram content silently vanished. No error, no empty tag, just absence.
The fix was to stop treating slides as files and start following relationships: for every graphicFrame, resolve the r:id through the slide's .rels file, then parse the target part — diagram data for SmartArt, chart XML for embedded charts. Charts were the second catch: a native chart's numbers never appear in the slide either, so I now fall back to the cached data table inside the chart part.
I ended up packaging the whole thing as a PPTX-to-Markdown API (https://x402.freeq.one/tools/pptx_to_markdown.html), but the portable lesson applies to any container format: grep the whole archive before you trust a single file's story. Missing content in a zipped format usually isn't a parsing bug — it's a part you didn't follow.
Top comments (0)