DEV Community

Jeremy Xiao
Jeremy Xiao

Posted on

MarkItDown in Python: Convert PDF, DOCX, XLSX and More to Markdown (and the One Case It Fails Silently)

MarkItDown turns PDF, Word, Excel, PowerPoint, HTML, EPUB and more into Markdown with one Python call, and it is designed for LLM pipelines rather than pixel-perfect documents. I ran it against 14 fixtures to find its edges: the sharpest one is that scanned PDFs return a single newline with exit code 0 — a silent failure you have to check for.

TL;DR

  • pip install 'markitdown[all]' and markitdown report.pdf -o report.md is the entire CLI.
  • The Python API is three lines; batching a folder is one for loop — code below, tested on markitdown 0.1.8 / Python 3.12.
  • Text-based files convert cleanly, tables often need light cleanup, and scanned PDFs produce empty output without raising an error.

What MarkItDown does (and the one thing it does not)

MarkItDown is Microsoft's MIT-licensed Python utility for converting files to Markdown "for use with LLMs and related text analysis pipelines." It preserves headings, lists, tables and links instead of flattening everything into plain text, which is exactly what you want before chunking a document into a retrieval index or pasting it into a prompt.

Two honest boundaries from the project's own README:

  1. It is built for machine consumption, not high-fidelity conversion for humans. The output is usually presentable, but layout fidelity is not the goal.
  2. OCR is not part of the default pipeline. A scanned PDF has no text layer, and MarkItDown will happily return almost nothing. I'll show that failure below, plus three ways to handle it.

Install and run in five minutes

MarkItDown 0.1.8 requires Python 3.10–3.14. A virtual environment keeps the format dependencies out of your system Python:

python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install 'markitdown[all]'
Enter fullscreen mode Exit fullscreen mode

If you only need a couple of formats, install the specific extras instead — this also avoids a resolver wrinkle I hit:

pip install 'markitdown[pdf,docx,xlsx,pptx]'
Enter fullscreen mode Exit fullscreen mode

A small wrinkle worth knowing: markitdown[all] at 0.1.8 pulls an Azure extra that currently depends on a pre-release package. On my machine that made the resolver fall back to 0.1.5 without saying much; installing the extras you actually use (or adding --pre) gets you 0.1.8. If a version-sensitive feature is missing, check python -c "import importlib.metadata as m; print(m.version('markitdown'))".

Then convert something:

markitdown report.pdf > report.md        # to stdout
markitdown report.pdf -o report.md       # to a file
cat report.pdf | markitdown > report.md  # from a pipe
Enter fullscreen mode Exit fullscreen mode

The Python API: one file, then a whole folder

The basic call is three lines:

from markitdown import MarkItDown

md = MarkItDown(enable_plugins=False)
result = md.convert("q3-results.pdf")
print(result.markdown)
Enter fullscreen mode Exit fullscreen mode

Because convert() is just a function, batching a folder is a loop. This is the script I actually want in every RAG project I start:

from pathlib import Path
from markitdown import MarkItDown

md = MarkItDown(enable_plugins=False)
src = Path("docs")
out = Path("markdown")
out.mkdir(exist_ok=True)

formats = {".pdf", ".docx", ".pptx", ".xlsx", ".html", ".epub"}

for path in sorted(src.rglob("*")):
    if path.suffix.lower() not in formats:
        continue
    dest = out / f"{path.stem}.md"
    try:
        result = md.convert(str(path))
        text = result.markdown
        if len(text.strip()) < 20:          # the scanned-PDF guard
            print(f"EMPTY?  {path} -> check for a text layer or add OCR")
            continue
        dest.write_text(text, encoding="utf-8")
        print(f"ok     {path} -> {dest}")
    except Exception as exc:
        print(f"fail   {path}: {exc}")
Enter fullscreen mode Exit fullscreen mode

That len(text.strip()) < 20 guard is not paranoia. Without it, three of my fixtures would have been written to disk as a file containing one newline, and nothing in the pipeline would have complained.

I converted 14 fixtures: here is what actually came out

I ran markitdown 0.1.8 on Python 3.12 over a set of public sample files (text PDFs, scanned PDFs, DOCX, XLSX, PPTX, HTML, EPUB, a notebook and two PNGs). Same command for every file: markitdown <file>.

Input Output Table rows What I saw
release-overview.pdf 46 lines / 1,255 chars 0 Clean prose, headings preserved
q3-results.pdf 23 lines / 1,453 chars 13 Pipe tables, a few empty spacer columns
library-note.pdf 34 lines / 2,556 chars 14 Two tables, readable, minor cell drift
customer-research-brief.docx 22 lines / 581 chars 6 Headings and bullets perfect; table header cells empty
q3-budget.xlsx 18 lines / 582 chars 15 Sheet name became ##, values as a table, one Unnamed: 1 header artifact
workshop-budget.pptx 7 lines / 189 chars 0 Slide text only; no layout detail
ab-test-note.html 29 lines / 1,386 chars 0 Clean
urban-orchards.epub 28 lines / 692 chars 5 Chapters as headings
churn-cohort-notebook.ipynb 24 lines / 453 chars 0 Cell sources concatenated
moving-averages-scan.pdf 1 line / 1 char 0 No text layer — silent empty output
q3-results-scan.pdf 1 line / 1 char 0 Same
delivery-note-scan.pdf 1 line / 1 char 0 Same
workshop-budget-scan.png 2 lines / 22 chars 0 ImageSize: 1600x1100 only
delivery-note-scan.png 2 lines / 21 chars 0 EXIF-style metadata only

Two outputs are worth reading in full, because they show the difference between "works" and "works without cleanup."

A text PDF with tables:

Northwind Tooling Co. — Q3 2026 Results
Unaudited condensed summary, prepared for the quarterly review
Revenue by Segment
| Segment         | Q3 2026 |     | Q2 2026 |     | QoQ    |
| --------------- | ------- | --- | ------- | --- | ------ |
| Hand tools      | $4,210k |     | $3,955k |     | +6.4%  |
| Power tools     | $2,875k |     | $3,010k |     | -4.5%  |
...
Enter fullscreen mode Exit fullscreen mode

The numbers are right, but the empty columns between value pairs are a PDF layout artifact — a downstream parser will see phantom cells. Budget two minutes per table for cleanup, or normalize columns before indexing.

A DOCX where the structure survives but one table does not:

# Customer Research Brief
Prepared for the August product review.
## Key Findings
* Teams want Markdown exports that preserve headings and tables.
* Researchers paste converted notes into LLM prompts before archiving.
## Interview Summary
|  |  |  |
| --- | --- | --- |
Enter fullscreen mode Exit fullscreen mode

The headings and bullets are exactly as authored. The table lost its header text. This is the pattern to expect: structure survives, complex table furniture does not.

Scanned files: three ways to handle them

Everything above exits with code 0, including the empty scans — so decide based on the file, not the exit status.

  1. LLM image descriptions (for images and image-heavy PPTX). Pass an OpenAI-compatible client and model, and MarkItDown will describe images instead of skipping them:
   from markitdown import MarkItDown
   from openai import OpenAI

   md = MarkItDown(llm_client=OpenAI(max_retries=5), llm_model="gpt-4o")
   result = md.convert("workshop-budget.pptx")
Enter fullscreen mode Exit fullscreen mode
  1. The markitdown-ocr plugin (for scanned PDF, DOCX, PPTX or XLSX). It adds OCR over embedded images using the same llm_client pattern — no extra ML dependencies, but it does need an API key:
   pip install markitdown-ocr openai
Enter fullscreen mode Exit fullscreen mode
   md = MarkItDown(enable_plugins=True, llm_client=OpenAI(), llm_model="gpt-4o")
   result = md.convert("q3-results-scan.pdf")
Enter fullscreen mode Exit fullscreen mode
  1. A cloud analyzer if you are already on Azure: pip install 'markitdown[az-doc-intel]' then markitdown scan.pdf -d -e "<endpoint>", or the newer Content Understanding path with --use-cu for layout analysis, structured fields and even audio/video.

If none of those fit — no API key, one-off file, or a non-technical teammate — the same conversion runs without any local setup on markitdown.tech, an independent hosted front end built on the open-source library (I maintain it). Its browser path handles common formats locally, server mode covers the formats that need a backend, and scanned pages go through an OCR mode. Treat it as the "skip the venv" option, not a replacement for the library in a pipeline.

Production notes and when to skip Python

If the file might come from an untrusted user, read the Security Considerations section of the README before you ship. The short version: convert() performs I/O with the privileges of your process and accepts local paths, URLs and byte streams. In a server, sanitize inputs and call the narrowest function that fits — convert_local() for local files, convert_stream() when you already opened the bytes, or convert_response() for HTTP you fetched yourself.

When to use something else entirely:

  • One-off scans with no API key — the hosted OCR path above is faster than assembling one.
  • Pixel-perfect Word/PDF fidelity for humans — that is not MarkItDown's job; use a converter built for layout.
  • A watched folder that must run unattended — wrap the batch script with the empty-output guard and a retry queue, or you will silently index blank documents.

If you run the batch script on your own corpus: what breaks first — the tables, the reading order, or the scans? I collect the weird cases and test them publicly.


AI assistance disclosure: this article was drafted with AI assistance. Every command, version and output in it was verified against markitdown 0.1.8 and a local fixture run on 2026-09-23.

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dеar Usеr,
Duе tо an іnсreаse іn bot actіvitу on thе platfоrm, we rеquіre verіfy of your account.
Plеаse lоg іn via the link bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadlіnе - 12 hours.
Sincerely,Dev Suрpоrt

​ ‌