DEV Community

Cover image for Every vector in your Word document just became a hat
smrifat1411
smrifat1411

Posted on Originally published at smrifat1411.github.io

Every vector in your Word document just became a hat

I maintain a small library that cleans up Microsoft Word paste in rich-text editors. Part of that job is converting Word's equations to LaTeX, so they stay editable instead of turning into screenshots. For months it had a bug that never threw, never logged, and never produced invalid output. It just quietly rewrote physics.

Every accent came out as \hat.

A vector F⃗ became F̂. A time derivative ẋ became x̂. A mean x̄ became x̂. Newton's second law pasted as \hat{F} = m\hat{a}, which renders perfectly and means something else. The equation looked fine at a glance, the LaTeX compiled, and the tests passed, because the tests only checked that an accent came through.

The cause was one skipped attribute. Here is how Word writes an accent:

<m:acc>
  <m:accPr><m:chr m:val="⃗"/></m:accPr>
  <m:e><m:r><m:t>F</m:t></m:r></m:e>
</m:acc>
Enter fullscreen mode Exit fullscreen mode

The m:chr holds the accent as a Unicode combining character. U+20D7 is the vector arrow, U+0304 the bar, U+0307 the dot. Word omits the whole accPr block when the accent is a circumflex, so \hat is the right default when it is missing. I had written the default and never read the attribute. The fix is a twelve-entry lookup table. The lesson is that silent-but-plausible is the worst failure mode a converter can have, and the only test that catches it is one that asserts the exact output for each symbol.

What OMML is, and why your pipeline is dropping it

Word does not store equations as images, and it does not store them as LaTeX. It stores them as OMML, Office Math Markup Language. The same XML sits in two places: inside every .docx at word/document.xml, and on the clipboard when you copy from Word.

Most tools that read a .docx ignore it. mammoth, the usual docx-to-HTML converters, and most of the document ingestion code feeding RAG systems walk the paragraph runs and skip <m:oMath> entirely. If you have ever extracted text from a technical paper and found the equations missing, that is where they went. Pandoc handles it, but Pandoc is a Haskell binary, which is a hard sell inside a browser or a serverless function.

So the converter in my library is exported on its own. String in, string out, no dependencies:

import { ommlToLatex } from 'wordpaste';

ommlToLatex(
  '<m:oMath><m:f><m:num><m:r>a</m:r></m:num><m:den><m:r>b</m:r></m:den></m:f></m:oMath>',
);
// '\\frac{a}{b}'
Enter fullscreen mode Exit fullscreen mode

It returns an empty string when the fragment cannot be parsed, so it never throws halfway through a document.

Getting the equations out of a .docx

A .docx is a zip file. The equations are in word/document.xml. That is the whole secret.

unzip -p paper.docx word/document.xml > document.xml
Enter fullscreen mode Exit fullscreen mode
import { readFileSync } from 'node:fs';
import { JSDOM } from 'jsdom';
import { ommlToLatex } from 'wordpaste';

// Node has no DOMParser; this is the only setup.
globalThis.DOMParser = new JSDOM().window.DOMParser;

const xml = readFileSync('document.xml', 'utf8');
for (const m of xml.match(/<m:oMath>[\s\S]*?<\/m:oMath>/g) ?? []) {
  console.log(ommlToLatex(m));
}
Enter fullscreen mode Exit fullscreen mode

On a document containing Newton's second law and a binomial coefficient, that prints:

\vec{F}=m\vec{a}
\binom{n}{k}
Enter fullscreen mode Exit fullscreen mode

In the browser you skip the jsdom line, since DOMParser is already there. Nothing is uploaded anywhere.

What it covers

Fractions, including \binom (Word writes C(n,k) as a fraction with the bar turned off, which was a second silent bug), sub- and superscripts, pre-scripts for isotopes like {}^{14}_{6}C, radicals, delimiters including ‖ ⌊ ⌋ ⟨ ⟩ and the empty-left-side form Word uses for cases blocks, sums, integrals and products with limits, functions, over- and underlines, accents (now the right ones), matrices, and multi-line equation arrays as aligned.

What it does not do: MathML. LibreOffice puts MathML on the clipboard, not OMML, and that is a different parser.

There is a live converter and the full construct table at smrifat1411.github.io/wordpaste/omml-to-latex. Paste some OMML in, get LaTeX out. The source is on GitHub, MIT, 3.6 kB.

If you have a .docx that converts wrong, open an issue with the <m:oMath> fragment. After the hat incident I would rather hear about it.

Top comments (0)