DEV Community

Hexloom Labs
Hexloom Labs

Posted on

Read the hidden author, comments and tracked changes in a .docx with plain JavaScript

A .docx file is a zip archive of XML files. That is why a Word document you send to a client can carry the name of its author, the name of the last person who saved it, every comment, and every tracked change, even when the page you see looks clean.

You can read all of that in the browser with no library. Modern browsers can inflate zip entries natively with DecompressionStream('deflate-raw'), so the only code you need is a small reader for the zip directory.

Step 1: read the zip directory

The directory sits at the end of the file. Find the end-of-central-directory marker, then walk the entries. Each entry gives you a name, a compression method and where the bytes start.

async function inflate(bytes) {
  const s = new Blob([bytes]).stream().pipeThrough(new DecompressionStream('deflate-raw'));
  return new Uint8Array(await new Response(s).arrayBuffer());
}

function readZip(buf) { // buf: Uint8Array -> Map(name -> () => Promise<string>)
  const dv = new DataView(buf.buffer, buf.byteOffset, buf.length);
  const td = new TextDecoder();
  let e = buf.length - 22;
  while (e >= 0 && dv.getUint32(e, true) !== 0x06054b50) e--;
  if (e < 0) throw new Error('not a zip file');
  let count = dv.getUint16(e + 10, true);
  let p = dv.getUint32(e + 16, true);
  const files = new Map();
  for (; count > 0; count--) {
    const method = dv.getUint16(p + 10, true);
    const size = dv.getUint32(p + 20, true);
    const nameLen = dv.getUint16(p + 28, true);
    const extraLen = dv.getUint16(p + 30, true);
    const commentLen = dv.getUint16(p + 32, true);
    const off = dv.getUint32(p + 42, true);
    const name = td.decode(buf.subarray(p + 46, p + 46 + nameLen));
    const start = off + 30 + dv.getUint16(off + 26, true) + dv.getUint16(off + 28, true);
    const raw = buf.subarray(start, start + size);
    files.set(name, async () => td.decode(method === 0 ? raw : await inflate(raw)));
    p += 46 + nameLen + extraLen + commentLen;
  }
  return files;
}
Enter fullscreen mode Exit fullscreen mode

Step 2: know which parts hold what

Part inside the zip What it holds
docProps/core.xml author (dc:creator), last saved by (cp:lastModifiedBy), created and modified dates, title
docProps/app.xml company, template name, total editing time
word/comments.xml every comment, with its author
word/document.xml the body, including <w:ins> (inserted text) and <w:del> (deleted text), each with an author and a date

Deleted text is the one people forget. A tracked deletion is still in the file until someone accepts the change, so the sentence you struck out is readable by anyone who opens the XML.

Step 3: pull the fields out

const tag = (xml, name) => (new RegExp(`<${name}[^>]*>([^<]*)</${name}>`).exec(xml) || [])[1] || '';
const authors = (xml, el) => [...new Set([...xml.matchAll(new RegExp(`<w:${el} [^>]*w:author="([^"]*)"`, 'g'))].map((m) => m[1]))];

async function report(buf) {
  const z = readZip(buf);
  const get = async (n) => (z.has(n) ? z.get(n)() : '');
  const [core, app, comments, doc] = await Promise.all(
    ['docProps/core.xml', 'docProps/app.xml', 'word/comments.xml', 'word/document.xml'].map(get));
  return {
    author: tag(core, 'dc:creator'),
    lastModifiedBy: tag(core, 'cp:lastModifiedBy'),
    created: tag(core, 'dcterms:created'),
    modified: tag(core, 'dcterms:modified'),
    company: tag(app, 'Company'),
    commentAuthors: authors(comments, 'comment'),
    insertionAuthors: authors(doc, 'ins'),
    deletionAuthors: authors(doc, 'del'),
  };
}
Enter fullscreen mode Exit fullscreen mode

In a page, feed it a dropped file: report(new Uint8Array(await file.arrayBuffer())). The file never leaves the tab, so you can run it on a draft you are not ready to share.

What this does not do

  • It uses regular expressions on XML, which is fine for a quick look and wrong for edge cases. Names with &amp; stay escaped, and an unusual namespace prefix will be missed. For anything you rely on, use DOMParser.
  • It does not read zip64 archives (files over 4 GB), and a password-protected .docx is not a zip at all, so readZip throws on it.
  • It only looks at four parts. Headers and footers, embedded images (which can carry their own EXIF data), custom XML, document-level hyperlinks to local file paths and docProps/custom.xml can also hold names and paths.
  • It only reads. Removing the data safely is a different job: you have to rewrite the zip and keep the package valid, or Word will report the file as corrupt.

If you only want the answer

I build Doc Scrubber, a free tool that does this and then writes a clean copy, in your browser with nothing uploaded. It reports what a Word, Excel, PowerPoint or PDF file holds before you remove anything, and it lists what stays in the file afterwards. The free tier handles 5 files at a time: https://hexloomlabs.com/doc-scrub/?ref=devto-0710-docx

Top comments (0)