A .docx file is a zip archive of XML files. That is why a Word document you send to a client can carry the name of its author, the name of the last person who saved it, every comment, and every tracked change, even when the page you see looks clean.
You can read all of that in the browser with no library. Modern browsers can inflate zip entries natively with DecompressionStream('deflate-raw'), so the only code you need is a small reader for the zip directory.
Step 1: read the zip directory
The directory sits at the end of the file. Find the end-of-central-directory marker, then walk the entries. Each entry gives you a name, a compression method and where the bytes start.
async function inflate(bytes) {
const s = new Blob([bytes]).stream().pipeThrough(new DecompressionStream('deflate-raw'));
return new Uint8Array(await new Response(s).arrayBuffer());
}
function readZip(buf) { // buf: Uint8Array -> Map(name -> () => Promise<string>)
const dv = new DataView(buf.buffer, buf.byteOffset, buf.length);
const td = new TextDecoder();
let e = buf.length - 22;
while (e >= 0 && dv.getUint32(e, true) !== 0x06054b50) e--;
if (e < 0) throw new Error('not a zip file');
let count = dv.getUint16(e + 10, true);
let p = dv.getUint32(e + 16, true);
const files = new Map();
for (; count > 0; count--) {
const method = dv.getUint16(p + 10, true);
const size = dv.getUint32(p + 20, true);
const nameLen = dv.getUint16(p + 28, true);
const extraLen = dv.getUint16(p + 30, true);
const commentLen = dv.getUint16(p + 32, true);
const off = dv.getUint32(p + 42, true);
const name = td.decode(buf.subarray(p + 46, p + 46 + nameLen));
const start = off + 30 + dv.getUint16(off + 26, true) + dv.getUint16(off + 28, true);
const raw = buf.subarray(start, start + size);
files.set(name, async () => td.decode(method === 0 ? raw : await inflate(raw)));
p += 46 + nameLen + extraLen + commentLen;
}
return files;
}
Step 2: know which parts hold what
| Part inside the zip | What it holds |
|---|---|
docProps/core.xml |
author (dc:creator), last saved by (cp:lastModifiedBy), created and modified dates, title |
docProps/app.xml |
company, template name, total editing time |
word/comments.xml |
every comment, with its author |
word/document.xml |
the body, including <w:ins> (inserted text) and <w:del> (deleted text), each with an author and a date |
Deleted text is the one people forget. A tracked deletion is still in the file until someone accepts the change, so the sentence you struck out is readable by anyone who opens the XML.
Step 3: pull the fields out
const tag = (xml, name) => (new RegExp(`<${name}[^>]*>([^<]*)</${name}>`).exec(xml) || [])[1] || '';
const authors = (xml, el) => [...new Set([...xml.matchAll(new RegExp(`<w:${el} [^>]*w:author="([^"]*)"`, 'g'))].map((m) => m[1]))];
async function report(buf) {
const z = readZip(buf);
const get = async (n) => (z.has(n) ? z.get(n)() : '');
const [core, app, comments, doc] = await Promise.all(
['docProps/core.xml', 'docProps/app.xml', 'word/comments.xml', 'word/document.xml'].map(get));
return {
author: tag(core, 'dc:creator'),
lastModifiedBy: tag(core, 'cp:lastModifiedBy'),
created: tag(core, 'dcterms:created'),
modified: tag(core, 'dcterms:modified'),
company: tag(app, 'Company'),
commentAuthors: authors(comments, 'comment'),
insertionAuthors: authors(doc, 'ins'),
deletionAuthors: authors(doc, 'del'),
};
}
In a page, feed it a dropped file: report(new Uint8Array(await file.arrayBuffer())). The file never leaves the tab, so you can run it on a draft you are not ready to share.
What this does not do
- It uses regular expressions on XML, which is fine for a quick look and wrong for edge cases. Names with
&stay escaped, and an unusual namespace prefix will be missed. For anything you rely on, useDOMParser. - It does not read zip64 archives (files over 4 GB), and a password-protected
.docxis not a zip at all, soreadZipthrows on it. - It only looks at four parts. Headers and footers, embedded images (which can carry their own EXIF data), custom XML, document-level hyperlinks to local file paths and
docProps/custom.xmlcan also hold names and paths. - It only reads. Removing the data safely is a different job: you have to rewrite the zip and keep the package valid, or Word will report the file as corrupt.
If you only want the answer
I build Doc Scrubber, a free tool that does this and then writes a clean copy, in your browser with nothing uploaded. It reports what a Word, Excel, PowerPoint or PDF file holds before you remove anything, and it lists what stays in the file afterwards. The free tier handles 5 files at a time: https://hexloomlabs.com/doc-scrub/?ref=devto-0710-docx
Top comments (0)