At the end of September a Nebraska TV station reported that a data center operator's state water and electricity filing had its figures "redacted" with a box that copy and paste undid: select the box, paste it into another document, and the numbers appear. Word files fail in a similar way, and in a few more.
A common way to "redact" a Word document is to select the sensitive text and give it a black highlight. It looks right on screen and in a PDF export. The text is still in the file.
A .docx is a zip archive, and the body lives in word/document.xml. Highlighting does not remove text from that file; it only adds formatting next to it. Anyone can select the black bar, copy it and paste it into a plain-text editor, or change the font colour and read it. Four things routinely survive when people think they have redacted a Word file:
-
Black on black. A run with
<w:highlight w:val="black"/>or run shading withw:fill="000000", plus a black font colour, shows as a bar. The<w:t>element with the words is still there. (With the font colour left on automatic, Word turns the text white on a dark background, so the text stays readable: only the black-on-black combination looks redacted.) -
Tracked deletions. Text deleted with Track Changes is wrapped in
<w:del>with a<w:delText>element. It disappears from the page view but stays in the file until someone accepts the change. -
Hidden text. A run with
<w:vanish/>in its properties (Font > Hidden) is not shown or printed but is in the XML. - A black rectangle drawn over the text. The shape is a separate drawing element anchored to the paragraph; the text under it is untouched.
A small check
This script finds the first three. It reads only word/document.xml, so it ignores headers, footers, footnotes and comments, and it does not look at white-on-white text, paragraph or table-cell shading, or shapes. Treat it as a way to see the problem, not as a release gate. I wrote this article with AI assistance; the output shown below is from running the script on the test files described.
import html, re, sys, zipfile
def text_of(xml):
return html.unescape("".join(re.findall(r"<w:(?:t|delText)(?: [^>]*)?>(.*?)</w:(?:t|delText)>", xml, re.S)))
def runs(xml):
# every <w:r>...</w:r> with its properties and visible/deleted text
for m in re.finditer(r"<w:r[ >].*?</w:r>", xml, re.S):
r = m.group(0)
props = re.search(r"<w:rPr>(.*?)</w:rPr>", r, re.S)
yield (props.group(1) if props else ""), text_of(r)
def check(path):
xml = zipfile.ZipFile(path).read("word/document.xml").decode("utf8")
found = []
for props, text in runs(xml):
if not text.strip():
continue
black_bg = 'w:highlight w:val="black"' in props or re.search(r'<w:shd [^>]*w:fill="000000"', props)
black_fg = re.search(r'<w:color w:val="000000"', props)
if black_bg and black_fg:
found.append(("black on black", text))
elif "<w:vanish" in props:
found.append(("hidden text", text))
for m in re.finditer(r"<w:del [^>]*>(.*?)</w:del>", xml, re.S):
found.append(("tracked deletion", text_of(m.group(1))))
return found
if __name__ == "__main__":
for kind, text in check(sys.argv[1]):
print(f"{kind}: {text!r}")
Run it as python docx_redaction_check.py report.docx. On a test file with one of each failure it prints:
black on black: 'SECRET-ALPHA'
black on black: 'SECRET-BETA'
hidden text: 'SECRET-HIDDEN'
tracked deletion: 'SECRET-DEL'
and on a file with only ordinary text and a yellow highlight it prints nothing.
Fixing it properly
Formatting can be undone, so the fix is to remove the text itself, not to cover it:
- Delete the sensitive words (or replace them with a neutral marker such as
[REDACTED]) while Track Changes is off, or accept or reject every tracked change first. - Remove hidden text, and check headers, footers, footnotes and comments too.
- Save a new copy and run a check on that copy, not on the original. For anything that must stay confidential, also read the file's text in a plain-text view before you send it.
For a PDF that has to carry the redaction, export from the cleaned document and then check the PDF the same way: copy-paste from it.
Where this check stops
Colours that come from a style or a theme, rather than being set on the run, are invisible to the script above. Text colour and background can also differ by a hair (for example 010101 on black), which an exact string match misses. A real checker compares contrast, resolves styles, and reads every part of the package. The browser version below also flags automatic font colour on a black highlight, because Word draws that text white while other programs draw it black; the script does not.
I put a browser version of this idea, with those extra cases and a fixed-copy download, at https://hexloomlabs.com/redaction-check/?ref=devto-redaction. It runs in the page and uploads nothing; after it checks your files it reports how many network requests it made, and you can confirm that in your browser's Network tab.
Top comments (0)