Windows hides file extensions by default. So an executable named invoice.pdf.exe shows up in Explorer as invoice.pdf. Give it a PDF icon and there is nothing left to tell it apart by looking.
An extension, though, is only a claim. A PDF normally carries %PDF- at or near the start (many readers accept it anywhere in the first 1024 bytes, which matters later). A Windows PE executable starts with MZ and has a PE header after it. A file can call itself anything it likes, but its first bytes are the contents themselves. They are not proof of anything on their own, but checking their structure against the claim catches inconsistencies that the name alone never shows.
This post is about a Windows app I built that reads those first bytes to determine a file's real type and compares it against the claim (the extension), so you can catch this specific class of disguise before you open the file. It is a pre-open sanity check for files at rest, not an antivirus (more on that boundary in the Limits section). I'll cover how the signature matching works, how it recovers the type of a file whose extension has been removed, how it guesses the programming language of a source file, and the hole I found in the open button right before release.
First published in September 2026, when the app was at v1.8. Revised in October 2026 to match v1.10.2, the version currently in the Store (both of its fixes came out of reviews of this post).
TL;DR
- The extension is a claim; the first bytes are the contents. Compare them and disguises fall out mechanically
- Keep "mismatch" and "danger" separate. "Danger" means a shape that leads to execution, or an executable disguised as something else
- For text with the extension gone, score features × weight × cap to recover the language (47% → 100% on my own 254 files; real-world numbers are in the follow-up)
- "The verdict is right" and "it's safe to open" are different questions. I found one hole in the open button, and one in signature ordering while writing this post. Both are fixed
What I built
Drop a file on the window and you get a card per file: the verdict, whether the claim matches the contents, and how to open it. Drop a whole folder and it checks everything inside, sorted with the most dangerous first.
The screenshot shows four of the sample files that ship with the app. invoice.pdf.exe is flagged as a double-extension disguise, vacation_photo.jpg is flagged because its contents are an executable, and payroll and stats, which have no extension at all, are identified from their contents as COBOL and MATLAB source.
The stack:
- Language: Python
- GUI: tkinter (plus tkinterdnd2 for drag and drop)
- Detection engine: standard library only (
struct,zipfile,json,re,ctypes) - Pillow, but only for the image preview
The engine has no third-party dependencies on purpose. Partly to keep the footprint small, but mostly because I wanted to be able to explain exactly what it checks by pointing at one file. file_kantei.py holds both the GUI and the engine, and the engine functions don't touch the GUI, so tests call them directly.
It runs fully offline and only ever reads files. Nothing about a file (contents or name) is sent anywhere, and checking a file does not count as opening it.
Why not libmagic /
file? libmagic answers "what is this file". The point of this tool is what comes after that: compare the claim with the contents, decide how dangerous the mismatch is, decide whether it's safe to open. Identification is just the front door. Shipping libmagic on Windows also means bundling a DLL and maintaining a magic database, a few dozen formats fit comfortably in the standard library, and I wanted to be able to explain every verdict by pointing at one file.
The source isn't published; the app is distributed through the Microsoft Store only. The snippets in this post are lifted from the actual engine and trimmed for the post (some formats and most of the error handling are left out).
The big picture: three stages
Each file goes through:
- Identify the contents. Match the leading bytes against a signature table. If nothing matches, check whether it is text.
- Compare with the claim. Does the extension expect the format we found? If not, it's a "mismatch". If an executable is dressed up as a document, it's "danger".
- Identity check for executables. For exe/dll: verify the digital signature, read the company/product name it claims, and look for build-tool traces.
The verdict is one of five:
| Verdict | Meaning | Example |
|---|---|---|
| ✅ Match (OK) | Claim and contents agree |
photo.jpg that is a JPEG |
| ⚠ Match (handle with care) | They agree, but the type runs when opened |
setup.exe that is an exe; a macro-enabled .xlsm
|
| ⚠ Caution | Claim and contents differ, or something else needs a look; no execution-related disguise found |
image.png that is actually a JPEG; a .docx with macros; a shortcut that runs a command |
| 🚨 Danger (do not open) | A shape that leads to execution, or an executable in disguise |
vacation_photo.jpg that is an exe; invoice.pdf.exe
|
| ❓ Unidentified | No signature matched and the extension is unknown | Proprietary formats |
Until v1.10.0 the Caution row was labelled "Mismatch". It was renamed in v1.10.1 because the same level now also covers things that aren't a mismatch at all, such as a shortcut that runs a command, or a script hidden with the Windows Script Encoder.
Keeping "mismatch" and "danger" separate is the whole point. A .png that is really a JPEG is a harmless, everyday mismatch (it happens all the time when saving from the web). Call that dangerous and nobody will trust the tool. The word "danger" is reserved for shapes that lead to execution or disguise an executable: the contents are executable (exe/dll/ELF/Flash) and the claim is not, a double extension like invoice.pdf.exe (below), and, since v1.9.0, a shortcut whose target or arguments look like an attack (see Limits).
1. Identifying the contents: signature matching
Start from the first 64 KB
The initial verdict comes from the first 64 KB alone, so its cost doesn't grow with the file size. Some formats then get a follow-up read: OLE files are searched up to 4 MB for stream names (below), the identity lines for executables and PDFs look at the first and last 2 MB, and zips are opened with zipfile, which reads the central directory (the list of entries) from the end of the file and then only the few small entries the check needs, each capped at a fixed number of bytes (the cap was added in v1.10.2; see the zip section). None of these read a whole multi-gigabyte file.
Most signatures are fixed bytes at offset 0, so it is a table walked top to bottom:
_SIGNATURES = [
(b"\x89PNG\r\n\x1a\n", "png"),
(b"\xff\xd8\xff", "jpeg"),
(b"GIF87a", "gif"), (b"GIF89a", "gif"),
(b"PK\x03\x04", "zip"),
(b"Rar!\x1a\x07", "rar"),
(b"7z\xbc\xaf\x27\x1c", "7z"),
(b"\xd0\xcf\x11\xe0\xa1\xb1\x1a\xe1", "ole"), # legacy Office, msi, msg
(b"MZ", "exe"),
(b"\x7fELF", "elf"),
(b"SQLite format 3\x00", "sqlite"),
(b"\x4c\x00\x00\x00\x01\x14\x02\x00", "lnk"),
...
]
A handful of formats are not "fixed bytes at offset 0", and those are handled before the table:
# RIFF container: same 4 bytes, and bytes 8-12 tell WAV / AVI / WebP apart
if head[:4] == b"RIFF":
kind = head[8:12] # b"WAVE" / b"AVI " / b"WEBP"
# ISO Base Media: "ftyp" at offset 4, then the brand: MP4 / MOV / HEIC / AVIF / M4A
if head[4:8] == b"ftyp":
brand = head[8:12] # b"qt " -> mov, b"heic" -> heic, b"avif" -> avif ...
# Executables (MZ + a valid PE header) are settled before any "search past offset 0" check.
# Put this after the PDF scan below and an exe with "%PDF-" in its DOS stub becomes a PDF.
if head[:2] == b"MZ" and _detect_pe_kind(head):
return _detect_pe_kind(head)
# PDF: the spec puts the header on the first line, but readers (Acrobat included)
# accept it anywhere in the first 1024 bytes, and files with a preamble exist
if b"%PDF-" in head[:1024]:
return "pdf"
# TAR: "ustar" lives at offset 257
if head[257:262] == b"ustar":
return "tar"
# ISO: "CD001" lives at offset 0x8001
if head[0x8001:0x8006] == b"CD001":
return "iso"
Digging into containers: telling docx / xlsx / apk apart inside a zip
Everything that starts with PK\x03\x04 is a zip, but the user doesn't want to hear "it's a zip". They want "it's a Word document". docx, xlsx, pptx, jar, apk, msix, epub and odt are all zips, so the tool looks at the entry names inside:
def _detect_zip_kind(path):
with zipfile.ZipFile(path) as zf:
names = zf.namelist()
# epub / odt / ods / odp name themselves in a "mimetype" entry.
# Read it while the zip is still open, and only the first 100 bytes
mime = b""
if "mimetype" in names:
with zf.open("mimetype") as f:
mime = f.read(100)
has_macro = any(n.endswith("vbaProject.bin") for n in names)
if any(n.startswith("word/") for n in names): return "docx", has_macro
if any(n.startswith("xl/") for n in names): return "xlsx", has_macro
if any(n.startswith("ppt/") for n in names): return "pptx", has_macro
if "AndroidManifest.xml" in names: return "apk", has_macro
if "AppxManifest.xml" in names: return "msix", has_macro
if b"epub" in mime: return "epub", has_macro
if b"opendocument.text" in mime: return "odt", has_macro
... # ods / odp the same way
if "META-INF/MANIFEST.MF" in names: return "jar", has_macro
return "zip", has_macro
Why zf.open(...).read(100) and not zf.read(...)[:100]? zf.read decompresses the whole entry before you slice it, so a crafted zip whose mimetype expands to gigabytes would eat the machine's memory. A reviewer of this post pointed out that pattern in an earlier version of this snippet, and the app had it too (in three places, the other two reading Office metadata). Fixed in v1.10.2.
As a side effect, the presence of vbaProject.bin tells us whether the document contains macros. That feeds one extra rule: a .docx / .xlsx / .pptx (extensions that mean no macros) with a macro project inside is raised to "Caution". Office saves macro-enabled files as .docm / .xlsm, so macros in an .xlsx are not something Office produces on its own.
Legacy Office and Outlook .msg: OLE compound files, split by stream name
.doc, .xls, .ppt, .msi and .msg all share one signature (OLE compound file). The stream names inside are stored as UTF-16LE, so the tool searches the first 4 MB for them:
if "WordDocument".encode("utf-16-le") in data: return "doc"
if "Workbook".encode("utf-16-le") in data: return "xls"
if "PowerPoint Document".encode("utf-16-le") in data: return "ppt"
if b"__substg1.0_" in data: return "msg" # Outlook message
exe vs dll: one bit in the PE header
For files starting with MZ, the offset at 0x3C points to PE\0\0, and bit 0x2000 of the Characteristics field that follows says whether it is a DLL:
def _detect_pe_kind(head):
e_lfanew = struct.unpack_from("<I", head, 0x3C)[0]
if head[e_lfanew:e_lfanew + 4] == b"PE\x00\x00":
characteristics = struct.unpack_from("<H", head, e_lfanew + 22)[0]
return "dll" if characteristics & 0x2000 else "exe"
return None
The None case exists because of a false positive I'll get to in a moment.
Is it text? NUL position and control-character ratio
If no signature matches, the tool checks whether the file is text:
def _sniff_text(head):
if head.startswith(b"\xef\xbb\xbf"): encodings = [("utf-8-sig", "UTF-8 (with BOM)")]
elif head.startswith(b"\xff\xfe"): encodings = [("utf-16-le", "UTF-16")]
elif head.startswith(b"\xfe\xff"): encodings = [("utf-16-be", "UTF-16")]
else:
nul = head.find(b"\x00")
if 0 <= nul < 4096: # an early NUL means binary
return None
if nul >= 4096: # a late NUL: judge what comes before it
head = head[:nul]
encodings = [("utf-8", "UTF-8"), ("cp932", "Shift_JIS (Japanese)")]
for enc, label in encodings:
try:
text = head.decode(enc)
except UnicodeDecodeError as e:
if e.start <= len(head) - 8: # broken in the middle: try the next encoding
continue
try: # only a multibyte char cut off at the 64 KB edge:
text = head[:e.start].decode(enc) # still counts as text
except UnicodeDecodeError:
continue
ctrl = sum(1 for ch in text if ord(ch) < 32 and ch not in "\t\r\n")
if text and ctrl / len(text) > 0.05: # more than 5% control characters: binary
return None
return (text, label)
return None # no encoding fits: not text
UTF-8 first, then CP932 (Shift_JIS): if UTF-8 fails partway through, the next encoding gets its turn. I'm in Japan, and a huge number of business CSV files here are still CP932, so that fallback is not optional for me. Swap in your own legacy encoding.
False positives found by sweeping 2,400 real files
Before release I ran roughly 2,400 real files from my own drives through it (PDFs, Office documents, source trees and build output of my own apps, a whole website) and fixed things until there were zero false warnings. Three came up:
An English text file starting with "BM" was identified as a BMP image. The BMP signature is just two bytes, BM, and an English text beginning with "BMW..." matches it. The BMP header has a reserved field at bytes 6-10 that must be zero, so the tool now checks that too.
A memo starting with "MZ" was identified as an executable. Same two-byte problem. If a file starts with MZ but has no PE header and decodes cleanly as text, it is treated as text. That is why _detect_pe_kind returns None.
Legitimate text with a NUL deep inside was identified as binary. Another tool of mine writes HTML reports that embed raw bytes it detected. "Any NUL means binary" rejected those. Now only a NUL within the first 4 KB counts as evidence of binary; a later NUL just truncates what gets examined.
All three are the same lesson: short signatures collide, and real files don't follow the spec.
And one more, found while writing this post. The %PDF- check above searches anywhere in the first 1024 bytes. In the original implementation it ran before the signature table. That means an executable with %PDF- written into its DOS stub (the do-whatever-you-like region from offset 0x40) was identified as a PDF, and named invoice.pdf it came back as "Match (OK)". I noticed it re-reading the code for this post and reproduced it. The fix is one line: if the file is MZ with a valid PE header, it is an executable before any "search past offset 0" rule gets a say. The fixed-offset signatures (RIFF, ftyp, ICO, BMP) can't coexist with MZ, so ordering only matters for the rules that search. Sweeping about 6,000 real files (about 1,000 of them executables) changed the verdict on zero legitimate files.
2. Comparing with the claim
Once the contents are known, the tool looks the extension up in a dictionary (about 180 entries) of what each extension expects:
EXT_DICT = {
"jpg": {"expect": {"jpeg"}, "desc": "JPEG image (photo)"},
"tgz": {"expect": {"gz"}, "desc": "Compressed archive"}, # aliases live in the set
"ai": {"expect": {"pdf"}, "desc": "Illustrator file"}, # contents are PDF
"py": {"expect": {"text"}, "desc": "Python program", "script": True},
...
}
The rules are evaluated in this order, and the first one that fires wins:
double_ext = (len(parts) >= 3 and ext in EXECUTABLE_EXTS and parts[-2] in DOCLIKE_EXTS)
exec_content = content_id in ("exe", "dll", "elf", "swf")
if double_ext: # invoice.pdf.exe
level = "danger"
elif content_id in ext_info["expect"]: # as claimed
level = "ok"
if ext in EXECUTABLE_EXTS or ext_info.get("script"): level = "caution"
if has_macro: level = "caution"
if ext in ("docx", "xlsx", "pptx") and has_macro: level = "warn"
elif exec_content and ext not in EXECUTABLE_EXTS: # executable dressed as a document
level = "danger"
else: # any other mismatch
level = "warn"
Why the double extension is checked first. In invoice.pdf.exe, the extension .exe and the exe contents match. Run the match rule first and it ends as "match (handle with care)". Even though claim and contents agree, the shape document-like extension + executable extension is the classic disguise, so it is declared dangerous before the match rule ever runs.
Why only "executable dressed as a document" is dangerous. As above, a .png holding a JPEG is harmless. But an executable calling itself .jpg, .pdf or .docx has no legitimate reason to exist. Double-clicking it won't run it, but it's the intermediate form of a well-known trick: smuggle it past scanners as a "picture", then rename it and run it later. So the line between "mismatch" and "danger" is not only "would opening it execute something". It is does it lead to execution, or does it disguise an executable.
3. When the extension is gone: classifying text
For binary formats, the signature works whether or not there is an extension. Text is the problem. A Python source file with .py removed and a memo with .txt removed are both "UTF-8 text".
So I built a text classifier.
Feature × weight × cap
For each language, a list of regexes for "things this language looks like", each with a weight and a cap on how many times it may count:
TEXT_KINDS = {
"py": ("py", "Python program", [
(r"^#!.*python", 10, 1),
(r"^\s*def \w+\(.*\)\s*(->.*)?:\s*$", 4, 2),
(r"^\s*(import|from) [\w.]+", 3, 2),
(r"^\s*if __name__ == ", 5, 1),
(r'\A(#[^\n]*\n)*\s*"""', 6, 1), # leading comments then a docstring
...]),
"go": ("go", "Go program", [
(r"^\s*package \w+\s*$", 4, 1), (r"^\s*func \w+\(", 4, 2),
(r"\bfmt\.\w+\(", 4, 1), (r":=", 2, 1), ...]),
...
}
Scoring is just "occurrences (up to the cap) × weight", summed. The cap is there so that a hundred print( calls don't become a hundred points. Let one feature dominate and a file in a different language that happens to use that token a lot will win.
def classify_text(text):
head = text[:24000]
raw = {}
for k, (ext, name, feats, base) in _COMPILED_KINDS.items():
sc = 0
for rx, w, cap in feats:
n = min(sum(1 for _ in rx.finditer(head)), cap)
sc += w * n
raw[k] = sc
...
scores.sort(reverse=True)
best, kind = scores[0]
second = scores[1][0] if len(scores) > 1 else 0
if best >= 7 and best - second >= 3:
return kind, "high"
if best >= 4 and best > second:
return kind, "mid"
...
Confidence has two levels. High (7+ points, and 3 clear of second place) lets the tool say "Original extension: .py. Add '.py' to the end of the name and it opens as usual." Mid (4+ points, sole leader) only gets "the contents look like a Python program", with no firm claim.
The window is the first 24,000 characters. Shorter windows missed files that begin with a long license header or docstring: the features never made it into the window and nothing scored.
Derived languages: TypeScript sits on top of JavaScript
A TypeScript file has every JavaScript feature (const, =>, console.log). Score JS and TS independently and a TS file ends up "JS 9, TS 9": a tie, no sole leader, no verdict at all.
So there is a notion of a derived kind. TS derives from JS; so does Google Apps Script. C++ derives from C, SCSS from CSS.
"ts": ("ts", "TypeScript program", [
(r"\w+\s*:\s*(string|number|boolean|any|void|unknown)(\[\])?\s*[,;)=|{]", 4, 2),
(r"^\s*(export\s+)?interface \w+\s*\{", 4, 1),
...], "js"), # <- fourth element: the base kind
A derived kind is only a candidate if at least one of its own features fires, and its score is then "base score + own score". One type annotation or interface and TS beats JS for certain; none, and TS isn't even in the race, so JS wins cleanly. If two derived kinds tie on the same base (GAS and TS both on JS), the tool returns the shared base, JS, at mid confidence.
Benchmark: 47% → 100% (on my own 254 files)
Accuracy is measured by taking real files from my drives, copying them with the extension stripped, running the tool, and checking whether it recovers the original extension. 254 files, 28 kinds.
| Stage | Accuracy | What changed |
|---|---|---|
| First version | 47.1% | — |
| Derived kinds | 84.6% | JS/TS/GAS ties were erasing the verdict |
| Narrower TS features | 93.3% | A comment saying "type: string" counted as a type annotation. Now only identifier: type followed by a delimiter counts |
| YAML fix | 97.2% | Lines of Japanese prose like "Usage: ..." were counted as YAML keys. Keys must be ASCII now |
| Markdown fix | 100% | A plain memo with bullet points was classified as Markdown. Bullets are capped, and only things you only write in Markdown (front matter, wiki links) count as evidence |
The last two fixes are about not calling an ordinary note "some language". A missed detection costs less trust than a false one. Files whose correct answer is "this was just plain text" are part of the benchmark too.
That 100% is on my own 254 files. I found errors, fixed the rules, and re-measured on the same 254 files, so it is not accuracy on independent data. When I later ran the same classifier over real files already on my PC, text classification scored 61.0% on 2,621 files. After fixes (v1.10.0) it reached 83.6%, but that number was also tuned while looking at those files. The miss rate on malicious files has not been measured. How I measured it, and what was broken, is in the follow-up: I ran my file-type detector over 8,900 real files. It was wrong in ways no test suite would have told me
Never put heavy weight on a generic feature
A trap I fell into when expanding to about 40 languages. I gave MATLAB's function name( five points, and the existing JavaScript tests broke, because JavaScript has function name( as well.
"matlab": ("m", "MATLAB/Octave program", [
# Only "function [out] = name(" is strong; a bare "function name(" looks like JS, so it's weak
(r"^\s*function \[?[\w, ]+\]?\s*=\s*\w+\(", 5, 1),
(r"^\s*function \w+\(", 1, 1),
(r"^\s*(clear all|clc|close all)\b", 6, 1),
...]),
Shapes shared across languages (function, a trailing ;, end) get one point. Only shapes unique to that language (function [out] = name(, clear all) get heavy weight. This is why every new language means re-running the tests for all languages.
A middle answer: "a program, language unknown"
A language that isn't in the table (or a short script with few features) scores nothing anywhere and falls through to "plain text". But a human glancing at if (...) {, return, trailing ; and // comments would say "that's code of some kind".
So there is a kind called code made of nothing but language-agnostic features. It never competes normally; only when nothing else reaches a verdict, and it has 10+ points, does the tool answer "a program (language not identified)". No language name, but enough of a warning that "adding an extension and double-clicking might execute this".
4. Identity check for executables: Authenticode, offline
When the contents are exe/dll, the tool calls the Windows APIs through ctypes and checks three things.
Digital signature. WinVerifyTrust. The important part is that it never goes online for revocation checks:
data = TRUST_DATA(
...,
2, # WTD_UI_NONE (no dialogs)
0, # no revocation checks
1, # WTD_CHOICE_FILE
...,
0x1000, # WTD_CACHE_ONLY_URL_RETRIEVAL (no network)
...)
rc = wintrust.WinVerifyTrust(None, byref(action), byref(data)) & 0xFFFFFFFF
The return code is collapsed into five states for the user:
| Return code | Shown as |
|---|---|
| 0 | Signed (valid) + signer name |
TRUST_E_NOSIGNATURE etc. |
No embedded signature → if the file is in the OS catalog (CryptCATAdmin*), "Windows catalog signature"; otherwise "unsigned" |
TRUST_E_BAD_DIGEST |
Signature does not match the contents (possible tampering) → verdict raised to "Caution" |
CERT_E_EXPIRED |
Expired |
| anything else | Untrusted publisher |
Being unsigned is not, on its own, a reason to flag a file. My own apps are unsigned. But "signed, and the contents don't match the signature" means a legitimate file was modified afterwards, and that is the one state that raises the verdict.
The claim. GetFileVersionInfoW gives the company name, product name and version. For an unsigned file that is self-reported, and the tool labels it "(self-reported)". An unsigned exe claiming to be from Microsoft is easy to spot once it's phrased that way.
The build. From the PE header: 64/32-bit, GUI or console subsystem, .NET or not (data directory 14, the COM descriptor). From markers in the first and last 2 MB: PyInstaller, Inno Setup, NSIS. "A Python (PyInstaller) GUI app, 64-bit, unsigned" is a useful thing to know when you ask the sender what they sent you.
5. The hole in the open button
Files with a "Match (OK)" verdict got an open button, labelled "📂 Safe: open it" at the time, that handed them to the usual application. For a file with no extension, the tool made a copy in a temp folder with the estimated extension added, and opened that (the original is never touched).
Right before release, while clicking through the sample folder, I noticed: strip the extension from a batch file, check it, and you get "Batch file (estimated .bat), Match (OK)", complete with the "Safe: open it" button. Click it, and a copy named *.bat goes to os.startfile. Which runs it.
The verdict was correct. There's no claim, so there's nothing to mismatch. But "open" is a different question from "what is it", and it has to be decided by what happens when you open it:
def _is_program_like(ext, text_kind):
words = ("program", "script", "executable")
if ext in EXECUTABLE_EXTS:
return True
info = EXT_DICT.get(ext or "")
if info and (info.get("script") or any(w in info["desc"] for w in words)):
return True
kind = TEXT_KINDS.get(text_kind or "")
if kind and any(w in kind[1] for w in words):
return True
return False
def can_open_safely(r):
if r.get("level") != "ok":
return False
ext = r.get("ext") or (r.get("guess_ext") or "") # no extension: check the estimated one
if not ext:
return False # no estimate either (Makefile etc.): nothing to open it with
return not _is_program_like(ext, r.get("text_kind"))
Rather than maintaining a hand-written list of executable extensions, the check scans the dictionary's own descriptions for the words "program", "script" and "executable". A hand-written list will be forgotten the next time a language is added. When there's no extension, the estimated one is checked, and the text classification result is checked as well.
No automated test caught this. There were tests for "is the verdict right?", but none for "what happens when the button is pressed?". It was found the boring way: build a sample folder, and press every button yourself before shipping.
The label was a promise too. A format that matches its extension is not proof that the file is harmless (a macro-free .docx can still carry a phishing link), so "Safe" claimed more than the check can back up. In v1.10.2 the button was renamed to "📂 Open with default app". It says what the button does, not what it guarantees.
Limits: this is not an antivirus
What it can't do:
It doesn't judge intent. It can't tell you whether an unsigned exe is safe or malicious. It doesn't read what a macro does. It doesn't look at JavaScript inside a PDF. What it knows is "what is this", "does the claim match", and "who signed it". It is not a replacement for antivirus software; it sits in front of it, answering "should I be opening this at all?".
It inspects files at rest only. (Added 2026-09-19, prompted by a reader's comment below.) It looks at headers, container structure, and extension/signature consistency. It does not observe anything that happens after a file is opened or executed. Process injection (malicious code written into the memory of a legitimate, signed process such as explorer.exe and run there), reflective DLL loading (a DLL loaded straight from memory without touching the disk), fileless delivery through scripts or macros (PowerShell, WMI, living-off-the-land binaries), and similar runtime techniques are entirely outside its reach, because in those cases there is either no file to inspect or the file on disk is genuinely benign. Watching for that is the job of EDR and the behavioral side of antivirus: API hooking, suspicious call chains (Word calling VirtualAllocEx into another process, an Office app spawning powershell), memory scanning. Windows Defender's Attack Surface Reduction rules cover some of it, such as blocking Office apps from creating child processes. This tool is a pre-open sanity check, not a replacement for antivirus or EDR, and I am deliberately not adding a resident monitor to it: that needs permissions and a threat model I am not equipped to get right, and a half-built one would be worse than none.
That said, even fileless attacks usually have a delivery artifact that is a file on disk, and the classic one is a shortcut (.lnk) wearing a document icon that launches powershell in a hidden window. That is still a file-at-rest check, so v1.9.0 added it. Shortcuts are parsed minimally as Shell Link (MS-SHLLINK): LinkInfo and StringData give the target, the arguments and the borrowed icon, and all three are shown on the card. A target of mshta / wscript / rundll32 and friends with any arguments, powershell / cmd with arguments that hide the window (-w hidden), pass the command encoded so it can't be read at a glance (-enc), or download and run code (DownloadString to fetch it, IEX to execute the resulting string), or a name like invoice.pdf.lnk whose target is not that document, is marked Danger. The real-file sweep earned its keep again here: Windows itself creates invoice.pdf.lnk entries in the Recent folder, so a naive double-extension rule produced 18 false alarms on my own PC, and the rule became "does the target match the document the name claims to be". Across 278 real shortcuts on this machine: 0 Danger, 9 Caution, all of them genuine command-running developer shortcuts. The same release also reads the Mark of the Web (the NTFS Zone.Identifier stream that browsers and mail clients attach) and shows "Source: downloaded from the internet". Even so, the tool sees the delivery artifact and nothing beyond it.
It doesn't look inside zips. A zip is "a zip (a Word document if the entries say so)". Files inside it aren't checked until you extract them, and for an encrypted zip only the entry names are visible.
Formats without a signature can't be identified. Proprietary formats come back as "unidentified". For those there is a "🌐 search the web" button that opens your default browser with a search for the extension (only the extension string is passed; never the file's contents or name).
Polyglots. A file crafted to satisfy two signatures at once (a valid PDF that is also a valid ZIP, for example) is reported as whichever format is checked first. The tool targets commodity disguises, not files built specifically to evade it.
Text classification is a guess. Short scripts, mixed languages and sparse code will miss. That's why confidence has two levels and the mid level never asserts.
The regression suite is 66 engine tests, 50 text-classification tests, 79 language tests, 35 media-info tests, 35 help tests, 19 English-UI tests, 15 GUI tests and 71 shortcut / Mark-of-the-Web tests, and all of it runs every time a language is added.
Takeaways
- The extension is a claim; the first bytes are the contents. Compare the two and disguises fall out mechanically
- Keep "mismatch" and "danger" separate. "Danger" means a shape that leads to execution, or an executable disguised as something else
- A double extension is dangerous even though claim and contents agree. Check it before the match rule
- Short signatures (
BM,MZ) collide. Sweep thousands of real files and fix until there are zero false warnings - For text with the extension gone: score features × weight × cap. Put derived languages on top of their base. Never give a generic shape heavy weight. Don't call an ordinary note "some language"
- "The verdict is right" and "it's safe to open" are different questions. Decide the open button by what happens when it's pressed
The tool is on the Microsoft Store. Everything it checks is free, and it runs fully offline.
https://apps.microsoft.com/detail/9PKG5KT1WXR8?hl=en-us&gl=US
If there's a format it doesn't recognize or a language you'd like added, leave a comment. New signatures and features are welcome.
About the author
Okinawa Software Lab. I lead in-house digital transformation at a small company in Okinawa, Japan. I build the tools we need ourselves, and I publish Windows apps on the Microsoft Store that follow the same principle: everything happens on your own PC.
- Website: https://okinawasoftwarelab.com/en/

Top comments (4)
Hi Okinawa Software Lab
Really enjoyed the write-up — the signature-matching walkthrough (especially the DOS-stub %PDF- bug and the double-extension ordering fix) is a great example of hardening through real-file sweeps rather than spec-reading alone.
One thing worth flagging clearly for readers, though: this class of tool (and honestly, most static/signature-based AV) only covers file-at-rest disguises. A meaningful share of real-world malware today never needs a mismatched extension at all, because it doesn't rely on the file being executed directly:
None of this is a knock on the engineering — the double-extension and container-sniffing work is genuinely solid for what it targets. But I think the framing ("catching invoice.pdf.exe before you open it") could undersell to readers how narrow the protection is. Modern EDR products lean on behavioral/runtime monitoring — API-call hooking, watching for suspicious call chains (e.g. a Word process suddenly calling VirtualAllocEx into another process), memory scanning for shellcode patterns — precisely because file inspection can't reach injection or fileless execution.
Might be worth adding an explicit line in the Limits section along those lines, so users don't walk away thinking this covers more of the threat landscape than it does. Happy to be wrong here — curious if you've thought about any lightweight runtime signals (even something as simple as flagging known LOLBins being spawned by Office processes) as a v2 direction.
Hi,
Thank you for taking the time to write this up — and for reading the article closely enough to pick out the DOS-stub %PDF- case and the double-extension ordering fix. Those were exactly the parts I was least sure anyone would notice.
You're right, and I don't think you're wrong on any of it. I should be upfront that I come to this from a non-security background: I build small tools for office staff at an automotive repair business, and this one grew out of a very specific problem — people receiving "invoice.pdf.exe" style attachments and having no way to check them without opening them. So the tool's scope was always "what can be learned about a file at rest, before it's opened." It was never meant to be an AV or EDR substitute, but I agree the current framing doesn't make that boundary clear enough, and a reader could easily take away more than the tool actually delivers.
I'll add an explicit paragraph to the Limits section, roughly:
"This tool inspects files at rest only — headers, container structure, and extension/signature consistency. It does not observe anything that happens after a file is opened or executed. Process injection, reflective loading, fileless delivery through scripts or macros, and similar runtime techniques are entirely outside its reach, because in those cases there is either no file to inspect or the file on disk is genuinely benign. It is a pre-open sanity check, not a replacement for antivirus or EDR."
I'll also soften the "catch it before you open it" wording so it reads as "catch this specific class of disguise," not "catch malware."
On the v2 question — I've thought about it, and my honest answer is that I'm going to deliberately not go there. Watching for things like an Office process spawning a LOLBin would need a resident component subscribing to process/ETW events, and that pulls in permissions and a threat model I'm not equipped to get right. It also cuts against the design constraints all of my tools share: fully offline, no background process, never modifies the original file. A half-built runtime monitor seems worse than none — it would either generate false positives that erode trust, or give users a false sense of coverage in exactly the area you're describing. That job belongs to proper EDR products, and I'd rather say so plainly than blur the line.
If you're willing, I'd appreciate a second look at the revised Limits wording once it's up. Thanks again — this is the kind of feedback that makes the tool more honest, which matters more to me than making it sound bigger.
Best regards,
Okinawa Software Lab
単にこの分野に一定の関心があり、経験があるというだけです。
セキュリティ分野にはまだ取り組むべき課題が数多くあり、最近ではAIの急速な発展に伴い、「AIセキュリティ」という新たな領域へと拡大しています。
私は日本語もネイティブレベルで話すことができます。
もしよろしければ、カジュアルな面談を通じて、経験の共有や今後の展望についてお話しできればと考えています。
中林さん
お返事ありがとうございます。いただいたご指摘は記事とアプリの次の版に反映しました。改めて感謝します。
面談のお誘いは、申し訳ありませんが今回は見送らせてください。記事やアプリについてお気づきの点があれば、引き続きこのコメント欄で教えていただけると助かります。
Okinawa Software Lab